Senior Software Engineer, Site Reliability & Security
New
J
JobgetherAI Infrastructure
United StatesFull-TimeSenior
SalaryCompetitive salary and stock options.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSDockerPythonGCPKubernetesTypeScriptGoPrometheusTerraform
Requirements
- 5+ years of professional experience working with a modern programming language such as Go, TypeScript, Python, or a comparable language.
- Strong understanding of computer science fundamentals, including algorithms, data structures, systems design, and distributed systems.
- Hands-on experience with cloud environments such as AWS or GCP and container technologies including Kubernetes and Docker.
- Strong Infrastructure as Code experience with tools such as Terraform, OpenTofu, or Pulumi.
- Experience with GitOps and continuous delivery platforms such as ArgoCD, Spacelift, Terraform Cloud, or comparable technologies.
- Demonstrated ability to build, troubleshoot, and operate Kubernetes environments in production.
- Strong debugging and problem-solving skills, particularly in complex production environments.
- Experience profiling applications and databases to identify and resolve performance and scalability issues.
- Fluency with UNIX/Linux command-line environments for log analysis, troubleshooting, and operational tasks.
- Experience with monitoring and observability tools such as Grafana and Prometheus.
- Strong understanding of site reliability, high availability, redundancy, incident management, and production operations.
- Ability to communicate clearly and collaborate effectively with engineering teams in a remote environment.
Responsibilities
- Build, operate, and continuously improve monitoring, tracing, alerting, and observability infrastructure to maintain strong system reliability and visibility.
- Drive platform security initiatives, with an emphasis on preventative controls, resilient architecture, and protecting production systems.
- Lead incident response and recovery efforts, including troubleshooting, root-cause analysis, and implementation of measures that prevent recurring issues.
- Design and maintain highly available, redundant, and scalable infrastructure capable of supporting resource-intensive applications and growing workloads.
- Participate in on-call operations, respond to production alerts, and help ensure critical services remain reliable and performant.
- Improve deployment and release processes to make code changes faster, simpler, safer, and more consistent.
- Develop platform capabilities that enable engineering teams to deliver products efficiently while maintaining high standards for reliability and security.
- Identify innovative approaches to load management, resource utilization, performance optimization, and system scalability.
- Partner with software engineers to promote reliable coding practices and help teams balance rapid delivery with operational excellence.
- Research and implement improvements to infrastructure, security, automation, and engineering processes as the platform evolves.
View Full Description & ApplyYou'll be redirected to the employer's site