Senior Software Engineer, Site Reliability & Security

New
J
JobgetherAI Infrastructure
United StatesFull-TimeSenior
SalaryCompetitive salary and stock options.
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSDockerPythonGCPKubernetesTypeScriptGoPrometheusTerraform

Requirements

  • 5+ years of professional experience working with a modern programming language such as Go, TypeScript, Python, or a comparable language.
  • Strong understanding of computer science fundamentals, including algorithms, data structures, systems design, and distributed systems.
  • Hands-on experience with cloud environments such as AWS or GCP and container technologies including Kubernetes and Docker.
  • Strong Infrastructure as Code experience with tools such as Terraform, OpenTofu, or Pulumi.
  • Experience with GitOps and continuous delivery platforms such as ArgoCD, Spacelift, Terraform Cloud, or comparable technologies.
  • Demonstrated ability to build, troubleshoot, and operate Kubernetes environments in production.
  • Strong debugging and problem-solving skills, particularly in complex production environments.
  • Experience profiling applications and databases to identify and resolve performance and scalability issues.
  • Fluency with UNIX/Linux command-line environments for log analysis, troubleshooting, and operational tasks.
  • Experience with monitoring and observability tools such as Grafana and Prometheus.
  • Strong understanding of site reliability, high availability, redundancy, incident management, and production operations.
  • Ability to communicate clearly and collaborate effectively with engineering teams in a remote environment.

Responsibilities

  • Build, operate, and continuously improve monitoring, tracing, alerting, and observability infrastructure to maintain strong system reliability and visibility.
  • Drive platform security initiatives, with an emphasis on preventative controls, resilient architecture, and protecting production systems.
  • Lead incident response and recovery efforts, including troubleshooting, root-cause analysis, and implementation of measures that prevent recurring issues.
  • Design and maintain highly available, redundant, and scalable infrastructure capable of supporting resource-intensive applications and growing workloads.
  • Participate in on-call operations, respond to production alerts, and help ensure critical services remain reliable and performant.
  • Improve deployment and release processes to make code changes faster, simpler, safer, and more consistent.
  • Develop platform capabilities that enable engineering teams to deliver products efficiently while maintaining high standards for reliability and security.
  • Identify innovative approaches to load management, resource utilization, performance optimization, and system scalability.
  • Partner with software engineers to promote reliable coding practices and help teams balance rapid delivery with operational excellence.
  • Research and implement improvements to infrastructure, security, automation, and engineering processes as the platform evolves.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive salary and stock options.
Apply Now