Site Reliability Engineer

New
J
JobgetherSoftware, AI
Based in United States, able to work effectively across relevant time zonesFull-TimeSenior
SalaryCompetitive compensation package with an organization-wide goal-based bonus program.
Apply NowOpens the employer's application page

Job Details

Languages
English
Experience
5+ years of experience building and operating modern infrastructure systems for senior-level candidates, or 8+ years for staff-level candidates.
Required Skills
AWSPostgreSQLElasticSearchKubernetesMongoDBGrafanaTerraformAnsibleGitHub Actions

Requirements

  • 5+ years of experience building and operating modern infrastructure systems for senior-level, or 8+ years for staff-level.
  • Strong experience with AWS cloud platforms.
  • Hands-on expertise with containerized environments, specifically Kubernetes.
  • Proficiency with infrastructure automation tools including Terraform and Ansible.
  • Experience with CI/CD technologies such as GitHub Actions or ArgoCD.
  • Experience managing production databases and data platforms such as MongoDB, PostgreSQL, or Elasticsearch.
  • Strong understanding of observability and monitoring tools like Grafana, Prometheus/Mimir, Loki, Tempo, or OpenTelemetry.
  • Experience supporting mission-critical production environments, including incident response and disaster recovery.
  • Solid understanding of networking concepts and protocols including DNS, HTTP, and TCP.
  • Ability to design simple, maintainable, and resilient systems.
  • Strong written and verbal communication skills in English.
  • Based in the United States and able to work effectively across relevant time zones.

Responsibilities

  • Automate database lifecycle management, including provisioning, scaling, failover processes, and retirement strategies.
  • Improve infrastructure security by reducing reliance on static credentials and implementing identity-based authentication.
  • Enhance system reliability by minimizing downtime, improving deployment processes, and strengthening disaster recovery.
  • Design and maintain multi-region recovery strategies to ensure fast and reliable failover.
  • Own and improve observability systems, including telemetry pipelines and monitoring platforms.
  • Maintain and optimize CI/CD workflows for efficient, reliable, and secure software delivery.
  • Manage cloud infrastructure and automation using Kubernetes, Terraform, and Ansible.
  • Act as a senior escalation point for production incidents and drive long-term solutions.
  • Collaborate with engineering teams to improve architecture, operational practices, and platform scalability.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive compensation package with an organization-wide goal-based bonus program.
Apply Now