Site Reliability Engineer
New
J
JobgetherSoftware, AI
Based in United States, able to work effectively across relevant time zonesFull-TimeSenior
SalaryCompetitive compensation package with an organization-wide goal-based bonus program.
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 5+ years of experience building and operating modern infrastructure systems for senior-level candidates, or 8+ years for staff-level candidates.
- Required Skills
- AWSPostgreSQLElasticSearchKubernetesMongoDBGrafanaTerraformAnsibleGitHub Actions
Requirements
- 5+ years of experience building and operating modern infrastructure systems for senior-level, or 8+ years for staff-level.
- Strong experience with AWS cloud platforms.
- Hands-on expertise with containerized environments, specifically Kubernetes.
- Proficiency with infrastructure automation tools including Terraform and Ansible.
- Experience with CI/CD technologies such as GitHub Actions or ArgoCD.
- Experience managing production databases and data platforms such as MongoDB, PostgreSQL, or Elasticsearch.
- Strong understanding of observability and monitoring tools like Grafana, Prometheus/Mimir, Loki, Tempo, or OpenTelemetry.
- Experience supporting mission-critical production environments, including incident response and disaster recovery.
- Solid understanding of networking concepts and protocols including DNS, HTTP, and TCP.
- Ability to design simple, maintainable, and resilient systems.
- Strong written and verbal communication skills in English.
- Based in the United States and able to work effectively across relevant time zones.
Responsibilities
- Automate database lifecycle management, including provisioning, scaling, failover processes, and retirement strategies.
- Improve infrastructure security by reducing reliance on static credentials and implementing identity-based authentication.
- Enhance system reliability by minimizing downtime, improving deployment processes, and strengthening disaster recovery.
- Design and maintain multi-region recovery strategies to ensure fast and reliable failover.
- Own and improve observability systems, including telemetry pipelines and monitoring platforms.
- Maintain and optimize CI/CD workflows for efficient, reliable, and secure software delivery.
- Manage cloud infrastructure and automation using Kubernetes, Terraform, and Ansible.
- Act as a senior escalation point for production incidents and drive long-term solutions.
- Collaborate with engineering teams to improve architecture, operational practices, and platform scalability.
View Full Description & ApplyYou'll be redirected to the employer's site