Senior Site Reliability Engineer – Telephony & Communications Platform

New
J
JobgetherCloud Infrastructure
Based in United StatesFull-TimeSenior
Salary175,000 - 195,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
8+ years of experience in software engineering, infrastructure engineering, or operations; 4+ years of hands-on Site Reliability Engineering experience
Required Skills
AWSPythonBashCI/CDTerraformDistributed Systems

Requirements

  • 8+ years of experience in software engineering, infrastructure engineering, or operations roles.
  • 4+ years of hands-on Site Reliability Engineering experience supporting production environments.
  • Strong experience with AWS services, including EC2, ECS/EKS, VPC, Route 53, IAM, Lambda, S3, and CloudWatch.
  • Advanced knowledge of Terraform, Infrastructure as Code, CI/CD pipelines, observability practices, and automation.
  • Strong scripting skills with languages such as Python, Bash, or PowerShell.
  • Proven experience supporting highly available, low-latency, and scalable production systems.
  • Strong understanding of distributed systems, cloud architecture, and reliability best practices.
  • Ability to troubleshoot complex technical issues and drive solutions independently.
  • Strong communication skills with the ability to collaborate effectively with engineering teams.
  • Experience with telephony or communication platforms such as Amazon Connect, Twilio, SIP/RTP, or other VoIP technologies is a plus.
  • Familiarity with voice quality metrics such as MOS, jitter, and packet loss is preferred.
  • Experience working in regulated environments such as SOC 2, HIPAA, or FedRAMP is a plus.
  • Exposure to Azure or Google Cloud Platform environments is beneficial.

Responsibilities

  • Own the reliability, scalability, and performance of key platform services, ensuring highly available production systems.
  • Design, build, and maintain AWS infrastructure using Terraform and Infrastructure as Code principles.
  • Develop and improve CI/CD pipelines, automation frameworks, monitoring solutions, and operational tooling.
  • Define and maintain reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
  • Lead incident response activities, perform root cause analysis, and implement long-term reliability improvements.
  • Improve observability through logging, metrics, alerting, and proactive system monitoring.
  • Collaborate with engineering teams to identify operational risks and implement scalable solutions.
  • Participate in a 24/7 on-call rotation and contribute to production support excellence.
  • Mentor engineers and promote best practices in reliability engineering, automation, and infrastructure management.
  • Support communications and telephony platform reliability, including improvements related to voice infrastructure when applicable.
View Full Description & ApplyYou'll be redirected to the employer's site
175,000 - 195,000 USD per year
Apply Now