Senior Site Reliability Engineer – Telephony & Communications Platform
New
J
JobgetherCloud Infrastructure
Based in United StatesFull-TimeSenior
Salary175,000 - 195,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of experience in software engineering, infrastructure engineering, or operations; 4+ years of hands-on Site Reliability Engineering experience
- Required Skills
- AWSPythonBashCI/CDTerraformDistributed Systems
Requirements
- 8+ years of experience in software engineering, infrastructure engineering, or operations roles.
- 4+ years of hands-on Site Reliability Engineering experience supporting production environments.
- Strong experience with AWS services, including EC2, ECS/EKS, VPC, Route 53, IAM, Lambda, S3, and CloudWatch.
- Advanced knowledge of Terraform, Infrastructure as Code, CI/CD pipelines, observability practices, and automation.
- Strong scripting skills with languages such as Python, Bash, or PowerShell.
- Proven experience supporting highly available, low-latency, and scalable production systems.
- Strong understanding of distributed systems, cloud architecture, and reliability best practices.
- Ability to troubleshoot complex technical issues and drive solutions independently.
- Strong communication skills with the ability to collaborate effectively with engineering teams.
- Experience with telephony or communication platforms such as Amazon Connect, Twilio, SIP/RTP, or other VoIP technologies is a plus.
- Familiarity with voice quality metrics such as MOS, jitter, and packet loss is preferred.
- Experience working in regulated environments such as SOC 2, HIPAA, or FedRAMP is a plus.
- Exposure to Azure or Google Cloud Platform environments is beneficial.
Responsibilities
- Own the reliability, scalability, and performance of key platform services, ensuring highly available production systems.
- Design, build, and maintain AWS infrastructure using Terraform and Infrastructure as Code principles.
- Develop and improve CI/CD pipelines, automation frameworks, monitoring solutions, and operational tooling.
- Define and maintain reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Lead incident response activities, perform root cause analysis, and implement long-term reliability improvements.
- Improve observability through logging, metrics, alerting, and proactive system monitoring.
- Collaborate with engineering teams to identify operational risks and implement scalable solutions.
- Participate in a 24/7 on-call rotation and contribute to production support excellence.
- Mentor engineers and promote best practices in reliability engineering, automation, and infrastructure management.
- Support communications and telephony platform reliability, including improvements related to voice infrastructure when applicable.
View Full Description & ApplyYou'll be redirected to the employer's site