Site Reliability Engineer
New
J
JobgetherSoftware, SaaS
Based in United StatesFull-TimeSenior
Salary$100,000 - $120,000 per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSSQLJavascriptC#Azure.NETCI/CD
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 5+ years of professional experience in a Site Reliability Engineering role.
- Strong understanding of SRE principles, reliability practices, and incident management processes.
- Experience with observability platforms such as New Relic, Datadog, Sumo Logic, or similar tools.
- Proficiency reading and writing code using technologies such as JavaScript, .NET, and SQL.
- Experience troubleshooting C#/.NET web applications and identifying performance issues.
- Familiarity with cloud platforms such as AWS or Azure, with AWS experience strongly preferred.
- Experience operating within public cloud environments and SaaS platforms.
- Strong understanding of cloud architecture patterns and operational best practices.
- Knowledge of CI/CD pipelines and Infrastructure as Code (IaC) environments is preferred.
- Experience troubleshooting Windows environments and SQL Server is preferred.
- Strong analytical, problem-solving, and data-driven decision-making skills.
Responsibilities
- Lead post-incident investigations and perform detailed root cause analyses to identify failures and prevent recurring issues.
- Create clear and actionable Root Cause Analysis (RCA) documentation for internal teams and customer delivery.
- Develop and implement preventative strategies to improve system reliability and reduce operational disruptions.
- Monitor and improve key reliability metrics, including time to resolution and incident response effectiveness.
- Configure and maintain observability tools to ensure accurate monitoring, alerting, and performance visibility.
- Build client-focused dashboards and alerts to proactively identify application and platform performance challenges.
- Collaborate with Engineering, Cloud Operations, and SRE teams to implement improvements that enhance scalability and stability.
- Contribute to automation initiatives that streamline incident response and operational workflows.
View Full Description & ApplyYou'll be redirected to the employer's site