Senior Site Reliability Engineer
New
J
JobgetherIT infrastructure
Based in United StatesFull-TimeSenior
Salary105,000 - 140,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- At least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role
- Required Skills
- AWSPythonTerraform
Requirements
- Have at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.
- Have hands-on expertise with Windows Server 2016 or later, Active Directory, IIS, and Microsoft SQL.
- Have cloud infrastructure skills; AWS experience is preferred.
- Have advanced scripting capabilities, including reusable modules and integrations with REST APIs.
- Have hands-on experience with Terraform and Puppet or Chef.
- Have experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or New Relic.
- Understand networking fundamentals, including DNS, TCP/IP, load balancing, and VPN technologies.
- Have a Bachelor's degree in Computer Science, Information Technology, or a related discipline, or equivalent professional experience; the degree is preferred.
- Relevant AWS, Microsoft, or HashiCorp Terraform certifications are desirable.
- Experience with Docker, Kubernetes, and CI/CD tools such as GitLab Pipelines, Jenkins, or GitHub Actions is a plus.
- Knowledge of security best practices, compliance frameworks, and log aggregation and analysis tools such as the ELK Stack or Splunk is desirable.
Responsibilities
- Design, implement, and maintain scalable production infrastructure using Infrastructure as Code practices.
- Develop automation scripts for operating system provisioning, configuration management, and recurring operational tasks.
- Implement and manage configuration management solutions across hybrid infrastructure environments.
- Monitor system health, performance, reliability, and availability using observability tools and operational practices.
- Establish, maintain, and enforce SLAs, SLOs, and error budgets for production services.
- Participate in an on-call rotation, respond to production incidents, restore services, and conduct root cause analysis.
- Partner with development teams to improve deployment pipelines, release processes, and service reliability.
- Create and maintain operational documentation, including procedures, runbooks, and architectural decisions.
- Lead or contribute to post-mortem reviews and implement corrective actions to prevent recurring incidents.
- Troubleshoot complex infrastructure and application issues across multiple technology layers.
View Full Description & ApplyYou'll be redirected to the employer's site