Senior Site Reliability Engineer

New
J
JobgetherIT infrastructure
Based in United StatesFull-TimeSenior
Salary105,000 - 140,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
At least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role
Required Skills
AWSPythonTerraform

Requirements

  • Have at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.
  • Have hands-on expertise with Windows Server 2016 or later, Active Directory, IIS, and Microsoft SQL.
  • Have cloud infrastructure skills; AWS experience is preferred.
  • Have advanced scripting capabilities, including reusable modules and integrations with REST APIs.
  • Have hands-on experience with Terraform and Puppet or Chef.
  • Have experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or New Relic.
  • Understand networking fundamentals, including DNS, TCP/IP, load balancing, and VPN technologies.
  • Have a Bachelor's degree in Computer Science, Information Technology, or a related discipline, or equivalent professional experience; the degree is preferred.
  • Relevant AWS, Microsoft, or HashiCorp Terraform certifications are desirable.
  • Experience with Docker, Kubernetes, and CI/CD tools such as GitLab Pipelines, Jenkins, or GitHub Actions is a plus.
  • Knowledge of security best practices, compliance frameworks, and log aggregation and analysis tools such as the ELK Stack or Splunk is desirable.

Responsibilities

  • Design, implement, and maintain scalable production infrastructure using Infrastructure as Code practices.
  • Develop automation scripts for operating system provisioning, configuration management, and recurring operational tasks.
  • Implement and manage configuration management solutions across hybrid infrastructure environments.
  • Monitor system health, performance, reliability, and availability using observability tools and operational practices.
  • Establish, maintain, and enforce SLAs, SLOs, and error budgets for production services.
  • Participate in an on-call rotation, respond to production incidents, restore services, and conduct root cause analysis.
  • Partner with development teams to improve deployment pipelines, release processes, and service reliability.
  • Create and maintain operational documentation, including procedures, runbooks, and architectural decisions.
  • Lead or contribute to post-mortem reviews and implement corrective actions to prevent recurring incidents.
  • Troubleshoot complex infrastructure and application issues across multiple technology layers.
View Full Description & ApplyYou'll be redirected to the employer's site
105,000 - 140,000 USD per year
Apply Now