Site Reliability Engineer (SRE) – Night Shift

P
PeratonNational Security
United States, 11pm – 7am Eastern Standard Time (EST)Full-TimeSenior
Salary$104,000 - $166,000 / year
Apply NowOpens the employer's application page

Job Details

Experience
Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience; 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
Required Skills
AWSPythonBashJenkinsKubernetesLinuxTerraformAnsibleGitLab

Requirements

  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms.
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration.
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk, and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets and alert rationalization.
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

Responsibilities

  • Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, and security.
  • Monitor services using SLIs, SLOs, dashboards, and alerts, and continuously improve detection and resolution of operational issues.
  • Partner with application teams to define observability requirements and implement metrics, logs, and traces.
  • Manage production incidents, including on-call response, troubleshooting, service restoration, and root-cause analysis.
  • Execute application and infrastructure releases through GitLab and Jenkins deployment pipelines.
  • Manage the operational lifecycle of infrastructure, including upgrades, patching, and technology refreshes.
  • Assess service resilience through capacity planning, failure-mode analysis, and disaster recovery testing.
  • Automate operational activities using an everything-as-code approach to improve efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
$104,000 - $166,000 / year
Apply Now