Site Reliability Engineer (SRE) - Evening Shift
New
P
PeratonNational Security
United States, 3pm - 11pm Eastern Standard Time (EST)Full-TimeMiddle
Salary$104,000 - $166,000 / year
Apply NowOpens the employer's application page
Job Details
- Experience
- Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience. 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
- Required Skills
- AWSPythonKubernetesCI/CDLinuxTerraformAnsible
Requirements
- Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
- Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
- 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
- Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms.
- Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
- Experience with CI/CD platforms GitLab and Jenkins.
- Proficient in Linux and Windows Server administration.
- Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk, and Open Telemetry.
- Demonstrated ownership of an SLI/SLO and alerting program.
- Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
- Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).
Responsibilities
- Operate and maintain production infrastructure services and applications to ensure availability, performance, and security.
- Monitor services using SLIs, SLOs, dashboards, and observability tools to improve detection and resolution of issues.
- Manage production incidents, including on-call response, troubleshooting, and root-cause analysis.
- Execute infrastructure releases through CI/CD pipelines including staging and production promotion.
- Automate operational tasks using an everything-as-code approach to improve efficiency and consistency.
- Collaborate with platform and application teams to define operational requirements and improve environment reliability.
View Full Description & ApplyYou'll be redirected to the employer's site