Senior Site Reliability Engineer
New
R
Red HatCloud Infrastructure
Source API remote eligibility restrictions: United StatesFull-TimeSenior
SalaryUSD 118600 - 195680 / year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience operating production services on Kubernetes / OpenShift; 3+ years of programming experience in Python, Go; 2+ years of experience of using cloud providers and technologies (Google, Azure, Amazon, etc.)
- Required Skills
- PythonKubernetesGoPrometheusCI/CDLinuxDatadog
Requirements
- 5+ years of experience operating production services on Kubernetes / OpenShift.
- 3+ years of programming experience in Python or Go.
- 2+ years of experience using cloud providers (Google, Azure, Amazon).
- Solid understanding of Linux systems administration (RHEL/Fedora preferred).
- Understanding of standard networking (TCP/IP, DNS, HTTP/TLS) and authentication protocols (LDAP).
- Experience with GitOps workflows for managing infrastructure or application configuration.
- Knowledge of SRE principles, including SLOs, error budgets, and toil measurement.
- Comfort with incident response and on-call responsibilities.
- Ability to work independently with minimal supervision.
Responsibilities
- Design, build, and manage large-scale infrastructure and platform services including public cloud, private cloud, and datacenter-based.
- Automate cloud infrastructure using scripting (Python and Go) and tools like auto scaling and load balancing.
- Implement and maintain intelligent infrastructure and application monitoring (e.g., Splunk, Prometheus, DataDog).
- Develop standardized CI/CD platform components using OpenShift Pipelines, Tekton, and GitLab.
- Apply Infrastructure as Code methodologies using GitOps practices with ArgoCD.
- Lead escalation support for high severity platform-impacting events and drive blameless postmortems.
- Mentor peers and participate in a regular on-call schedule.
View Full Description & ApplyYou'll be redirected to the employer's site