Senior Site Reliability Engineer
New
J
JobgetherHealthcare Technology
Based in United StatesFull-TimeSenior
SalaryCompetitive base salary range of $160,000–$208,000 USD.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSDockerPythonGCPKubernetesAzureGogRPCPrometheusLinuxHelm
Requirements
- 5+ years of programming experience with proficiency in languages such as Python, Go, or Shell scripting.
- Strong experience with containerization technologies including Docker, Containerd, and Kubernetes.
- Hands-on experience with CNCF technologies such as Helm, gRPC, and Prometheus.
- Experience managing public cloud environments including AWS, GCP, or Azure.
- Strong understanding of networking fundamentals, including TCP/IP, UDP, DNS, routing, firewalls, and load balancing.
- Solid Linux system administration experience and knowledge of Linux architecture principles.
- Understanding of Site Reliability Engineering concepts including monitoring, automation, performance optimization, and incident management.
- Experience building and maintaining reliable infrastructure platforms for production workloads.
- Ability to work independently, identify challenges proactively, and drive solutions with limited supervision.
- Strong communication and collaboration skills with the ability to work effectively across technical teams.
- Adaptability and willingness to learn new technologies in a fast-changing environment.
Responsibilities
- Design and maintain systems for declarative application and infrastructure lifecycle management, including CI/CD pipelines, continuous deployment, and service inventory management.
- Build, improve, and automate infrastructure processes to reduce manual effort and operational complexity.
- Manage and support containerized workloads, including Kubernetes-based environments and cloud infrastructure platforms.
- Monitor, troubleshoot, and resolve infrastructure issues while minimizing downtime and improving system reliability.
- Develop automation tools and workflows that improve deployment efficiency, system performance, and operational consistency.
- Contribute to the strategic direction of Site Reliability Engineering practices and align technical initiatives with broader business objectives.
- Collaborate with engineering teams, technical leads, and data professionals to develop scalable infrastructure solutions.
- Improve monitoring, alerting, performance tuning, and incident response processes.
- Support infrastructure reliability across different compute, storage, and networking environments.
- Promote a collaborative engineering culture focused on innovation, knowledge sharing, and continuous improvement.
View Full Description & ApplyYou'll be redirected to the employer's site