Senior Site Reliability Engineer, Production Engineering
New
J
JobgetherCloud Infrastructure
India, Global 24/7 reliability operations environment with flexibility to work split-weekend shifts.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years
- Required Skills
- PythonJenkinsKubernetesGoLinux
Requirements
- 7+ years of experience administering large-scale production Kubernetes environments in high-availability cloud or data-center settings.
- Bachelor's degree in Computer Science, Engineering, Mathematics, or equivalent professional experience.
- Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management.
- Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
- Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core networking.
- Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure.
- Experience with CI/CD technologies and tools such as Jenkins and ArgoCD.
- Strong understanding of observability, incident management, reliability engineering, and production operations.
- Excellent analytical and troubleshooting skills under pressure.
- Strong communication and interpersonal skills to present technical information and influence stakeholders.
Responsibilities
- Support production Kubernetes services as part of a global 24/7 production engineering operation.
- Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure.
- Automate operational processes to reduce manual tasks and improve engineering efficiency.
- Use monitoring, observability, and alerting to proactively detect, prevent, and respond to production incidents.
- Analyze system logs and metrics to troubleshoot complex issues and determine root causes.
- Lead incident management calls to coordinate the resolution of critical production issues.
- Engage with cross-functional teams and service owners to resolve complex technical incidents.
- Develop and improve monitoring and reliability mechanisms in collaboration with development teams.
- Perform systems administration and security monitoring across large-scale infrastructure.
- Contribute to the architecture, deployment, and improvement of Kubernetes environments.
View Full Description & ApplyYou'll be redirected to the employer's site