Senior Site Reliability Engineer, Production Engineering

New
J
JobgetherCloud Infrastructure
India, Global 24/7 reliability operations environment with flexibility to work split-weekend shifts.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
7+ years
Required Skills
PythonJenkinsKubernetesGoLinux

Requirements

  • 7+ years of experience administering large-scale production Kubernetes environments in high-availability cloud or data-center settings.
  • Bachelor's degree in Computer Science, Engineering, Mathematics, or equivalent professional experience.
  • Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management.
  • Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
  • Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core networking.
  • Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure.
  • Experience with CI/CD technologies and tools such as Jenkins and ArgoCD.
  • Strong understanding of observability, incident management, reliability engineering, and production operations.
  • Excellent analytical and troubleshooting skills under pressure.
  • Strong communication and interpersonal skills to present technical information and influence stakeholders.

Responsibilities

  • Support production Kubernetes services as part of a global 24/7 production engineering operation.
  • Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure.
  • Automate operational processes to reduce manual tasks and improve engineering efficiency.
  • Use monitoring, observability, and alerting to proactively detect, prevent, and respond to production incidents.
  • Analyze system logs and metrics to troubleshoot complex issues and determine root causes.
  • Lead incident management calls to coordinate the resolution of critical production issues.
  • Engage with cross-functional teams and service owners to resolve complex technical incidents.
  • Develop and improve monitoring and reliability mechanisms in collaboration with development teams.
  • Perform systems administration and security monitoring across large-scale infrastructure.
  • Contribute to the architecture, deployment, and improvement of Kubernetes environments.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now