Senior Site Reliability Engineer

New
L
Level AIAI, Contact Centers
Remote, India, EST (8 PM to 4 AM IST)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
4-5 years
Required Skills
PythonGCPKubernetesGoRustCI/CDTerraform

Requirements

  • 4-5 years of hands-on systems experience.
  • Production experience in Python, Go, or Rust.
  • Deep knowledge of Kubernetes at scale, including scheduler behavior, HPA/VPA, and node pool design.
  • Experience with cost-aware autoscaling tools like Cast AI or Karpenter.
  • Cloud and on-premise infrastructure experience, specifically with GCP and IaC tools like Terraform.
  • Familiarity with GPU workload throughput profiling, inference server tuning, and utilization metrics.
  • Strong observability skills including metrics, traces, logs, and SLO management.
  • Demonstrated history of FinOps and converting infrastructure choices into cost outcomes.
  • Ability to independently handle platform-security workstreams.
  • Must be able to work in the EST time zone (8 PM to 4 AM IST).

Responsibilities

  • Own infrastructure cost efficiency, Kubernetes right-sizing, and cost telemetry.
  • Run structured experimentation programs on on-premise GPU clusters.
  • Build tooling, dashboards, and processes to enable backend teams to manage their own reliability and cost budgets.
  • Instrument systems to ensure comprehensive coverage for cost-at-scale and reliability.
  • Execute defined platform-security workstreams.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now