Senior Site Reliability Engineer
New
L
Level AIAI, Contact Centers
Remote, India, EST (8 PM to 4 AM IST)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 4-5 years
- Required Skills
- PythonGCPKubernetesGoRustCI/CDTerraform
Requirements
- 4-5 years of hands-on systems experience.
- Production experience in Python, Go, or Rust.
- Deep knowledge of Kubernetes at scale, including scheduler behavior, HPA/VPA, and node pool design.
- Experience with cost-aware autoscaling tools like Cast AI or Karpenter.
- Cloud and on-premise infrastructure experience, specifically with GCP and IaC tools like Terraform.
- Familiarity with GPU workload throughput profiling, inference server tuning, and utilization metrics.
- Strong observability skills including metrics, traces, logs, and SLO management.
- Demonstrated history of FinOps and converting infrastructure choices into cost outcomes.
- Ability to independently handle platform-security workstreams.
- Must be able to work in the EST time zone (8 PM to 4 AM IST).
Responsibilities
- Own infrastructure cost efficiency, Kubernetes right-sizing, and cost telemetry.
- Run structured experimentation programs on on-premise GPU clusters.
- Build tooling, dashboards, and processes to enable backend teams to manage their own reliability and cost budgets.
- Instrument systems to ensure comprehensive coverage for cost-at-scale and reliability.
- Execute defined platform-security workstreams.
View Full Description & ApplyYou'll be redirected to the employer's site