Senior Site Reliability Engineer

New
L
Level AIAI, SaaS
India, Remote, EST (8 PM to 4 AM IST)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
4-5 years
Required Skills
PythonGCPKubernetesGoRustCI/CDTerraform

Requirements

  • 4-5 years of hands-on systems experience.
  • Production experience in Python and Go or Rust.
  • Ability to own backend services end-to-end and reason about code.
  • Deep expertise in Kubernetes, including HPA/VPA, resource requests, and cost-aware autoscaling.
  • Proficiency in GCP, Terraform, CI/CD, and hybrid/on-prem GPU infrastructure.
  • Familiarity with GPU throughput profiling, batching, and inference server tuning.
  • Experience with observability stacks: metrics, traces, logs, and SLOs.
  • Proven track record in FinOps and measurable cost reduction.
  • Ability to independently handle platform-security workstreams.
  • Must be able to work in the EST time zone (approx. 8 PM to 4 AM IST).

Responsibilities

  • Own the reduction of Kubernetes overprovisioning and maintain cost telemetry for FinOps.
  • Run structured experimentation programs on on-premise GPU clusters for throughput optimization.
  • Develop tooling, dashboards, and processes to enable backend teams to own cost and reliability budgets.
  • Ensure reliability instrumentation is correctly implemented for new and offline flows.
  • Execute defined platform-security workstreams.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now