Senior Site Reliability Engineer
New
L
Level AIAI, SaaS
India, Remote, EST (8 PM to 4 AM IST)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 4-5 years
- Required Skills
- PythonGCPKubernetesGoRustCI/CDTerraform
Requirements
- 4-5 years of hands-on systems experience.
- Production experience in Python and Go or Rust.
- Ability to own backend services end-to-end and reason about code.
- Deep expertise in Kubernetes, including HPA/VPA, resource requests, and cost-aware autoscaling.
- Proficiency in GCP, Terraform, CI/CD, and hybrid/on-prem GPU infrastructure.
- Familiarity with GPU throughput profiling, batching, and inference server tuning.
- Experience with observability stacks: metrics, traces, logs, and SLOs.
- Proven track record in FinOps and measurable cost reduction.
- Ability to independently handle platform-security workstreams.
- Must be able to work in the EST time zone (approx. 8 PM to 4 AM IST).
Responsibilities
- Own the reduction of Kubernetes overprovisioning and maintain cost telemetry for FinOps.
- Run structured experimentation programs on on-premise GPU clusters for throughput optimization.
- Develop tooling, dashboards, and processes to enable backend teams to own cost and reliability budgets.
- Ensure reliability instrumentation is correctly implemented for new and offline flows.
- Execute defined platform-security workstreams.
View Full Description & ApplyYou'll be redirected to the employer's site