Senior Site Reliability Engineer

New
L
LumaAI/Infrastructure
SF Bay Area / Remote (US)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonBashAirflowGoLinuxTerraform

Requirements

  • 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep, hands-on Linux expertise.
  • Experience with containerized systems and low-level performance debugging.
  • Working experience with Terraform, Airflow, and Ray.
  • Strong experience with AWS or OCI.
  • Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
  • Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
  • Comfort in a less-structured, fast-paced environment.

Responsibilities

  • Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
  • Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply, at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure.
  • Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures.
  • Help achieve and maintain security certifications like SOC 2 and ISO.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now