Senior Site Reliability Engineer

New
R
RunwareAI Infrastructure
United Kingdom. France. Germany. Spain. Italy. SwedenFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PHPPythonKubernetesMySQLGoRedisDistributed Systems

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, or Platform Engineering role
  • Strong understanding of distributed systems
  • Ability to debug across applications, databases, queues, containers, networking and infrastructure
  • Experience designing and operating observability systems using metrics, logs and distributed tracing
  • Deep understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Experience with Kubernetes, containers, IaC and automated deployment practices
  • Ability to write software and automation using Python, Go, or PHP
  • Experience with high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments

Responsibilities

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads
  • Participate in engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now