Senior Site Reliability Engineer
New
R
RunwareAI Infrastructure
United Kingdom. France. Germany. Spain. Italy. SwedenFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PHPPythonKubernetesMySQLGoRedisDistributed Systems
Requirements
- Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, or Platform Engineering role
- Strong understanding of distributed systems
- Ability to debug across applications, databases, queues, containers, networking and infrastructure
- Experience designing and operating observability systems using metrics, logs and distributed tracing
- Deep understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
- Experience with Kubernetes, containers, IaC and automated deployment practices
- Ability to write software and automation using Python, Go, or PHP
- Experience with high-throughput or low-latency APIs and distributed systems
- Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
- Experience with RabbitMQ or other distributed messaging and queueing systems
- Experience operating MySQL, Redis, ClickHouse or similar production data systems
- Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
Responsibilities
- Own and improve the reliability, availability and performance of critical production services across the Runware platform
- Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
- Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads
- Participate in engineering on-call rotation
- Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
- Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
- Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements
View Full Description & ApplyYou'll be redirected to the employer's site