Site Reliability Engineering (SRE) Leader
New
P
PatSnapSaaS, AI, IP
Remote, UKFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Fluent in English; Mandarin is highly desirable.
- Experience
- At least 8 years of experience in DevOps, SRE, or infrastructure operations.
- Required Skills
- AWSDockerArtificial IntelligenceKubernetesCI/CDDevOpsSaaS
Requirements
- Bachelor’s degree in Computer Science or a related field.
- At least 8 years of experience in DevOps, SRE, or infrastructure operations.
- Proven experience leading technical teams and managing production environments at scale.
- Strong expertise in cloud platforms (AWS preferred).
- Proficiency with Kubernetes, Docker, CI/CD pipelines, and Infrastructure as Code.
- Deep understanding of distributed systems and high-availability architectures.
- Experience driving automation and operational excellence initiatives.
- Hands-on experience with AI tools such as ChatGPT, Claude, GitHub Copilot, or Codex.
- Strong problem-solving, leadership, and stakeholder management skills.
- Fluent in English; Mandarin is highly desirable.
Responsibilities
- Build, lead, and develop the UK SRE team, establishing operational standards and reliability goals.
- Ensure high availability, stability, security, and performance of business-critical SaaS platforms.
- Define and drive the operational strategy for the global platform.
- Lead major incident management as the senior escalation point for critical production events.
- Establish and monitor reliability metrics including SLIs, SLOs, and operational KPIs.
- Drive automation across infrastructure, deployments, and monitoring workflows.
- Champion the adoption of AI-powered operations to enhance engineering productivity.
- Partner with Engineering, Product, Security, and Infrastructure teams to improve platform scalability.
- Lead disaster recovery planning and operational resilience initiatives.
View Full Description & ApplyYou'll be redirected to the employer's site