Site Reliability Engineer Technical Lead

New
J
JobgetherSecurity & IT
Based in United StatesFull-TimeLead
Salary100,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
8+ years of IT experience
Required Skills
AWSPythonGCPKubernetesMicrosoft Azure

Requirements

  • 8+ years of IT experience, including significant experience in a senior-level SRE, infrastructure engineering, systems engineering, or closely related role.
  • Deep understanding and hands-on application of SRE principles within enterprise-scale environments.
  • Strong experience with cloud platforms such as AWS, Google Cloud Platform (GCP), and/or Microsoft Azure.
  • Solid expertise with microservices architectures and container orchestration technologies, particularly Kubernetes.
  • Demonstrated ability to write production-quality code, particularly with Python or comparable programming languages.
  • Proven experience designing and implementing automation to reduce manual operational work.
  • Experience building and maintaining internal tooling that improves infrastructure and operational processes.
  • Strong expertise in observability, including monitoring, alerting, logging, and distributed tracing.
  • Experience with technologies such as Dynatrace, Splunk, ELK Stack, or similar observability platforms.
  • Strong analytical skills and the ability to interpret system metrics to proactively identify reliability and performance issues.
  • Exceptional communication, leadership, and interpersonal skills.
  • Ability to collaborate effectively with senior leadership.

Responsibilities

  • Drive the operational excellence, reliability, scalability, and performance of critical production systems.
  • Apply Site Reliability Engineering principles to enterprise-level systems and continuously identify opportunities to improve resilience and availability.
  • Lead technical aspects of incident response, troubleshooting complex production issues and developing sustainable solutions to prevent recurrence.
  • Design and implement automation that eliminates repetitive manual work, reduces operational toil, and improves engineering efficiency.
  • Develop and maintain internal tools and automated workflows that support scalable, reliable, and self-healing infrastructure.
  • Work across cloud environments such as AWS, GCP, and Azure, as well as microservices and containerized platforms including Kubernetes.
  • Build, maintain, and optimize observability capabilities covering monitoring, alerting, logging, and distributed tracing.
  • Analyze operational metrics and system performance data to guide performance tuning, capacity planning, and reliability improvements.
  • Collaborate effectively with engineering teams and senior stakeholders to communicate technical risks, recommendations, and solutions.
View Full Description & ApplyYou'll be redirected to the employer's site
100,000 - 150,000 USD per year
Apply Now