Senior Site Reliability Engineer

New
C
ClickHouseCloud Infrastructure
US-RemoteFull-TimeSenior
Salary$150K - $230K
Apply NowOpens the employer's application page

Job Details

Experience
At least 8 years of experience in Site Reliability Engineering or a related field.
Required Skills
AWSPythonGCPKubernetesAzureClickhouseGoTerraformAnsible

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Previous experience using ClickHouse in production.
  • Strong problem-solving and production debugging skills.
  • Understanding of distributed databases and SQL.

Responsibilities

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems.
  • Establish and manage service level objectives (SLOs) and service level agreements (SLAs).
  • Monitor and alert on infrastructure components across Dataplane, Control Plane, and Core services.
  • Manage incident response processes, including blameless post-mortem analysis.
  • Plan and drive Chaos engineering initiatives.
  • Manage on-call processes to resolve performance issues and minimize downtime.
View Full Description & ApplyYou'll be redirected to the employer's site
$150K - $230K
Apply Now