Senior Site Reliability Engineer

New
C
ClickHouseCloud Data Analytics
EMEA(Remote)Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
At least 8 years of experience in Site Reliability Engineering or a related field.
Required Skills
AWSPythonGCPKubernetesAzureClickhouseGoTerraformAnsible

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Previous experience using ClickHouse in production environments.
  • Hands-on experience with Go and/or Python programming languages.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL.
  • Hands-on experience with container orchestration tools like Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Solid production debugging skills and problem-solving abilities.
  • Strong communication and interpersonal skills.

Responsibilities

  • Collaborate with various engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  • Ensure infrastructure components like Dataplane and Control Plane have monitoring and alerting for timely incident resolution.
  • Refine incident response processes and conduct blameless post-mortem analysis for outages.
  • Continuously improve the reliability and performance of ClickHouse services.
  • Plan, enable, and drive Chaos initiatives across engineering teams.
  • Manage on-call processes and coordinate escalation to resolve performance issues.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now