Senior Site Reliability Engineer
New
C
ClickHouseCloud Infrastructure
We can only hire for candidates with a valid work authorization status in the UK, Netherlands, Sweden, Germany, France, Spain, Italy, Portugal, Poland or Czech Republic at this time.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- At least 8 years of experience in Site Reliability Engineering or a related field.
- Required Skills
- AWSPythonGCPKubernetesAzureClickhouseGoTerraformAnsible
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- At least 8 years of experience in Site Reliability Engineering or a related field.
- Previous experience using ClickHouse in production.
- Hands-on experience with Go and/or Python.
- Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
- Excellent understanding of distributed databases and SQL.
- Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
- Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
- Strong problem-solving and production debugging skills.
- Excellent communication and interpersonal skills.
Responsibilities
- Collaborate with various engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
- Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
- Ensure infrastructure components (Dataplane, Control Plane, ClickHouse Core) have monitoring and alerting for incident detection.
- Enhance incident response processes and conduct post-mortem analysis for outages.
- Continuously improve the reliability and performance of ClickHouse services.
- Plan, enable, and drive Chaos initiatives across engineering teams.
- Manage on-call processes to respond to performance and reliability issues.
View Full Description & ApplyYou'll be redirected to the employer's site