Site Reliability Engineer
T
TinybirdData Infrastructure
Spain, EU timezoneFull-TimeSenior
Salary€58,000 - €97,000 a year
Apply NowOpens the employer's application page
Job Details
- Languages
- English and Spanish
- Required Skills
- AWSPythonSQLGCPKubernetesC++ClickhouseCI/CDDistributed Systems
Requirements
- Strong experience designing, building, and running distributed cloud architectures.
- Deep knowledge of Kubernetes, including cluster operations and custom controllers.
- Experience with autoscaling mechanisms like KEDA and Karpenter.
- Skilled in AWS and GCP environments.
- Ability to read and understand existing codebase (Python and C++).
- Experience debugging production incidents and improving observability.
- SQL proficiency for querying analytical data.
- Ability to communicate clearly in writing for documentation and asynchronous work.
- Fluent in English and Spanish.
- Willingness to participate in on-call rotations.
- Located in an EU timezone.
Responsibilities
- Enhance high availability and elasticity to handle growth automatically.
- Boost observability capabilities, including telemetry, dashboards, and alerting.
- Improve disaster recovery tools and incident discovery processes.
- Handle Kubernetes lifecycle tasks, cluster management, and autoscaling.
- Extract peak performance from ClickHouse and identify system bottlenecks.
- Transform manual processes into repeatable, well-managed systems.
- Strengthen CI/CD foundations to support confident deployments.
- Participate in on-call rotations to resolve incidents and support service health.
View Full Description & ApplyYou'll be redirected to the employer's site