Weekend DevOps Engineer
New
S
Sporty
Must be based in Europe or Asia or LatAM, Core working hours are 10am-3pm in your local time zone, with flexibility outside of this.Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years DevOps / platform engineering experience
- Required Skills
- AWSPythonBashKubernetesGoPrometheusTerraformHelm
Requirements
- Have 3+ years of DevOps or platform engineering experience.
- Be based in Europe, Asia, or LatAm.
- Have experience independently leading project planning and deployment.
- Have experience with cloud platforms, especially AWS, and using cloud resources to meet team and production demand.
- Have a strong understanding of Kubernetes and container orchestration; EKS, ArgoCD, and Helm experience is highly valued.
- Have Infrastructure-as-Code experience, particularly with Terraform.
- Be proficient in scripting and automation with Bash, Python, or Golang.
- Have hands-on observability experience covering metrics, logs, distributed traces, and profiling.
- Have experience with real user monitoring (RUM).
- Have on-call and incident-response experience, including production triage, post-mortems, and follow-up actions.
- Have experience defining SLIs and SLOs and applying them to reliability work.
- Have solid networking knowledge, especially TCP/IP and HTTP, and experience with high-volume HTTP systems and high availability.
- Understand caching, including CDN, HTTP cache, Redis, or Memcached.
- Have strong troubleshooting skills, including Linux OS diagnosis and parameter optimisation.
Responsibilities
- Improve infrastructure and processes across deployed countries and streamline deployment to new countries.
- Improve Kubernetes platform stability and efficiency, resource utilisation, costs, and environment provisioning through GitOps-first practices.
- Monitor and maintain cloud infrastructure using autoscaling, alerting pipelines, and Grafana dashboards for metrics, logs, traces, and RUM.
- Own weekend on-call operations, triage and respond to production incidents, perform root cause analysis, and lead post-incident reviews.
- Design and maintain actionable alert pipelines that reduce alert fatigue, waterfall alerting, and notification flooding.
- Define and maintain SLIs and SLOs for critical services and use them to guide reliability improvements and on-call prioritisation.
- Take ownership of cloud operations and liaise with external security agencies for annual audits while performing internal security sweeps.
- Help reconfigure architecture to enable rapid deployments to new countries.
- Mentor less experienced team members.
View Full Description & ApplyYou'll be redirected to the employer's site