Weekend DevOps Engineer

New
S
Sporty
Must be based in Europe or Asia or LatAM, Core working hours are 10am-3pm in your local time zone, with flexibility outside of this.Full-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years DevOps / platform engineering experience
Required Skills
AWSPythonBashKubernetesGoPrometheusTerraformHelm

Requirements

  • Have 3+ years of DevOps or platform engineering experience.
  • Be based in Europe, Asia, or LatAm.
  • Have experience independently leading project planning and deployment.
  • Have experience with cloud platforms, especially AWS, and using cloud resources to meet team and production demand.
  • Have a strong understanding of Kubernetes and container orchestration; EKS, ArgoCD, and Helm experience is highly valued.
  • Have Infrastructure-as-Code experience, particularly with Terraform.
  • Be proficient in scripting and automation with Bash, Python, or Golang.
  • Have hands-on observability experience covering metrics, logs, distributed traces, and profiling.
  • Have experience with real user monitoring (RUM).
  • Have on-call and incident-response experience, including production triage, post-mortems, and follow-up actions.
  • Have experience defining SLIs and SLOs and applying them to reliability work.
  • Have solid networking knowledge, especially TCP/IP and HTTP, and experience with high-volume HTTP systems and high availability.
  • Understand caching, including CDN, HTTP cache, Redis, or Memcached.
  • Have strong troubleshooting skills, including Linux OS diagnosis and parameter optimisation.

Responsibilities

  • Improve infrastructure and processes across deployed countries and streamline deployment to new countries.
  • Improve Kubernetes platform stability and efficiency, resource utilisation, costs, and environment provisioning through GitOps-first practices.
  • Monitor and maintain cloud infrastructure using autoscaling, alerting pipelines, and Grafana dashboards for metrics, logs, traces, and RUM.
  • Own weekend on-call operations, triage and respond to production incidents, perform root cause analysis, and lead post-incident reviews.
  • Design and maintain actionable alert pipelines that reduce alert fatigue, waterfall alerting, and notification flooding.
  • Define and maintain SLIs and SLOs for critical services and use them to guide reliability improvements and on-call prioritisation.
  • Take ownership of cloud operations and liaise with external security agencies for annual audits while performing internal security sweeps.
  • Help reconfigure architecture to enable rapid deployments to new countries.
  • Mentor less experienced team members.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now