Platform Engineer

New
C
CloudLinuxLinux infrastructure
Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English - upper-intermediate or higher
Experience
Senior-level experience in infrastructure, platform or site reliability engineering
Required Skills
KubernetesGrafanaPrometheusTerraformAnsible

Requirements

  • Have senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for keeping at least one production service running.
  • Administer and debug Linux systems on bare metal and virtual machines.
  • Have production Kubernetes experience delivered through GitOps, including personally performing cluster upgrades.
  • Use Ansible and Terraform or OpenTofu for infrastructure as code, with changes reviewed in merge requests.
  • Have production GitLab administration and GitLab CI experience; deep experience with another CI system is acceptable if you can demonstrate equivalent depth.
  • Have working knowledge of Prometheus and Grafana, including running them for a team, writing alert rules and dashboards, and reading PromQL.
  • Write technical explanations for engineers outside your team, such as runbooks, notices, and responses to requests.
  • Use AI engineering assistants such as Claude and Codex to provide context, break down tasks, design agent loops, and delegate scoped end-to-end execution with clear stop conditions; explain, debug, test, and verify their output.
  • Have strong communication and interpersonal skills for clarifying product-team needs, agreeing scope, priority, and timing, and keeping people informed.
  • Have upper-intermediate or higher English.
  • Nice to have: alerting design using SLOs, burn-rate alerts, and data-sized thresholds; MicroVM isolation for CI; S3-compatible object storage operations; AWS cost work; operating Sentry or Kafka-, ClickHouse-, and Redis-backed applications; or Python or Go for exporters and small internal services.

Responsibilities

  • Run the observability platform, keep it healthy, onboard teams, monitor cost and capacity, and maintain its alerting.
  • Run GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills.
  • Maintain other services with appropriate monitoring and runbooks.
  • Research options, select designs, and deploy requested services from scratch as code, with monitoring, backups, and documentation.
  • Respond to developers' requests about access, onboarding, pipeline problems, exporters, and dashboards, and turn recurring requests into self-service.
  • Diagnose and mitigate incidents, restore service safely, complete root-cause analysis and post-mortems, and deliver prevention or detection improvements.
  • Ship changes as code reviewed in merge requests, planning and checking before each change.
  • Write runbooks, onboarding guides, maintenance notices, and status updates for engineers outside the team.
  • Delegate collection and drafting to AI agents, review their output, and record learnings for the team.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now