Platform Engineer
New
C
CloudLinuxLinux infrastructure
Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English - upper-intermediate or higher
- Experience
- Senior-level experience in infrastructure, platform or site reliability engineering
- Required Skills
- KubernetesGrafanaPrometheusTerraformAnsible
Requirements
- Have senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for keeping at least one production service running.
- Administer and debug Linux systems on bare metal and virtual machines.
- Have production Kubernetes experience delivered through GitOps, including personally performing cluster upgrades.
- Use Ansible and Terraform or OpenTofu for infrastructure as code, with changes reviewed in merge requests.
- Have production GitLab administration and GitLab CI experience; deep experience with another CI system is acceptable if you can demonstrate equivalent depth.
- Have working knowledge of Prometheus and Grafana, including running them for a team, writing alert rules and dashboards, and reading PromQL.
- Write technical explanations for engineers outside your team, such as runbooks, notices, and responses to requests.
- Use AI engineering assistants such as Claude and Codex to provide context, break down tasks, design agent loops, and delegate scoped end-to-end execution with clear stop conditions; explain, debug, test, and verify their output.
- Have strong communication and interpersonal skills for clarifying product-team needs, agreeing scope, priority, and timing, and keeping people informed.
- Have upper-intermediate or higher English.
- Nice to have: alerting design using SLOs, burn-rate alerts, and data-sized thresholds; MicroVM isolation for CI; S3-compatible object storage operations; AWS cost work; operating Sentry or Kafka-, ClickHouse-, and Redis-backed applications; or Python or Go for exporters and small internal services.
Responsibilities
- Run the observability platform, keep it healthy, onboard teams, monitor cost and capacity, and maintain its alerting.
- Run GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills.
- Maintain other services with appropriate monitoring and runbooks.
- Research options, select designs, and deploy requested services from scratch as code, with monitoring, backups, and documentation.
- Respond to developers' requests about access, onboarding, pipeline problems, exporters, and dashboards, and turn recurring requests into self-service.
- Diagnose and mitigate incidents, restore service safely, complete root-cause analysis and post-mortems, and deliver prevention or detection improvements.
- Ship changes as code reviewed in merge requests, planning and checking before each change.
- Write runbooks, onboarding guides, maintenance notices, and status updates for engineers outside the team.
- Delegate collection and drafting to AI agents, review their output, and record learnings for the team.
View Full Description & ApplyYou'll be redirected to the employer's site