Senior Platform Engineer, GitLab Orbit

New
G
GitLabSoftware Development
Remote, Canada; Remote, United StatesFull-TimeSenior
Salary139,200 - 235,200 USD per year
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSGCPKubernetesClickhouseRustTerraformHelmDistributed Systems

Requirements

  • Experience designing, building, and operating production backend services with strong Rust skills or clear evidence of ability to ramp up in Rust.
  • Experience with distributed-system design, including concurrency, failure handling, consistency, messaging, data partitioning, scalability, and multi-tenant isolation.
  • Hands-on knowledge of AWS, GCP, or both, including cloud networking, identity management, compute, and object storage.
  • Experience deploying and troubleshooting applications on Kubernetes with Helm.
  • Experience contributing to repeatable, reviewable infrastructure changes using Terraform or similar infrastructure as code tools.
  • Experience improving the reliability, observability, maintainability, and on-call readiness of backend services.
  • Ability to diagnose issues across application, data, orchestration, and infrastructure layers.
  • Strong system design skills, including making and explaining architectural decisions and aligning trade-offs with product and platform needs.
  • Ability to learn and apply new languages as needed, such as Ruby, Go, or TypeScript and Vue.

Responsibilities

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve the deployment, monitoring, and operations of GitLab Orbit across GitLab.com, Dedicated, and Self-Managed deployments using Kubernetes, Helm, Terraform, and cloud services.
  • Automate recurring operational work and build tools that make deployments, upgrades, recovery, capacity management, and service maintenance safer and more efficient.
  • Strengthen observability by improving metrics, logs, traces, dashboards, and alerts, and collaborating with site reliability engineering teams.
  • Investigate production issues, address underlying causes, and write backend code that handles concurrency, partial failures, and multi-tenant isolation.
  • Build and improve the graph query engine, SDLC indexing pipelines, and API/MCP surfaces.
  • Design reliable, scalable, and cost-aware data workflows using systems such as Amazon S3, ClickHouse, NATS, and Siphon.
View Full Description & ApplyYou'll be redirected to the employer's site
139,200 - 235,200 USD per year
Apply Now