Staff Engineer (Core & MLOps)
New
J
JobgetherWeb data infrastructure
Based in the United StatesFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of experience building scalable distributed backend systems
- Required Skills
- PythonJavaKafkaKubernetesgRPCTerraformMLOps
Requirements
- Bring 10+ years of experience building scalable distributed backend systems.
- Have a strong track record creating internal platforms or core libraries adopted across engineering organizations.
- Demonstrate advanced Java expertise, including reactive frameworks such as Vert.x or Netty, and strong Python proficiency.
- Have deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
- Bring hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
- Have experience designing automated telemetry pipelines, materialized views, feature stores, or other production-data feedback systems.
- Demonstrate reliability engineering experience with SLO/SLI definition, blast-radius analysis, fault tolerance, and service contracts.
- Have strong technical writing skills and the ability to communicate architectural concepts and drive alignment across teams.
- Have written and interpersonal communication skills suited to a globally distributed, remote-first environment.
- Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
- MLOps experience with model serving, performance monitoring, or production drift detection is advantageous.
- Familiarity with zero-trust networking or service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
Responsibilities
- Architect and evolve control and context planes, including service and schema registries, SLO enforcement, health-aware routing, and automated canary releases.
- Maintain and improve Java and Python client libraries, workload specifications, Helm charts, and deployment pipelines.
- Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning, and schema evolution.
- Operate and improve platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, Valkey, and database modernization initiatives.
- Lead architectural strategy through Requests for Discussion on workflow and gateway orchestration, multi-cluster routing, and automated failover.
- Establish reliability practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
- Participate in shared infrastructure on-call rotations, lead incident post-mortems, and turn operational insights into platform improvements.
- Mentor engineers across squads, review architectural proposals, and establish engineering practices.
View Full Description & ApplyYou'll be redirected to the employer's site