Staff Software Engineer, Distributed Systems
New
O
OrbisDistributed data infrastructure
Remote - US / Middle EastFull-TimeStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 7+ years building distributed systems in production
- Required Skills
- PythonKubernetesGogRPCGitHub ActionsDatadogDistributed Systems
Requirements
- Have 7+ years of production distributed-systems experience.
- Have a track record of shaping whole platform or infrastructure areas, rather than only individual features.
- Have deep hands-on Go experience in high-throughput concurrent systems, including goroutines, bounded channels, context-based cancellation, and graceful shutdown under load.
- Have working knowledge of message queues and streaming systems such as NATS/JetStream or Kafka, including delivery semantics, retention, and eviction strategies.
- Have experience designing state machines and idempotent operations that recover deterministically after crashes or partial failures.
- Have experience with gRPC/protobuf service boundaries and cross-system transfer protocols, including schema versioning and bidirectional session management.
- Understand distributed-systems failure semantics, including supervised restarts with backoff and deterministic reconciliation loops.
- Have set technical direction and be able to explain a subsystem you built and what you would do differently.
- Be able to travel up to 25% for customer engagement, architecture reviews, or team collaboration, including travel to the Middle East.
- Preferred: experience operating embedded streaming or data infrastructure without an external broker or open control ports.
- Preferred: experience with regulated or high-reliability customers, or national security, intelligence, or defense environments.
- Active Secret-or-higher security clearance; Top Secret preferred.
Responsibilities
- Define architecture across Catalyst’s distributed data plane, including event streaming, queuing, and state machines.
- Maintain and expand Go and Python frameworks and libraries for batch and streaming data movement plugins.
- Design high-performance, low-latency architectures for specialized workloads such as TAK protocols and RTSP media streams.
- Manage component and dependency lifecycles, using automated tooling to identify vulnerabilities and support safe upgrades.
- Define, monitor, and improve reliability targets and SLOs using Datadog, and lead incident response and postmortems.
- Architect and maintain CI/CD pipelines with GitHub Actions and develop Infrastructure-as-Code using tools such as Pulumi or Ansible.
- Operate distributed services with Docker and improve Kubernetes deployments across the fleet.
- Architect the distributed artifact registry and scoped-sync mechanism for content-addressed distribution of platform components.
- Document and socialize engineering procedures and standards using Notion and GitHub.
- Travel for customer engagement, architecture reviews, or team collaboration, including travel to the Middle East.
View Full Description & ApplyYou'll be redirected to the employer's site