Site Reliability Engineer - Low-Latency Trading Systems
New
O
Ondo FinanceFintech, Blockchain
Remote (US)Full-TimeSenior
SalaryCompetitive compensation including but not limited to salary, future token rights, and/or equity
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSKubernetesGoPostgresPrometheusRustLinuxDatadogNetworking
Requirements
- 5+ years in SRE, production engineering, or infrastructure roles.
- Meaningful experience supporting real-time or latency-sensitive systems.
- Strong programming ability in Go or Rust, with a willingness to work in both.
- Deep, hands-on experience running stateful, latency-sensitive workloads on Kubernetes and AWS.
- Strong observability instincts including fluent PromQL and structured-log analysis.
- Solid Linux internals and networking fundamentals.
- Experience designing alerts with high signal and low noise.
- Sound judgment under pressure.
- Clear written communication skills for incident documentation.
Responsibilities
- Debug production incidents end to end including stale market data feeds, exchange rate limits, WebSocket disconnects, and latency regressions.
- Harden market data ingestion from providers and venue-native feeds, including staleness detection and failover.
- Build reconciliation and data-integrity tooling across live gauges, Postgres, and S3 data lakes.
- Operate and evolve multi-region Kubernetes clusters on AWS (EKS) using GitOps workflows.
- Develop and refine observability systems using Prometheus, Grafana, and Datadog.
- Participate in an on-call rotation covering US equity market hours and 24/7 crypto venues.
View Full Description & ApplyYou'll be redirected to the employer's site