Senior DevOps / SRE Engineer
New
M
MLabsFinTech infrastructure
Source API remote eligibility restrictions: United States, Based in US to GMT timezonesFull-TimeSenior
SalaryUSD 120000 - 150000 / year
Apply NowOpens the employer's application page
Job Details
- Required Skills
- Node.jsPostgreSQLPythonAWS EKSKubernetesTypeScriptClickhouseGoRedis
Requirements
- Write production-grade code in Go, Python, and Node.js/TypeScript; Go is strongly preferred for runtime services.
- Have a track record building complex backend services, including APIs, job scheduling, and distributed systems, from scratch.
- Understand low-latency environments, websocket management, and sub-second condition evaluation.
- Have production experience with Postgres, Redis, and at least one analytical database such as ClickHouse, TimescaleDB, or BigQuery.
- Have hands-on experience deploying, scaling, and debugging production workloads on Kubernetes (AWS EKS).
- Be able to design, build, deploy, and maintain systems throughout their lifecycle.
- Model-serving and inference optimization experience with tools such as vLLM, TGI, or TensorRT-LLM is preferred.
- Experience with exchange APIs, wallet operations, or on-chain infrastructure is preferred.
- Familiarity with Model Context Protocol (MCP) or multi-agent platforms is preferred.
Responsibilities
- Lead the evolution of the per-agent process into a centralized Go service using Postgres state and real-time websocket price feeds.
- Build a YAML-configurable Scanner Gateway for scoring and filtering signals.
- Develop and maintain the RatchetStop Backend for sub-second evaluation and websocket-based order execution.
- Manage the Model Context Protocol server and the Redis, Postgres, and ClickHouse data pipeline.
- Migrate agents from third-party platforms to a custom-hosted environment with isolated workspaces and state persistence.
- Evaluate and implement self-hosted inference to replace external LLM APIs.
- Architect telemetry systems that capture agent decisions and scores.
- Build CI/CD pipelines for zero-downtime rollouts.
- Manage AWS/EKS environments using Infrastructure-as-Code.
- Own fleet operational health, monitoring, alerting, and incident response.
View Full Description & ApplyYou'll be redirected to the employer's site