Software Engineer, AI Training Data & Evals Lab

New
J
JobgetherAI infrastructure
Listing location: US; Workplace type: Remote; Structured job location: US, meaningful overlap with U.S. time zonesFull-Time
Salary7,000 - 10,000 USD per month
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSNode.jsPythonGCPTypeScriptGoRESTful APIsDistributed Systems

Requirements

  • Have strong software engineering fundamentals and professional experience with Node.js and TypeScript.
  • Have strong coding ability in Python and/or Go.
  • Have experience building, deploying, and owning production systems, including APIs, backend services, and data pipelines.
  • Understand distributed systems, scalability, reliability, and engineering trade-offs.
  • Have experience with AWS or GCP and infrastructure technologies such as containers and Kubernetes.
  • Have shipped and maintained production systems that other people depend on, not only prototypes.
  • Have strong written communication skills and the ability to collaborate in an asynchronous, distributed environment.
  • Be comfortable owning ambiguous technical problems, making engineering decisions, and following projects through to production.
  • Experience with evaluation frameworks, experimentation platforms, or machine-learning tooling is a plus.
  • Experience with data pipelines, workflow orchestration, or internal platforms for research and operations teams is a plus.
  • Experience in early-stage environments or high-ownership B2B SaaS and platform teams is a plus.

Responsibilities

  • Build and maintain evaluation harnesses to measure AI models and agents on real-world tasks.
  • Improve evaluation reliability, coverage, and signal quality through rubrics, task design support, and scoring approaches.
  • Develop tools that let researchers and operators run experiments without repeatedly rebuilding workflows.
  • Build and maintain APIs and backend services for human-in-the-loop workflows, task routing, and quality control.
  • Improve data pipelines that turn expert work into structured training and evaluation datasets.
  • Strengthen observability, scalability, and operational reliability through logging, metrics, monitoring, and debugging.
  • Write maintainable production code and participate in code reviews, architecture discussions, and technical design decisions.
  • Document technical decisions and system behavior for engineers and collaborators.
  • Own core systems from development through production operation and deliver meaningful improvements.
View Full Description & ApplyYou'll be redirected to the employer's site
7,000 - 10,000 USD per month
Apply Now