Senior AI Backend Engineer - Agent Evaluation & Quality

New
S
SallaAI / E-commerce
Makkah, Makkah Province, Saudi ArabiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
PythonTypeScriptCI/CDLLMLangChain

Requirements

  • 5+ years of software engineering experience.
  • Recent hands-on experience building with LLMs and agents.
  • Proficiency in production-grade Python or TypeScript.
  • Experience with clean API and system design.
  • Experience with testing frameworks and CI/CD pipelines.
  • Experience building with orchestration frameworks like LangGraph or LangChain.
  • Deep understanding of RAG, tool/function calling, and agent behavior.
  • Experience running LLM systems in production, including managing reliability, latency, cost, and observability.
  • Measurement mindset with ability to reason about metrics, calibration, and experimentation.

Responsibilities

  • Own the evaluation stack by designing and building LLM-as-judge systems.
  • Calibrate evaluation systems against human labels to ensure quality measurement.
  • Build per-PR evaluation harnesses and regression detection integrated into CI/CD.
  • Develop user simulators to generate test coverage and adversarial cases.
  • Convert production signals into improvement loops for evaluation sets.
  • Collaborate with product teams to define concrete, measurable quality criteria.
  • Contribute to agent development and hardening based on evaluation insights.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now