Senior AI Backend Engineer - Agent Evaluation & Quality
New
S
SallaAI / E-commerce
Makkah, Makkah Province, Saudi ArabiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonTypeScriptCI/CDLLMLangChain
Requirements
- 5+ years of software engineering experience.
- Recent hands-on experience building with LLMs and agents.
- Proficiency in production-grade Python or TypeScript.
- Experience with clean API and system design.
- Experience with testing frameworks and CI/CD pipelines.
- Experience building with orchestration frameworks like LangGraph or LangChain.
- Deep understanding of RAG, tool/function calling, and agent behavior.
- Experience running LLM systems in production, including managing reliability, latency, cost, and observability.
- Measurement mindset with ability to reason about metrics, calibration, and experimentation.
Responsibilities
- Own the evaluation stack by designing and building LLM-as-judge systems.
- Calibrate evaluation systems against human labels to ensure quality measurement.
- Build per-PR evaluation harnesses and regression detection integrated into CI/CD.
- Develop user simulators to generate test coverage and adversarial cases.
- Convert production signals into improvement loops for evaluation sets.
- Collaborate with product teams to define concrete, measurable quality criteria.
- Contribute to agent development and hardening based on evaluation insights.
View Full Description & ApplyYou'll be redirected to the employer's site