Senior Research Engineer (Agentic Behavior)

J
Amsterdam, Netherlands; Belgrade, Serbia; Berlin, Germany; Limassol, Cyprus; London, United Kingdom; Madrid, Spain; Munich, Germany; Prague, Czech Republic; Warsaw, Poland; Yerevan, ArmeniaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
At least three years (Python engineering)
Required Skills
PythonSQLKotlinPyTorch

Requirements

  • Hands-on experience building evaluation or analysis pipelines for LLMs or AI coding agents in a research or production setting.
  • Strong Python engineering skills (at least three years) with experience in data-heavy or ML-adjacent codebases.
  • Proficiency in data analysis at scale, including querying large datasets (SQL/Athena) and building data pipelines.
  • Ability to own projects end-to-end, from problem identification to shipping a fix.
  • Product-aware mindset with an interest in developer workflows.
  • Familiarity with Kotlin or a strong willingness to develop deep Kotlin expertise.
  • Experience with post-training LLMs (SFT, RLHF, DPO, GRPO) is a plus.
  • Experience with deep learning frameworks like PyTorch and LLM training stacks (TRL, verl).
  • Knowledge of AI agent development frameworks and workflows.
  • Familiarity with evaluation tools like Inspect AI, Promptfoo, or LM-evaluation-harness.
  • Experience with experiment tracking tools such as Weights & Biases or MLflow.
  • Contribution to or maintenance of open-source projects, specifically benchmarks or evaluation tools.

Responsibilities

  • Design and implement tooling to systematically capture, classify, and analyze errors made by AI coding agents.
  • Build observability pipelines over agentic traces from various coding agents.
  • Design and maintain evaluation pipelines measuring Kotlin code generation quality, including correctness, idiomaticity, and test coverage.
  • Develop simulation environments for realistic Kotlin developer tasks like KMP projects and Gradle management.
  • Experiment with post-training techniques like SFT, DPO, and GRPO to improve model handling of Kotlin patterns.
  • Collaborate with model providers such as Anthropic, OpenAI, and Google to translate research findings into model improvements.
  • Build and maintain open-source benchmarks that measure AI coding agent performance on Kotlin tasks.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now