Senior Research Engineer (Agentic Behavior)
J
JetBrainsAI/ML
Amsterdam, Netherlands; Belgrade, Serbia; Berlin, Germany; Limassol, Cyprus; London, United Kingdom; Madrid, Spain; Munich, Germany; Prague, Czech Republic; Warsaw, Poland; Yerevan, ArmeniaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- At least three years (Python engineering)
- Required Skills
- PythonSQLKotlinPyTorch
Requirements
- Hands-on experience building evaluation or analysis pipelines for LLMs or AI coding agents in a research or production setting.
- Strong Python engineering skills (at least three years) with experience in data-heavy or ML-adjacent codebases.
- Proficiency in data analysis at scale, including querying large datasets (SQL/Athena) and building data pipelines.
- Ability to own projects end-to-end, from problem identification to shipping a fix.
- Product-aware mindset with an interest in developer workflows.
- Familiarity with Kotlin or a strong willingness to develop deep Kotlin expertise.
- Experience with post-training LLMs (SFT, RLHF, DPO, GRPO) is a plus.
- Experience with deep learning frameworks like PyTorch and LLM training stacks (TRL, verl).
- Knowledge of AI agent development frameworks and workflows.
- Familiarity with evaluation tools like Inspect AI, Promptfoo, or LM-evaluation-harness.
- Experience with experiment tracking tools such as Weights & Biases or MLflow.
- Contribution to or maintenance of open-source projects, specifically benchmarks or evaluation tools.
Responsibilities
- Design and implement tooling to systematically capture, classify, and analyze errors made by AI coding agents.
- Build observability pipelines over agentic traces from various coding agents.
- Design and maintain evaluation pipelines measuring Kotlin code generation quality, including correctness, idiomaticity, and test coverage.
- Develop simulation environments for realistic Kotlin developer tasks like KMP projects and Gradle management.
- Experiment with post-training techniques like SFT, DPO, and GRPO to improve model handling of Kotlin patterns.
- Collaborate with model providers such as Anthropic, OpenAI, and Google to translate research findings into model improvements.
- Build and maintain open-source benchmarks that measure AI coding agent performance on Kotlin tasks.
View Full Description & ApplyYou'll be redirected to the employer's site