AI Benchmark Engineer | Native Language Specialist

New
J
JobgetherAI Engineering
BrazilContractMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Native or near-native fluency in the target language, and strong English proficiency for technical communication and collaboration.
Experience
At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
Required Skills
PythonSoftware EngineeringPrompt Engineering

Requirements

  • At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
  • Technical experience through work at established technology organizations or graduation from a strong engineering university.
  • Native or near-native fluency in the target language.
  • Strong English proficiency for technical communication and collaboration.
  • Strong Python skills and proficiency in standard shell scripting and data processing workflows.
  • Extensive experience with Terminal/CLI-based development environments.
  • Familiarity with coding agents or AI-assisted development tools.
  • Solid understanding of multilingual text-processing challenges including encoding/decoding, Unicode normalization, and locale-dependent casing.
  • Attention to detail and ability to create deterministic, reproducible, technically rigorous evaluation tasks.
  • Ability to work independently and maintain consistent quality across project-based assignments.
  • Commitment of at least 2 hours per day or 10 hours per week.
  • Ability to provide an updated CV in English and successfully complete a GenAI assessment.

Responsibilities

  • Design and engineer challenging, realistic benchmark tasks that evaluate coding agents and large language models in multilingual terminal environments.
  • Create authentic task environments using datasets, files, prompts, and other assets written in your native language.
  • Identify model failure points related to native-language prompting, translation gaps, multilingual reasoning, and language-specific software workflows.
  • Develop robust reference implementations and highly reliable, deterministic verifier scripts.
  • Analyze execution logs and calibrate task difficulty using standardized Terminal-Bench configurations.
  • Participate in a rigorous multi-layer quality process covering task creation, human review, and final audit.
  • Validate grammatical accuracy, linguistic authenticity, technical correctness, and overall benchmark integrity.
  • Apply deep knowledge of multilingual text processing to identify edge cases involving Unicode normalization and locale behavior.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now