AI Benchmark Engineer | Native Language Specialist
New
J
JobgetherAI Engineering
BrazilContractMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Native or near-native fluency in the target language, and strong English proficiency for technical communication and collaboration.
- Experience
- At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
- Required Skills
- PythonSoftware EngineeringPrompt Engineering
Requirements
- At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
- Technical experience through work at established technology organizations or graduation from a strong engineering university.
- Native or near-native fluency in the target language.
- Strong English proficiency for technical communication and collaboration.
- Strong Python skills and proficiency in standard shell scripting and data processing workflows.
- Extensive experience with Terminal/CLI-based development environments.
- Familiarity with coding agents or AI-assisted development tools.
- Solid understanding of multilingual text-processing challenges including encoding/decoding, Unicode normalization, and locale-dependent casing.
- Attention to detail and ability to create deterministic, reproducible, technically rigorous evaluation tasks.
- Ability to work independently and maintain consistent quality across project-based assignments.
- Commitment of at least 2 hours per day or 10 hours per week.
- Ability to provide an updated CV in English and successfully complete a GenAI assessment.
Responsibilities
- Design and engineer challenging, realistic benchmark tasks that evaluate coding agents and large language models in multilingual terminal environments.
- Create authentic task environments using datasets, files, prompts, and other assets written in your native language.
- Identify model failure points related to native-language prompting, translation gaps, multilingual reasoning, and language-specific software workflows.
- Develop robust reference implementations and highly reliable, deterministic verifier scripts.
- Analyze execution logs and calibrate task difficulty using standardized Terminal-Bench configurations.
- Participate in a rigorous multi-layer quality process covering task creation, human review, and final audit.
- Validate grammatical accuracy, linguistic authenticity, technical correctness, and overall benchmark integrity.
- Apply deep knowledge of multilingual text processing to identify edge cases involving Unicode normalization and locale behavior.
View Full Description & ApplyYou'll be redirected to the employer's site