LLM Application Engineer
B
Bolder AppsAI software
Workable locations: Argentina. Colombia. Mexico. Uzbekistan. Georgia, US hours overlap through roughly 5 PM EST when live coordination is neededContractSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English at C1 or above
- Required Skills
- PythonGCPPrompt Engineering
Requirements
- Have shipped LLM applications in production, including prompts, structured outputs, retries, observability, and failure handling.
- Have hands-on experience with Google Gemini, including multimodal text, image, or document workflows and structured extraction.
- Have practical experience with at least one other major LLM stack, such as OpenAI or Anthropic.
- Have experience extracting structured data from messy inputs such as HTML, PDFs, images, and email-like content.
- Have experience with classification or taxonomy systems built on LLM outputs.
- Use offline evaluations, regression suites, and production quality metrics with clear acceptance criteria.
- Understand token budgeting, model tiers, caching, and batching, and can explain cost tradeoffs.
- Have Python backend experience on serverless cloud, such as Cloud Functions, and document stores such as Firestore or similar GCP patterns.
- Have English proficiency at C1 or above for client-adjacent debugging with a PM.
- Be available for US hours overlap through roughly 5 PM EST when live coordination is needed.
- Bring ownership habits, including honest estimates, raising blockers early, and completing releases.
Responsibilities
- Own production LLM pipelines end to end, from ingestion and multimodal model calls to structured records and storage.
- Design prompt and schema strategies to produce consistent, product-ready results.
- Build classification and filtering layers, including taxonomy mapping, audience filters, deduplication, and cleanup logic.
- Define and run evaluation harnesses, including golden sets, regression suites, and online metrics.
- Track and report against quality targets for completeness, duplicates, incorrect inclusions, image presence, and related product SLAs.
- Optimize token usage, model tiering, caching, and batching to manage cost and latency.
- Harden long-running asynchronous jobs with timeouts, partial recovery, memory limits, and safe production deployments.
- Partner with Flutter/mobile and QA teams on field contracts, review queues, and incident debugging.
- Document architecture and runbooks, and recommend model changes, fallbacks, or schema improvements.
View Full Description & ApplyYou'll be redirected to the employer's site