LLM Application Engineer

B
Bolder AppsAI software
Workable locations: Argentina. Colombia. Mexico. Uzbekistan. Georgia, US hours overlap through roughly 5 PM EST when live coordination is neededContractSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English at C1 or above
Required Skills
PythonGCPPrompt Engineering

Requirements

  • Have shipped LLM applications in production, including prompts, structured outputs, retries, observability, and failure handling.
  • Have hands-on experience with Google Gemini, including multimodal text, image, or document workflows and structured extraction.
  • Have practical experience with at least one other major LLM stack, such as OpenAI or Anthropic.
  • Have experience extracting structured data from messy inputs such as HTML, PDFs, images, and email-like content.
  • Have experience with classification or taxonomy systems built on LLM outputs.
  • Use offline evaluations, regression suites, and production quality metrics with clear acceptance criteria.
  • Understand token budgeting, model tiers, caching, and batching, and can explain cost tradeoffs.
  • Have Python backend experience on serverless cloud, such as Cloud Functions, and document stores such as Firestore or similar GCP patterns.
  • Have English proficiency at C1 or above for client-adjacent debugging with a PM.
  • Be available for US hours overlap through roughly 5 PM EST when live coordination is needed.
  • Bring ownership habits, including honest estimates, raising blockers early, and completing releases.

Responsibilities

  • Own production LLM pipelines end to end, from ingestion and multimodal model calls to structured records and storage.
  • Design prompt and schema strategies to produce consistent, product-ready results.
  • Build classification and filtering layers, including taxonomy mapping, audience filters, deduplication, and cleanup logic.
  • Define and run evaluation harnesses, including golden sets, regression suites, and online metrics.
  • Track and report against quality targets for completeness, duplicates, incorrect inclusions, image presence, and related product SLAs.
  • Optimize token usage, model tiering, caching, and batching to manage cost and latency.
  • Harden long-running asynchronous jobs with timeouts, partial recovery, memory limits, and safe production deployments.
  • Partner with Flutter/mobile and QA teams on field contracts, review queues, and incident debugging.
  • Document architecture and runbooks, and recommend model changes, fallbacks, or schema improvements.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now