Staff Applied AI Engineer, Product & Agent Performance

New
J
JobgetherHealthcare AI
Based in the United StatesFull-TimeStaff
Salary175,000 - 200,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
8+ years of production software engineering experience, including at least 3 years of hands-on ownership of ML, LLM, or agentic systems in production.
Required Skills
Machine LearningSoftware EngineeringPrompt EngineeringLLM

Requirements

  • 8+ years of production software engineering experience.
  • At least 3 years of hands-on ownership of ML, LLM, or agentic systems in production.
  • Professional experience working with AI systems in healthcare, finance, or another regulated environment.
  • Demonstrated ability to diagnose agent failures and implement system-level improvements.
  • Strong understanding of evaluating AI failures based on severity, risk, and cost.
  • Hands-on experience designing and implementing RAG architectures.
  • Experience with production-grounded evaluation frameworks.
  • Experience developing fallback mechanisms, human-in-the-loop workflows, or escalation logic.
  • Practical familiarity with AWS AI/ML services, including Bedrock and SageMaker.
  • Strong software engineering foundations and product-level decision-making capability.
  • Ability to challenge launch decisions based on safety or performance standards.

Responsibilities

  • Design and improve agent behavior across live, long-horizon, multi-turn, and multi-agent workflows.
  • Architect retrieval and context strategies to ensure agents remain grounded in reliable information.
  • Design memory and state-management approaches for retaining, summarizing, or discarding information.
  • Develop prompt and context templates using few-shot examples and structured formats for consistent agent behavior.
  • Build production-representative evaluation suites to measure accuracy, reliability, latency, and cost.
  • Create evaluation rubrics and performance thresholds that account for risk and impact of failures.
  • Design validation and escalation mechanisms to route high-risk cases to human review.
  • Maintain product-level AI documentation, including model cards and intended-use guidance.
  • Translate production failures and performance data into actionable improvements for cross-functional teams.
View Full Description & ApplyYou'll be redirected to the employer's site
175,000 - 200,000 USD per year
Apply Now