Staff Applied AI Engineer, Product & Agent Performance
New
A
ArcadiaHealthcare
Remote (USA)Full-TimeStaff
Salary175,000 - 200,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years of production software engineering experience, including 3+ years of hands-on ownership of ML, LLM, or agentic systems
- Required Skills
- Machine LearningPrompt EngineeringLLM
Requirements
- 8+ years of production software engineering experience
- 3+ years of hands-on ownership of ML, LLM, or agentic systems in production
- Experience in healthcare, finance, or another regulated industry
- Ability to diagnose agent failures and attribute fixes to instruction, retrieval, context, or memory design
- Hands-on experience with RAG architecture
- Experience with production-grounded evaluation frameworks
- Experience with fallback or human-in-the-loop logic for automated systems
- Working familiarity with AWS AI/ML services (Bedrock, SageMaker)
- Evidence-led judgment and ability to manage launch decisions
- Experience with long-horizon, multi-turn or multi-agent workflows
- Experience with product-level AI documentation such as model cards
Responsibilities
- Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks
- Design retrieval and context architecture to ensure agents remain grounded in real data
- Design memory and state handling across multi-turn and multi-agent flows
- Create context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding
- Improve performance through prompting, tool-use strategy, and context construction via direct experimentation
- Build and run evaluations against production conditions to measure performance, regressions, and failure modes
- Author evaluation rubrics, quality heuristics, and thresholds to monitor production behavior
- Design and validate escalation paths that route agents to human review based on uncertainty
- Design for cost-aware performance alongside latency, reliability, and accuracy
- Evaluate and sign off on model changes by baselining current behavior and making go/no-go decisions
View Full Description & ApplyYou'll be redirected to the employer's site