Software Engineer, AI Systems
New
H
Haven Safety AISafety AI
Workplace type: Remote. This role is based in the USFull-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonFastAPILangChain
Requirements
- Have AI or ML engineering experience, including shipping LLM systems that real users depend on.
- Have hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.
- Use a repeatable approach to LLM evaluation, such as representative datasets, regression testing, LLM-as-judge techniques, or human review loops.
- Have experience assembling context for LLMs and deciding what to retrieve and why.
- Have owned systems from deployment through monitoring and incident response, including diagnosing and fixing a failure or regression.
- Be comfortable working across model providers and explaining tradeoffs in quality, latency, cost, context, and operational risk.
- Have strong Python production habits with FastAPI, asynchronous services, testing, observability, and maintainable interfaces.
- Be comfortable with Neo4j and Cypher or a comparable graph store, and able to ramp up quickly on graph data modeling; this is strongly preferred.
- Deep graph experience, including Cypher, schema evolution, MERGE patterns, embeddings, and live knowledge graph operations, is a nice-to-have.
- Enterprise AI security experience, including prompt-injection awareness, context-leak prevention, tenant isolation, role-based access, and policy-layer separation, is a nice-to-have.
- Experience with Azure production AI services and platforms such as Azure AI Search, Pinecone, MongoDB Atlas, pgvector, or Elasticsearch is a nice-to-have.
- Prior experience operating in an enterprise-level environment is strongly preferred.
Responsibilities
- Build and operate multi-step LLM pipelines coordinating model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.
- Extend the coordinated agent team and orchestration layer for incident analysis, review, and enterprise learning.
- Design context retrieval across Neo4j graph traversal, vector search, and hybrid retrieval.
- Build evaluation datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution.
- Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards.
- Select models across OpenAI, Anthropic, and Google based on task requirements.
- Partner with product and knowledge engineering and help shape the AI roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site