Software Engineer, AI Systems

New
H
Haven Safety AISafety AI
Listing location: Canada; Workplace type: Remote; Structured job location: CanadaFull-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonFastAPILangChain

Requirements

  • Have AI or ML engineering experience, including shipping LLM systems that real users depend on.
  • Have hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.
  • Use a repeatable approach to LLM evaluation, such as representative datasets, regression testing, LLM-as-judge techniques, or human review loops.
  • Have experience assembling context for LLMs and deciding what to retrieve, how much, and why.
  • Have owned systems from deployment through monitoring and incident response, including diagnosing and fixing a failure or regression.
  • Be comfortable working across model providers and explaining tradeoffs in quality, latency, cost, context, and operational risk.
  • Have strong Python production skills, including FastAPI, asynchronous services, testing, observability, and maintainable interfaces.
  • Be comfortable with Neo4j and Cypher or a comparable graph store, and able to ramp up on graph data modeling.
  • Deep graph experience, including Cypher, schema evolution, MERGE patterns, embeddings, and live knowledge graphs, is preferred.
  • Experience with enterprise AI security, Azure production AI services, or similar search and vector platforms is a plus.

Responsibilities

  • Build and operate multi-step LLM pipelines coordinating model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.
  • Extend the coordinated agent team and orchestration layer for incident analysis, review, and enterprise learning.
  • Design context retrieval across Neo4j graph traversal, vector search, and hybrid retrieval.
  • Build evaluation datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution.
  • Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards.
  • Select models across OpenAI, Anthropic, and Google based on task requirements.
  • Partner with product and knowledge engineering and help shape the AI roadmap.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now