Senior GenAI Full-Stack Engineer

New
C
CoduranceAI customer support
BrazilFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
Node.jsTypeScriptNest.jsReactRESTful APIs

Requirements

  • Have proven experience building and operating production LLM-powered systems such as chatbots, AI assistants, agent copilots, RAG systems, or LLM orchestration platforms.
  • Bring strong TypeScript and Node.js engineering skills, including TypeScript strict-mode fluency.
  • Have production AI experience with prompt engineering, RAG pipelines, agent design, tool calling, model evaluation, observability, and failure-mode analysis.
  • Be comfortable working across NestJS APIs, React UIs, databases, infrastructure, and production operations.
  • Be able to evaluate tradeoffs between model quality, latency, reliability, throughput, and cost.
  • Be able to troubleshoot AI systems across prompts, retrieval pipelines, model configuration, infrastructure, and application code.
  • Apply state-machine thinking to complex asynchronous workflows; XState or similar experience is a strong signal.
  • Understand REST API design, asynchronous patterns such as queues and events, and caching strategies.
  • Have a strong testing culture, including unit, integration, and contract tests.
  • Have experience working in a monorepo with multiple interconnected services.
  • Hands-on MCP experience or experience building tool-use agentic workflows is a strong plus.
  • Experience with LLM observability and evaluation platforms, operating AI workloads at scale, or evaluating multiple foundation models and providers is a strong plus.

Responsibilities

  • Design and extend production LLM applications and agentic workflows using NestJS, XState v5, and the OpenAI SDK.
  • Build workflows for RAG, intent detection, clarification, fulfillment, escalation, tool use, and human-in-the-loop processes.
  • Build and maintain the conversation-machine substrate, including guard and action registries, flow validation with ajv, DB-driven flow configurations, and design-time tooling in Epicenter admin.
  • Build and evolve the Epic Support Assistant and Agent Support Assistant.
  • Integrate with MCP servers for tool use and agentic behaviors.
  • Evaluate, benchmark, and tune models across providers including OpenAI, Gemini, and Anthropic.
  • Troubleshoot production LLM issues such as hallucinations, retrieval failures, prompt regressions, model drift, latency bottlenecks, and provider outages.
  • Build resilience mechanisms including retries, fallback routing, caching, streaming, rate limiting, and provider routing.
  • Instrument and tune model quality using Langfuse, evaluation datasets, A/B testing, prompt versioning, and production telemetry.
  • Manage asynchronous workloads with BullMQ, caching with Redis, and PostgreSQL persistence via Kysely.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now