Senior AI Infrastructure Engineer
New
J
JobgetherAI Infrastructure
Fully remote work opportunity within the United States.Full-TimeSenior
Salary$155,000 to $180,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonSQLGoCI/CDPrompt Engineering
Requirements
- 6+ years of software engineering experience, including 1+ years focused on AI/ML infrastructure, LLM platforms, or agentic systems.
- Strong proficiency in Go (Golang) for backend engineering, with experience using Python and SQL.
- Experience designing, building, and operating production APIs and developer-facing platforms at scale.
- Strong understanding of multi-tenant architectures, including rate limiting, tenant isolation, and secure shared infrastructure.
- Hands-on experience with LLM APIs such as OpenAI or Anthropic, including prompt engineering and function/tool calling.
- Practical experience building agent workflows involving context management, state handling, caching strategies, and multi-step execution.
- Experience with MCP or similar tool-layer frameworks for AI agents and platform integrations.
- Familiarity with agent-to-agent communication protocols and durable workflow systems.
- Experience creating automated evaluation frameworks for AI systems, including testing suites, regression detection, and CI/CD quality controls.
- Ability to collaborate effectively with engineering, product, and operational stakeholders.
- Strong technical communication skills with experience creating design documents, system diagrams, and postmortems.
Responsibilities
- Design and implement multi-tenant AI infrastructure, including tenant isolation, rate limiting, and caller attribution systems.
- Build and maintain durable agent runtimes supporting long-running workflows, task lifecycle management, and failure recovery.
- Develop and enhance agent orchestration systems, including context management, memory, tool usage, and multi-agent coordination.
- Create evaluation-driven development frameworks with automated testing, regression detection, quality gates, and continuous improvement processes.
- Extend AI communication interfaces and enable agent-to-agent interactions across internal and customer-facing systems.
- Develop progressive discovery strategies for tools, agents, skills, and contextual information using semantic filtering and intelligent retrieval approaches.
- Build and maintain API-first infrastructure that supports both traditional applications and AI-powered consumers.
- Contribute to LLM abstraction layers that support flexibility across multiple model providers.
- Partner with engineering, product, and operations teams to align technical solutions with business needs.
- Participate in production support, incident response, postmortems, and reliability improvements.
View Full Description & ApplyYou'll be redirected to the employer's site