AI Infrastructure Engineer – Agents & ML Systems
New
H
HavocDefense Technology
RemoteFull-TimeMiddle
Salary$175K - $200K; $175K – $200K
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- PostgreSQLPythonKubernetesTypeScriptC++GoLLMMLOpsDistributed Systems
Requirements
- Bachelor's degree in Computer Science, Engineering, Machine Learning, Data Science, Applied Mathematics, or related field.
- 3+ years of professional experience in software engineering, infrastructure, ML infrastructure, or backend systems.
- Strong programming experience in Python, TypeScript, Go, C++, or similar languages.
- Experience building production software systems, APIs, services, or internal platforms.
- Experience with or strong interest in LLM applications, AI agents, RAG pipelines, or AI developer tools.
- Familiarity with modern AI infrastructure concepts like embeddings, vector search, prompt management, and model serving.
- Strong understanding of production engineering fundamentals including reliability, testing, and maintainability.
- Knowledge of security best practices for AI systems, such as mitigating prompt injection and secrets management.
- Ability to work effectively across diverse engineering and product teams.
- Strong debugging skills and experience with complex distributed systems.
- U.S. citizenship is required.
- Ability to obtain and maintain a security clearance.
Responsibilities
- Build internal AI infrastructure that connects LLMs and AI agents with internal tools, APIs, data sources, telemetry, and engineering workflows.
- Develop and maintain agentic AI systems for task automation, data analysis, simulation, and internal productivity.
- Create and implement tool integration and connector infrastructure for AI agents using MCP and other emerging standards.
- Build pipelines for retrieval, RAG, context management, document processing, and internal knowledge search.
- Support ML infrastructure workflows including data preparation, dataset curation, experiment tracking, model evaluation, and deployment.
- Build evaluation frameworks for agent performance, tool-use reliability, model quality, and regression testing.
- Develop observability, logging, tracing, and monitoring tools for AI agents and ML pipelines.
- Secure agentic AI systems using least-privilege tool access, sandboxed execution, and robust secrets management.
- Partner with cross-functional teams to identify and implement high-value AI use cases.
View Full Description & ApplyYou'll be redirected to the employer's site