Senior Software Engineer — Infra Agent Systems
New
J
JobgetherInfrastructure AI
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonKubernetesTypeScriptGoRustDistributed Systems
Requirements
- Bachelor’s degree or equivalent professional experience in Computer Science, Engineering, or a related technical field.
- 5+ years of professional experience building production backend systems, distributed systems, infrastructure platforms, or similarly complex software.
- Strong systems design capabilities and demonstrated experience taking significant systems from initial architecture through production.
- Deep expertise in at least one relevant area, such as AI agent systems, orchestration, tool use, evaluation, grounding, knowledge graphs, graph data modeling, search, retrieval, ranking, RAG, or semantic search.
- Strong backend engineering skills, including API design, service boundaries, data modeling, and integrations across complex technical environments.
- Experience with Kubernetes, GitOps practices such as ArgoCD, infrastructure-as-code, and cloud platforms.
- Proficiency in one or more relevant programming languages, such as Go, TypeScript, Python, or Rust, with the ability to work across multiple languages when required.
- Strong analytical and problem-solving abilities, with an interest in solving ambiguous and technically challenging infrastructure problems.
- Ability to own systems in production, balancing engineering quality, reliability, operational requirements, and delivery speed.
Responsibilities
- Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments.
- Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents.
- Build fleet intelligence capabilities that combine telemetry, infrastructure state, operational knowledge, and historical incidents to improve agent decision-making.
- Integrate agent systems with observability, incident management, ticketing, fleet inventory, source control, communication platforms, and internal infrastructure through reliable APIs.
- Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support.
- Improve agent quality and reliability through evaluations, retrieval optimization, better tools, and continuous feedback from production environments.
- Convert insights and knowledge generated through production use into reliable, reviewed software, workflows, and automation.
- Contribute to the evolution of platform architecture and engineering practices as autonomous infrastructure capabilities scale.
View Full Description & ApplyYou'll be redirected to the employer's site