Senior Software Engineer — Infra Agent Systems

New
J
JobgetherInfrastructure AI
IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
PythonKubernetesTypeScriptGoRustDistributed Systems

Requirements

  • Bachelor’s degree or equivalent professional experience in Computer Science, Engineering, or a related technical field.
  • 5+ years of professional experience building production backend systems, distributed systems, infrastructure platforms, or similarly complex software.
  • Strong systems design capabilities and demonstrated experience taking significant systems from initial architecture through production.
  • Deep expertise in at least one relevant area, such as AI agent systems, orchestration, tool use, evaluation, grounding, knowledge graphs, graph data modeling, search, retrieval, ranking, RAG, or semantic search.
  • Strong backend engineering skills, including API design, service boundaries, data modeling, and integrations across complex technical environments.
  • Experience with Kubernetes, GitOps practices such as ArgoCD, infrastructure-as-code, and cloud platforms.
  • Proficiency in one or more relevant programming languages, such as Go, TypeScript, Python, or Rust, with the ability to work across multiple languages when required.
  • Strong analytical and problem-solving abilities, with an interest in solving ambiguous and technically challenging infrastructure problems.
  • Ability to own systems in production, balancing engineering quality, reliability, operational requirements, and delivery speed.

Responsibilities

  • Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments.
  • Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents.
  • Build fleet intelligence capabilities that combine telemetry, infrastructure state, operational knowledge, and historical incidents to improve agent decision-making.
  • Integrate agent systems with observability, incident management, ticketing, fleet inventory, source control, communication platforms, and internal infrastructure through reliable APIs.
  • Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support.
  • Improve agent quality and reliability through evaluations, retrieval optimization, better tools, and continuous feedback from production environments.
  • Convert insights and knowledge generated through production use into reliable, reviewed software, workflows, and automation.
  • Contribute to the evolution of platform architecture and engineering practices as autonomous infrastructure capabilities scale.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now