AI Agent Engineer, Cloud Infrastructure
New
O
OrbisAI cloud infrastructure
Remote-USFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- Must have 5 years of proven experience translating user requirements into wireframes, workflows and such. Must have 1-year proven experience utilizing Claude or similar commercial offering to build AI agents.
- Required Skills
- PythonGCPAzureRESTful APIsPrompt Engineering
Requirements
- Have 5 years of proven experience translating user requirements into wireframes, workflows, and similar deliverables.
- Have 1 year of proven experience using Claude or a similar commercial offering to build AI agents.
- Demonstrate experience building a functioning multi-agent workflow and explaining its design, handoffs, context sharing, failures, and fixes.
- Have practical command of prompt engineering, including retrieval-augmented generation (RAG) and tool or function calling.
- Explain technical work, system behavior, limitations, and failure modes to non-technical audiences.
- Elicit requirements from stakeholders who cannot provide a specification and reach a working solution.
- Troubleshoot agent workflows built with platforms such as Claude, GPT-class models, LangChain, AutoGen, or CrewAI.
- Have working proficiency with APIs and enough scripting ability to build and debug workflows independently, including JSON payloads or Python scripts.
- Have strong written communication skills for producing documentation, runbooks, and reports.
- Have hands-on production experience with at least one major cloud platform (GCP, Azure, or Cloudflare) and working familiarity with a second.
- Have practical experience managing infrastructure as code using Terraform, Pulumi, or an equivalent.
- Have working command of cloud networking, identity, and secrets management, and treat security posture and run cost as design constraints.
Responsibilities
- Design, build, and deploy production multi-agent workflows for triage, data retrieval, summarization, reporting, and escalation routing.
- Translate process-owner requirements into workflow specifications and explain system behavior, reliability, and limitations to non-technical stakeholders.
- Monitor production workflows, diagnose failures across prompts, tools, data, and orchestration, and resolve them.
- Track AI workload costs and identify ways to keep execution costs defensible as volume grows.
- Define accuracy standards, human review checkpoints, and escalation criteria, and build checks to verify those standards.
- Author and maintain runbooks and SOPs for operating and troubleshooting workflows.
- Integrate workflows across Catalyst, Pulse, and Discovery, with engineering support where needed.
- Build and maintain reporting and data exports for resolution rate, escalation rate, output accuracy, response time, and employee satisfaction.
- Design, deploy, and maintain cloud infrastructure across GCP, AWS, Cloudflare, and Azure for enterprise workloads and client delivery environments.
- Manage infrastructure as code and own cloud networking, identity, secrets, security posture, run costs, and deployment and release paths.
View Full Description & ApplyYou'll be redirected to the employer's site