Staff Engineer, AI Operations & Governance, Workplace AI
New
J
JobgetherInformation Technology
Remote work arrangement within the United States.Full-TimeStaff
SalaryAnnual salary range of $200,000–$300,000
Apply NowOpens the employer's application page
Job Details
- Experience
- 4–7+ years
- Required Skills
- RESTful APIsDevOpsJSONMLOps
Requirements
- 4–7+ years of experience in platform engineering, DevOps/SRE, ML/AI operations, or technical SaaS operations.
- Strong knowledge of APIs, integrations, infrastructure-as-code, JSON, YAML, and automation.
- Hands-on experience with monitoring and observability tools for production system diagnostics.
- Practical experience with AI platforms, LLM providers, RPA solutions, or workflow engines.
- Proficiency in secure secrets management, access controls, and least-privilege security practices.
- Strong technical documentation skills and ability to communicate with technical and non-technical audiences.
- Proven ability to manage production systems with high ownership and initiative.
- Experience in ML/AI operations (deployment, evaluation, drift monitoring) is preferred.
- Familiarity with prompt engineering, LLM policies, and guardrails is an advantage.
- Experience in AI governance, compliance, or risk frameworks is beneficial.
- Experience collaborating with Security, Legal, and business stakeholders.
- Prior experience with incident response and on-call rotations is a plus.
Responsibilities
- Manage day-to-day configurations for AI platforms and orchestration tools, including models, routing, guardrails, and prompt libraries.
- Design and maintain connectors and integrations with SaaS applications, data sources, and workflow tools.
- Build and maintain observability across AI workflows, including monitoring metrics, logs, and alerts.
- Implement secure secrets management, API key rotation, and least-privilege access controls.
- Execute production changes and model updates while maintaining change records and rollback procedures.
- Conduct technical benchmarking of AI models and tools to support governance decisions.
- Translate technical operational findings into risk and reliability insights for stakeholders.
- Serve as a technical incident responder by triaging failures, investigating root causes, and implementing mitigations.
View Full Description & ApplyYou'll be redirected to the employer's site