Senior Site Reliability Engineer
New
M
MOZNEnterprise AI
Cairo, Cairo Governorate, EgyptFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years
- Required Skills
- AWSPythonGCPKubernetesAzureGrafanaPrometheusDatadog
Requirements
- 3+ years building production software with LLMs including agentic workflows, function calling, and RAG.
- Hands-on experience shipping work with agentic coding tools like Claude Code, OpenAI Codex, or Kimi K2/K3.
- Strong Python proficiency for building agent tooling and API wrappers.
- Real-world SRE experience, including serving as a primary on-call responder and leading incident response.
- Application-level debugging skills with the ability to read service code and implement fixes.
- Proven hands-on experience with Kubernetes and cloud providers (AWS/GCP/OCI/Azure).
- Fluency with observability tools such as Prometheus, Grafana, Datadog, or ELK.
- Knowledge of safety guardrails for autonomous systems including permissioning and rollback paths.
Responsibilities
- Carry a normal on-call rotation and act as a hands-on responder for incidents.
- Investigate, fix, and document incidents including debugging application code and shipping fixes directly into service repos.
- Design, build, and ship LLM-based agents that integrate with Kubernetes, cloud APIs, and observability stacks.
- Define tool interfaces and build safe wrappers for agent interactions with production systems.
- Set guardrails for autonomous agent actions and manage human-in-the-loop workflows.
- Own agent evaluation by defining metrics and building test suites against historical incidents.
- Maintain security and compliance standards, including audit trails and least-privilege access.
View Full Description & ApplyYou'll be redirected to the employer's site