Senior Site Reliability Engineer

New
M
MOZNEnterprise AI
Cairo, Cairo Governorate, EgyptFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years
Required Skills
AWSPythonGCPKubernetesAzureGrafanaPrometheusDatadog

Requirements

  • 3+ years building production software with LLMs including agentic workflows, function calling, and RAG.
  • Hands-on experience shipping work with agentic coding tools like Claude Code, OpenAI Codex, or Kimi K2/K3.
  • Strong Python proficiency for building agent tooling and API wrappers.
  • Real-world SRE experience, including serving as a primary on-call responder and leading incident response.
  • Application-level debugging skills with the ability to read service code and implement fixes.
  • Proven hands-on experience with Kubernetes and cloud providers (AWS/GCP/OCI/Azure).
  • Fluency with observability tools such as Prometheus, Grafana, Datadog, or ELK.
  • Knowledge of safety guardrails for autonomous systems including permissioning and rollback paths.

Responsibilities

  • Carry a normal on-call rotation and act as a hands-on responder for incidents.
  • Investigate, fix, and document incidents including debugging application code and shipping fixes directly into service repos.
  • Design, build, and ship LLM-based agents that integrate with Kubernetes, cloud APIs, and observability stacks.
  • Define tool interfaces and build safe wrappers for agent interactions with production systems.
  • Set guardrails for autonomous agent actions and manage human-in-the-loop workflows.
  • Own agent evaluation by defining metrics and building test suites against historical incidents.
  • Maintain security and compliance standards, including audit trails and least-privilege access.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now