Senior Site Reliability Engineer (Performance and Scalability)

New
D
DigitalZoneTechnology
Egypt. United Arab Emirates. Turkey. PolandFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPHPTypeScriptGoPostgres

Requirements

  • 5+ years in SRE, platform, or backend engineering.
  • Strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
  • Track record of scaling systems through real traffic spikes.
  • Experience designing and running load and failure testing programs.
  • Deep AWS experience.
  • Solid grasp of Postgres performance and scaling.
  • Fluency with observability tooling.
  • Fluency with infrastructure-as-code.
  • Proficiency in scripting with Go, TypeScript, or similar languages.
  • Calm, systematic approach to incident management.

Responsibilities

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes.
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services.
  • Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services.
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
  • Lead incident response and blameless postmortems, driving systemic fixes upstream.
  • Partner with engineering teams early on capacity and resilience, acting as a multiplier to make them self-sufficient.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now