Senior Site Reliability Engineer (Performance and Scalability)

D
DigitalZoneTechnology
Egypt. United Arab Emirates. Turkey. PolandFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPHPPostgreSQLTypeScriptGo

Requirements

  • 5+ years of experience in SRE, platform, or backend engineering.
  • Proven experience with production ownership of large-scale systems operating at 10s of thousands of requests per minute.
  • Track record of scaling systems through real traffic spikes.
  • Experience designing and running load and failure testing programs.
  • Deep AWS experience.
  • Solid grasp of Postgres performance and scaling.
  • Fluency with observability tooling.
  • Fluency with infrastructure-as-code.
  • Proficiency in scripting with Go, TypeScript, or similar languages.
  • Strong communication skills with a focus on influencing and enabling other teams.

Responsibilities

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation.
  • Establish load and failure testing as a standard engineering practice by providing teams with frameworks, tooling, and runbooks.
  • Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services.
  • Harden Postgres and AWS infrastructure for performance and availability while reducing toil via automation and IaC.
  • Lead incident response and blameless postmortems, driving systemic fixes upstream.
  • Partner with engineering teams on capacity and resilience to enable self-sufficiency in scaling systems.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now