Senior Site Reliability Engineer (Performance and Scalability)
New
D
DigitalZoneTechnology
Egypt. United Arab Emirates. Turkey. PolandFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPHPTypeScriptGoPostgres
Requirements
- 5+ years in SRE, platform, or backend engineering.
- Strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
- Track record of scaling systems through real traffic spikes.
- Experience designing and running load and failure testing programs.
- Deep AWS experience.
- Solid grasp of Postgres performance and scaling.
- Fluency with observability tooling.
- Fluency with infrastructure-as-code.
- Proficiency in scripting with Go, TypeScript, or similar languages.
- Calm, systematic approach to incident management.
Responsibilities
- Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes.
- Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services.
- Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services.
- Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
- Lead incident response and blameless postmortems, driving systemic fixes upstream.
- Partner with engineering teams early on capacity and resilience, acting as a multiplier to make them self-sufficient.
View Full Description & ApplyYou'll be redirected to the employer's site