Senior Site Reliability Engineer (Performance and Scalability)
D
DigitalZoneTechnology
Egypt. United Arab Emirates. Turkey. PolandFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPHPPostgreSQLTypeScriptGo
Requirements
- 5+ years of experience in SRE, platform, or backend engineering.
- Proven experience with production ownership of large-scale systems operating at 10s of thousands of requests per minute.
- Track record of scaling systems through real traffic spikes.
- Experience designing and running load and failure testing programs.
- Deep AWS experience.
- Solid grasp of Postgres performance and scaling.
- Fluency with observability tooling.
- Fluency with infrastructure-as-code.
- Proficiency in scripting with Go, TypeScript, or similar languages.
- Strong communication skills with a focus on influencing and enabling other teams.
Responsibilities
- Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation.
- Establish load and failure testing as a standard engineering practice by providing teams with frameworks, tooling, and runbooks.
- Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services.
- Harden Postgres and AWS infrastructure for performance and availability while reducing toil via automation and IaC.
- Lead incident response and blameless postmortems, driving systemic fixes upstream.
- Partner with engineering teams on capacity and resilience to enable self-sufficiency in scaling systems.
View Full Description & ApplyYou'll be redirected to the employer's site