Senior Site Reliability Engineer
New
L
LaravelCloud infrastructure
Workable locations: United States. Argentina. Brazil. United Kingdom. Portugal. Denmark, Remote, between Central Europe and US East timezones for optimal collaboration with the team.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerPHPBashKubernetesGoPrometheusTerraform
Requirements
- Bring deep experience with Linux system administration.
- Have experience with cloud platforms, specifically AWS.
- Be proficient with Kubernetes and Docker.
- Have experience managing infrastructure via Terraform.
- Be able to solve problems with software and scripting, for example PHP, Bash, or Go.
- Have experience with SLO, SLI, and SLA definition, capacity planning, and performance tuning.
- Be committed to documentation, cross-team collaboration, and an automation-first mindset.
- Laravel framework experience and familiarity with Laravel Cloud, Forge, or Vapor are highly preferred.
- Experience with Prometheus, Grafana Mimir, and Grafana Loki for metrics storage and alerting is a bonus.
Responsibilities
- Establish SRE as a core function at Laravel, building the fundamentals from the ground up.
- Design, build, and maintain multi-region Kubernetes infrastructure and global distributed systems.
- Solve operational challenges through software to reduce manual intervention for product teams.
- Design and implement monitoring, logging, and alerting systems using tools such as Prometheus, Grafana, and Loki.
- Partner with product leads and SecOps to make reliability a shared responsibility.
- Select an SLO monitoring solution and guide teams in identifying SLIs and using SLOs.
- Work with engineering teams to review existing SLAs, create SLOs, and educate teams on error budgets.
View Full Description & ApplyYou'll be redirected to the employer's site