Senior Site Reliability Engineer

New
L
LaravelCloud infrastructure
Workable locations: United States. Argentina. Brazil. United Kingdom. Portugal. Denmark, Remote, between Central Europe and US East timezones for optimal collaboration with the team.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSDockerPHPBashKubernetesGoPrometheusTerraform

Requirements

  • Bring deep experience with Linux system administration.
  • Have experience with cloud platforms, specifically AWS.
  • Be proficient with Kubernetes and Docker.
  • Have experience managing infrastructure via Terraform.
  • Be able to solve problems with software and scripting, for example PHP, Bash, or Go.
  • Have experience with SLO, SLI, and SLA definition, capacity planning, and performance tuning.
  • Be committed to documentation, cross-team collaboration, and an automation-first mindset.
  • Laravel framework experience and familiarity with Laravel Cloud, Forge, or Vapor are highly preferred.
  • Experience with Prometheus, Grafana Mimir, and Grafana Loki for metrics storage and alerting is a bonus.

Responsibilities

  • Establish SRE as a core function at Laravel, building the fundamentals from the ground up.
  • Design, build, and maintain multi-region Kubernetes infrastructure and global distributed systems.
  • Solve operational challenges through software to reduce manual intervention for product teams.
  • Design and implement monitoring, logging, and alerting systems using tools such as Prometheus, Grafana, and Loki.
  • Partner with product leads and SecOps to make reliability a shared responsibility.
  • Select an SLO monitoring solution and guide teams in identifying SLIs and using SLOs.
  • Work with engineering teams to review existing SLAs, create SLOs, and educate teams on error budgets.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now