Senior Site Reliability Engineer

New
U
UpsunCloud platform
Location: Remote • Australia; we are currently focused on hiring for this role in Western Australia., One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps
Required Skills
AWSPythonGCPMicrosoft AzureGoGrafanaPrometheusLinuxTerraform

Requirements

  • Have 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps.
  • Have proven experience owning reliability for production platforms at scale.
  • Demonstrate strong proficiency in Go or Python for building automation tools, custom controllers, or SRE platform components.
  • Bring advanced hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
  • Have deep expertise with cloud providers such as AWS, GCP, Azure, or OpenStack.
  • Have experience with cloud SDK-based custom tooling and declarative infrastructure tools such as Terraform.
  • Be able to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives with minimal guidance.
  • Participate in an on-call rotation: one week every 4–5 weeks, 02:00–10:00 UTC, including a weekend shift.
  • Bonus: experience with custom orchestration, edge, storage, or operational tooling.
  • Bonus: experience with Docker and production Kubernetes cluster management or containerized deployments.
  • Bonus: familiarity with PaaS architectures or developer-facing cloud platforms.

Responsibilities

  • Architect system monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack, and establish actionable SLIs/SLOs.
  • Design and implement automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
  • Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases.
  • Guide high-priority incident triage, lead blameless post-mortems, and implement preventative measures.
  • Partner with product and software engineering teams to incorporate SRE practices into product roadmaps.
  • Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
  • Own reliability engineering workstreams across multi-cloud environments.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now