Senior Site Reliability Engineer

New
U
UpsunCloud platform
Location: Remote • Australia; we are currently focused on hiring for this role in Western Australia., One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps
Required Skills
AWSPythonGCPAzureGoPrometheusTerraformAnsible

Requirements

  • Have 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps.
  • Have proven experience owning reliability for production platforms at scale.
  • Be proficient in Go or Python for building automation tools, custom controllers, or SRE platform components beyond basic shell scripting.
  • Have advanced hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
  • Have deep expertise with cloud providers such as AWS, GCP, Azure, or OpenStack.
  • Have experience with cloud SDK tooling and declarative infrastructure tools such as Terraform for distributed systems.
  • Demonstrate the ability to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives with minimal guidance.
  • Participate in on-call rotation one week every 4–5 weeks, including a weekend shift from 02:00 to 10:00 UTC.
  • Bonus: experience with custom-built orchestration, edge, storage, or operational tooling.
  • Bonus: experience with Docker and production Kubernetes cluster management or containerized deployment architectures.
  • Bonus: familiarity with PaaS architectures or developer-facing cloud platforms.

Responsibilities

  • Architect monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack, and establish actionable SLIs and SLOs.
  • Design and implement automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
  • Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases.
  • Guide high-priority incident triage, lead blameless post-mortems, and implement preventative measures.
  • Partner with product and software engineering teams to incorporate SRE practices into product roadmaps.
  • Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
  • Own reliability engineering workstreams across multi-cloud environments and the software delivery lifecycle.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now