Senior Site Reliability Engineer
New
U
UpsunCloud platform
Location: Remote • Australia; we are currently focused on hiring for this role in Western Australia., One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps
- Required Skills
- AWSPythonGCPAzureGoPrometheusTerraformAnsible
Requirements
- Have 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps.
- Have proven experience owning reliability for production platforms at scale.
- Be proficient in Go or Python for building automation tools, custom controllers, or SRE platform components beyond basic shell scripting.
- Have advanced hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
- Have deep expertise with cloud providers such as AWS, GCP, Azure, or OpenStack.
- Have experience with cloud SDK tooling and declarative infrastructure tools such as Terraform for distributed systems.
- Demonstrate the ability to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives with minimal guidance.
- Participate in on-call rotation one week every 4–5 weeks, including a weekend shift from 02:00 to 10:00 UTC.
- Bonus: experience with custom-built orchestration, edge, storage, or operational tooling.
- Bonus: experience with Docker and production Kubernetes cluster management or containerized deployment architectures.
- Bonus: familiarity with PaaS architectures or developer-facing cloud platforms.
Responsibilities
- Architect monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack, and establish actionable SLIs and SLOs.
- Design and implement automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
- Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases.
- Guide high-priority incident triage, lead blameless post-mortems, and implement preventative measures.
- Partner with product and software engineering teams to incorporate SRE practices into product roadmaps.
- Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
- Own reliability engineering workstreams across multi-cloud environments and the software delivery lifecycle.
View Full Description & ApplyYou'll be redirected to the employer's site