Senior Site Reliability Engineer
New
U
UpsunCloud platform
Location: Remote • Australia; we are currently focused on hiring for this role in Western Australia., One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.Full-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps
- Required Skills
- AWSPythonGCPMicrosoft AzureGoGrafanaPrometheusLinuxTerraform
Requirements
- Have 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps.
- Have proven experience owning reliability for production platforms at scale.
- Demonstrate strong proficiency in Go or Python for building automation tools, custom controllers, or SRE platform components.
- Bring advanced hands-on knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
- Have deep expertise with cloud providers such as AWS, GCP, Azure, or OpenStack.
- Have experience with cloud SDK-based custom tooling and declarative infrastructure tools such as Terraform.
- Be able to anticipate operational risks, make architectural trade-offs, and lead infrastructure initiatives with minimal guidance.
- Participate in an on-call rotation: one week every 4–5 weeks, 02:00–10:00 UTC, including a weekend shift.
- Bonus: experience with custom orchestration, edge, storage, or operational tooling.
- Bonus: experience with Docker and production Kubernetes cluster management or containerized deployments.
- Bonus: familiarity with PaaS architectures or developer-facing cloud platforms.
Responsibilities
- Architect system monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack, and establish actionable SLIs/SLOs.
- Design and implement automated infrastructure and workflows using Terraform and Ansible across AWS, GCP, and Azure.
- Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases.
- Guide high-priority incident triage, lead blameless post-mortems, and implement preventative measures.
- Partner with product and software engineering teams to incorporate SRE practices into product roadmaps.
- Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration.
- Own reliability engineering workstreams across multi-cloud environments.
View Full Description & ApplyYou'll be redirected to the employer's site