Platform Site Reliability Engineer
New
F
First DuePublic safety software
Remote - US OnlyFull-TimeSenior
Salary$165,000 + Bonus
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a related role.
- Required Skills
- AWSPythonBashKubernetesCI/CDTerraformNetworking
Requirements
- Have 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a related role.
- Bring strong experience managing and supporting production cloud environments.
- Have experience building and maintaining CI/CD pipelines and deployment automation.
- Understand infrastructure-as-code principles and tooling.
- Have experience troubleshooting production systems and resolving complex operational issues.
- Know networking, security, scalability, and high-availability architectures.
- Have strong scripting or automation experience using Python, Bash, PowerShell, or similar.
- Have experience working with software engineering teams throughout the software development lifecycle.
- Be able to balance operational stability with delivery speed and business priorities.
- Preferred: Experience supporting high-growth SaaS platforms and cloud providers such as AWS, Azure, or Google Cloud Platform.
- Preferred: Experience with Kubernetes, Docker, Terraform, Pulumi, CloudFormation, CI/CD platforms, or observability tools.
- Preferred: Experience with database operations, database scaling, query optimization, BCDR, or zero-downtime deployments.
Responsibilities
- Design, implement, and maintain scalable, secure, highly available cloud infrastructure.
- Build and manage infrastructure-as-code solutions for repeatable, reliable deployments.
- Design and maintain CI/CD pipelines, automate operational tasks, and improve deployment processes.
- Monitor production systems and improve availability, performance, and reliability.
- Participate in incident response, troubleshooting, root cause analysis, and post-incident reviews.
- Develop and maintain monitoring, alerting, logging, and observability solutions.
- Support disaster recovery, backup, business continuity, vulnerability remediation, patch management, and access controls.
- Partner with Engineering, Product, QA, and Security teams on platform initiatives and operational standards.
- Document infrastructure and operational processes, and contribute to architectural reviews and platform strategy.
View Full Description & ApplyYou'll be redirected to the employer's site