Senior Site Reliability Engineer, NetBox Delivery
New
J
JobgetherNetwork infrastructure software
Fully remote work within the United States.Full-TimeSenior
SalaryUS base salary: $180,000–$195,000 USD. Equity opportunity. Bonus eligibility.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience in software engineering, platform engineering, site reliability engineering, or a closely related discipline.
- Required Skills
- AWSPostgreSQLDjangoKubernetesGrafanaPrometheusTerraformGitHub ActionsHelm
Requirements
- Have 5+ years of professional experience in software engineering, platform engineering, site reliability engineering, or a closely related discipline.
- Write robust, maintainable production software and own systems throughout their lifecycle.
- Bring production experience with Django and PostgreSQL at scale, including schema design, migration risks, and query-performance optimization.
- Have containerization expertise, including base image design, Python dependency management, vulnerability scanning, image signing, and software supply-chain security practices.
- Have hands-on experience with AWS, including EC2, VPC, IAM, and RDS.
- Have experience with Kubernetes and Helm, GitHub Actions, ArgoCD or FluxCD, Terraform, Prometheus, and Grafana.
- Have experience building and operating systems in AI-augmented development environments, including tools such as Claude Code and reliable agentic development workflows.
- Lead initiatives across engineering teams, communicate technical proposals clearly, and drive complex migrations or changes to completion.
- Demonstrate strong troubleshooting, analytical, and incident-management skills, with a focus on finding root causes.
- Bring clear written and verbal communication skills and a collaborative approach to cross-functional engineering work.
Responsibilities
- Own and improve the software build and release pipeline, from base image creation through availability across cloud and self-managed environments.
- Design and maintain reliable release handoffs between core engineering and deployment teams.
- Improve application performance and production reliability, including startup behavior, database performance, and PostgreSQL optimization.
- Build and maintain observability across the application and release pipeline, including monitoring, alerting, metrics, and service-level objectives.
- Investigate significant performance and reliability issues, identify root causes, and contribute fixes to application code when appropriate.
- Strengthen software supply-chain security through vulnerability scanning, image security, and artifact integrity controls.
- Support build and delivery pipeline compliance initiatives, including SOC 2 requirements.
- Participate in an on-call rotation, lead incident response, and conduct postmortems that translate findings into engineering improvements.
- Drive cross-team technical initiatives from RFCs and proposals through implementation and migration.
- Help shape the practices, tooling, and operating model of a new engineering team.
View Full Description & ApplyYou'll be redirected to the employer's site