Senior DevOps Engineer, Observability
New
N
NetBox LabsNetwork observability
Location: LATAM, Remote
Secondary Locations: US Remote, UK, RemoteFull-TimeSenior
Salary$75K - $185K; $75K – $185K • Offers Equity • Offers Bonus • Multiple Ranges; Remote, LATAM: Remote, LATAM $75K – $85K; Remote, UK: Remote, UK £85K – £100K • Offers Equity • Offers Bonus; Remote, US: Remote, US $155K – $185K • Offers Equity • Offers Bonus
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in DevOps, SRE, or platform engineering roles; 2+ years of experience at a B2B software startup.
- Required Skills
- AWSPythonKubernetesGoGrafanaPrometheusTerraformGitHub ActionsHelm
Requirements
- Have 5+ years of experience in DevOps, SRE, or platform engineering roles.
- Have 2+ years of experience at a B2B software startup.
- Bring strong experience with AWS, including EC2, VPC, IAM, and RDS.
- Have experience with EKS/Kubernetes and infrastructure as code using Terraform and Helm.
- Have experience with CI/CD pipelines and automation tooling; GitHub Actions is preferred.
- Be familiar with observability tools such as Prometheus, Grafana, Mimir, or Loki, or similar tools.
- Be proficient in Python, Go, or shell scripting.
- Have experience with Change Data Capture (CDC) and event streaming systems, or with scaling large, multi-tenant observability systems including ingest, analysis, and alerting.
- Be comfortable working in a fast-paced, ambiguous startup environment.
- Have strong communication and documentation skills.
- Experience in a SOC 2-compliant or security-focused environment is a plus.
- Familiarity with MQTT, AMQP, NetBox, network automation tooling, open-source contributions, or AI tools is a plus.
Responsibilities
- Design, build, and maintain infrastructure systems supporting SaaS and on-premise engineering needs.
- Operate and optimize AWS and other cloud infrastructure for cost efficiency, security, and performance.
- Contribute to internal platform tooling, GitHub Actions CI/CD automation, and developer self-service capabilities.
- Enhance observability and incident response systems, including monitoring, alerting, and SLOs.
- Collaborate with stream-aligned product teams to understand their needs and improve internal platforms.
- Help enforce and improve security and compliance standards, including SOC 2 controls.
- Contribute to documentation, onboarding materials, and internal support processes.
- Participate in the on-call rotation.
View Full Description & ApplyYou'll be redirected to the employer's site