Senior DevOps Engineer, Observability
New
J
JobgetherB2B software
Fully remote working arrangement for candidates based in Brazil/LATAM.Full-TimeSenior
Salary75,000 - 85,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field; at least 2 years of experience working within a B2B software startup or similarly fast-paced technology environment.
- Required Skills
- AWSPythonKubernetesGoGrafanaPrometheusCI/CDTerraformGitHub ActionsHelm
Requirements
- Have 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field.
- Have at least 2 years of experience in a B2B software startup or similarly fast-paced technology environment.
- Bring hands-on AWS experience, including EC2, VPC, IAM, RDS, and EKS.
- Have solid Kubernetes experience and infrastructure-as-code expertise using Terraform, Helm, or comparable tools.
- Have experience designing and maintaining CI/CD pipelines and automation tooling; GitHub Actions is preferred.
- Be familiar with observability platforms such as Prometheus, Grafana, Mimir, or Loki.
- Be proficient in at least one scripting or programming language, such as Python, Go, or shell scripting.
- Have experience with either CDC and event-streaming systems or large, multi-tenant observability platforms covering ingestion, analysis, and alerting.
- Have strong communication, collaboration, troubleshooting, and technical documentation skills.
- Be able to work effectively amid ambiguity and evolving priorities.
- SOC 2 or security-focused environment experience is a plus; familiarity with MQTT, AMQP, network automation, open-source technologies, or AI-assisted development tools is beneficial.
Responsibilities
- Design, build, and maintain infrastructure supporting SaaS and on-premise engineering environments.
- Operate and optimize AWS infrastructure, including EC2, VPC, IAM, RDS, and EKS, for cost efficiency, security, scalability, reliability, and performance.
- Develop and improve infrastructure-as-code using Terraform and Helm.
- Build and maintain CI/CD pipelines and developer automation, particularly using GitHub Actions.
- Contribute to internal platform tooling and developer self-service capabilities.
- Improve monitoring, alerting, instrumentation, incident response, and SLO management.
- Collaborate with stream-aligned product teams to understand infrastructure needs and improve internal platforms.
- Support security and compliance initiatives, including implementing and maintaining SOC 2 controls.
- Create and maintain technical documentation, onboarding resources, and internal support processes.
- Participate in an on-call rotation and contribute to infrastructure initiatives involving CDC, event streaming, or large-scale multi-tenant observability systems.
View Full Description & ApplyYou'll be redirected to the employer's site