Senior DevOps Engineer, Observability

New
J
JobgetherB2B software
Fully remote working arrangement for candidates based in Brazil/LATAM.Full-TimeSenior
Salary75,000 - 85,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field; at least 2 years of experience working within a B2B software startup or similarly fast-paced technology environment.
Required Skills
AWSPythonKubernetesGoGrafanaPrometheusCI/CDTerraformGitHub ActionsHelm

Requirements

  • Have 5+ years of professional experience in DevOps, Site Reliability Engineering (SRE), platform engineering, or a closely related field.
  • Have at least 2 years of experience in a B2B software startup or similarly fast-paced technology environment.
  • Bring hands-on AWS experience, including EC2, VPC, IAM, RDS, and EKS.
  • Have solid Kubernetes experience and infrastructure-as-code expertise using Terraform, Helm, or comparable tools.
  • Have experience designing and maintaining CI/CD pipelines and automation tooling; GitHub Actions is preferred.
  • Be familiar with observability platforms such as Prometheus, Grafana, Mimir, or Loki.
  • Be proficient in at least one scripting or programming language, such as Python, Go, or shell scripting.
  • Have experience with either CDC and event-streaming systems or large, multi-tenant observability platforms covering ingestion, analysis, and alerting.
  • Have strong communication, collaboration, troubleshooting, and technical documentation skills.
  • Be able to work effectively amid ambiguity and evolving priorities.
  • SOC 2 or security-focused environment experience is a plus; familiarity with MQTT, AMQP, network automation, open-source technologies, or AI-assisted development tools is beneficial.

Responsibilities

  • Design, build, and maintain infrastructure supporting SaaS and on-premise engineering environments.
  • Operate and optimize AWS infrastructure, including EC2, VPC, IAM, RDS, and EKS, for cost efficiency, security, scalability, reliability, and performance.
  • Develop and improve infrastructure-as-code using Terraform and Helm.
  • Build and maintain CI/CD pipelines and developer automation, particularly using GitHub Actions.
  • Contribute to internal platform tooling and developer self-service capabilities.
  • Improve monitoring, alerting, instrumentation, incident response, and SLO management.
  • Collaborate with stream-aligned product teams to understand infrastructure needs and improve internal platforms.
  • Support security and compliance initiatives, including implementing and maintaining SOC 2 controls.
  • Create and maintain technical documentation, onboarding resources, and internal support processes.
  • Participate in an on-call rotation and contribute to infrastructure initiatives involving CDC, event streaming, or large-scale multi-tenant observability systems.
View Full Description & ApplyYou'll be redirected to the employer's site
75,000 - 85,000 USD per year
Apply Now