Site Reliability Engineer

New
C
CI&TCloud infrastructure
Listing location: Brazil; Workplace type: Remote; Structured job location: BrazilFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role.
Required Skills
AWSPythonBashGoGrafanaTerraform

Requirements

  • Have 3+ years of experience in site reliability engineering, DevOps, platform engineering, or production-focused software engineering.
  • Have hands-on experience administering and building with a platform such as New Relic, Grafana, Splunk, or Dynatrace, including dashboards, monitors, log management, and APM.
  • Demonstrate troubleshooting and root-cause analysis skills using logs, traces, and metrics.
  • Be proficient in at least one scripting or programming language, such as Python, TypeScript/JavaScript, Go, or Bash.
  • Be comfortable reading application code to understand failures.
  • Have experience operating services on AWS, GCP, or Azure, with a solid grasp of networking, containers, and Linux fundamentals.
  • Be familiar with infrastructure as code tools such as Terraform, CloudFormation, or Pulumi.
  • Be familiar with CI/CD tooling such as GitHub Actions.
  • Have experience with on-call responsibilities, incident response, and post-incident review processes.
  • Have clear written and verbal communication skills, including explaining reliability concerns and trade-offs to non-technical stakeholders.

Responsibilities

  • Operate the monitoring platform, maintaining dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring.
  • Analyze logs, error tracking, traces, and metrics to identify, triage, reproduce, and resolve failures, regressions, and anomalies.
  • Tune alert thresholds, monitor logic, and notification routing to reduce noise while ensuring customer-impacting issues reach the right people.
  • Instrument services with metrics, structured logs, and distributed traces, and partner with engineers to improve code observability.
  • Write post-incident reviews, track remediation items, and apply lessons learned to monitors, runbooks, and system design.
  • Define, measure, and report service level indicators for key customer-facing services.
  • Improve the reliability, scalability, and cost efficiency of cloud infrastructure, CI/CD pipelines, and release processes, automating operational work.
  • Maintain runbooks, escalation paths, and operational documentation.
  • Partner with enterprise InfoSec to remediate infrastructure security risks, including SSL/TLS, security headers, and DNS configuration findings.
  • Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now