Site Reliability Engineer
New
C
CI&TCloud infrastructure
Listing location: Brazil; Workplace type: Remote; Structured job location: BrazilFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role.
- Required Skills
- AWSPythonBashGoGrafanaTerraform
Requirements
- Have 3+ years of experience in site reliability engineering, DevOps, platform engineering, or production-focused software engineering.
- Have hands-on experience administering and building with a platform such as New Relic, Grafana, Splunk, or Dynatrace, including dashboards, monitors, log management, and APM.
- Demonstrate troubleshooting and root-cause analysis skills using logs, traces, and metrics.
- Be proficient in at least one scripting or programming language, such as Python, TypeScript/JavaScript, Go, or Bash.
- Be comfortable reading application code to understand failures.
- Have experience operating services on AWS, GCP, or Azure, with a solid grasp of networking, containers, and Linux fundamentals.
- Be familiar with infrastructure as code tools such as Terraform, CloudFormation, or Pulumi.
- Be familiar with CI/CD tooling such as GitHub Actions.
- Have experience with on-call responsibilities, incident response, and post-incident review processes.
- Have clear written and verbal communication skills, including explaining reliability concerns and trade-offs to non-technical stakeholders.
Responsibilities
- Operate the monitoring platform, maintaining dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring.
- Analyze logs, error tracking, traces, and metrics to identify, triage, reproduce, and resolve failures, regressions, and anomalies.
- Tune alert thresholds, monitor logic, and notification routing to reduce noise while ensuring customer-impacting issues reach the right people.
- Instrument services with metrics, structured logs, and distributed traces, and partner with engineers to improve code observability.
- Write post-incident reviews, track remediation items, and apply lessons learned to monitors, runbooks, and system design.
- Define, measure, and report service level indicators for key customer-facing services.
- Improve the reliability, scalability, and cost efficiency of cloud infrastructure, CI/CD pipelines, and release processes, automating operational work.
- Maintain runbooks, escalation paths, and operational documentation.
- Partner with enterprise InfoSec to remediate infrastructure security risks, including SSL/TLS, security headers, and DNS configuration findings.
- Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features.
View Full Description & ApplyYou'll be redirected to the employer's site