Senior Site Reliability Engineer (SRE) - Cassandra & AWS

N
NortalSite reliability engineering
Location: Latin America; Full remote workFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
Advanced English level is required
Experience
5+ years of experience working in Site Reliability Engineering, DevOps, platform engineering or a similar production infrastructure role.
Required Skills
AWSCI/CDDevOpsDistributed Systems

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field.
  • At least 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, or a similar production infrastructure role.
  • Apache Cassandra experience, including data modeling, replication, consistency, tuning, and multi-datacenter deployments.
  • Experience operating and scaling distributed systems in production environments.
  • Strong understanding of caching architectures, performance optimization, and high-availability systems.
  • Hands-on AWS experience and familiarity with cloud-native architecture.
  • Experience designing and implementing automated CI/CD pipelines and deployment strategies for highly available systems.
  • Experience with infrastructure, configuration, and secrets as code, and a preference for automated, repeatable processes.
  • Strong understanding of observability, monitoring, alerting, performance metrics, and production troubleshooting.
  • Experience with performance testing, capacity planning, and identifying system bottlenecks.
  • Strong programming or scripting experience for automation and tooling; Java/Spring Boot experience is strongly preferred.
  • Ability to work independently, own ambiguous technical problems, provide technical leadership, and establish engineering standards and guardrails.
  • Advanced English level is required.

Responsibilities

  • Design, build, and operate highly available, fault-tolerant distributed systems and caching platforms.
  • Evaluate architecture and identify opportunities to scale the platform toward six times current traffic while maintaining performance and resiliency.
  • Improve caching strategies, including TTLs, refresh patterns, cache placement, performance, and downstream dependency management.
  • Evolve the platform toward active-active resiliency and validate failure, recovery, and capacity scenarios.
  • Design and implement automated CI/CD pipelines with rolling, blue/green, or canary deployments, automated checks, and quality gates.
  • Establish infrastructure, configuration, secrets, and application deployment practices using an everything-as-code approach.
  • Build performance and load-testing capabilities and establish performance gates for releases.
  • Develop monitoring, alerting, and observability for cache latency, throughput, availability, errors, and reliability indicators.
  • Define and improve SLIs, SLOs, error budgets, and operational KPIs where appropriate.
  • Troubleshoot production issues, conduct root-cause analysis, guide engineering teams, and drive improvements through blameless post-incident reviews.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now