Senior Site Reliability Engineer (SRE) - Cassandra & AWS
N
NortalSite reliability engineering
Location: Latin America; Full remote workFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- Advanced English level is required
- Experience
- 5+ years of experience working in Site Reliability Engineering, DevOps, platform engineering or a similar production infrastructure role.
- Required Skills
- AWSCI/CDDevOpsDistributed Systems
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field.
- At least 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, or a similar production infrastructure role.
- Apache Cassandra experience, including data modeling, replication, consistency, tuning, and multi-datacenter deployments.
- Experience operating and scaling distributed systems in production environments.
- Strong understanding of caching architectures, performance optimization, and high-availability systems.
- Hands-on AWS experience and familiarity with cloud-native architecture.
- Experience designing and implementing automated CI/CD pipelines and deployment strategies for highly available systems.
- Experience with infrastructure, configuration, and secrets as code, and a preference for automated, repeatable processes.
- Strong understanding of observability, monitoring, alerting, performance metrics, and production troubleshooting.
- Experience with performance testing, capacity planning, and identifying system bottlenecks.
- Strong programming or scripting experience for automation and tooling; Java/Spring Boot experience is strongly preferred.
- Ability to work independently, own ambiguous technical problems, provide technical leadership, and establish engineering standards and guardrails.
- Advanced English level is required.
Responsibilities
- Design, build, and operate highly available, fault-tolerant distributed systems and caching platforms.
- Evaluate architecture and identify opportunities to scale the platform toward six times current traffic while maintaining performance and resiliency.
- Improve caching strategies, including TTLs, refresh patterns, cache placement, performance, and downstream dependency management.
- Evolve the platform toward active-active resiliency and validate failure, recovery, and capacity scenarios.
- Design and implement automated CI/CD pipelines with rolling, blue/green, or canary deployments, automated checks, and quality gates.
- Establish infrastructure, configuration, secrets, and application deployment practices using an everything-as-code approach.
- Build performance and load-testing capabilities and establish performance gates for releases.
- Develop monitoring, alerting, and observability for cache latency, throughput, availability, errors, and reliability indicators.
- Define and improve SLIs, SLOs, error budgets, and operational KPIs where appropriate.
- Troubleshoot production issues, conduct root-cause analysis, guide engineering teams, and drive improvements through blameless post-incident reviews.
View Full Description & ApplyYou'll be redirected to the employer's site