Senior Software Engineer (SRE), Cloud Compute

J
JobgetherCloud Infrastructure
BrazilFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSKubernetesDistributed SystemsNetworking

Requirements

  • Proven experience as a backend, platform, infrastructure, or Site Reliability Engineer, preferably at mid-to-senior level.
  • Strong hands-on experience operating and evolving Kubernetes infrastructure, particularly Amazon EKS, in production environments.
  • Practical experience with AWS services such as VPC, IAM, load balancers, Auto Scaling, and Elastic Beanstalk.
  • Solid understanding of distributed systems and networking concepts, including TCP/IP, DNS, proxies, and inter-service communication.
  • Experience working with production systems and reliability practices, including availability, resilience, incident response, and on-call operations.
  • Experience migrating workloads between compute platforms or modernizing legacy infrastructure is highly valuable.
  • Familiarity with GitOps and Infrastructure as Code, particularly ArgoCD, Pulumi, or equivalent technologies.
  • Ability to work comfortably with systems undergoing modernization while balancing technical debt, reliability, and delivery needs.
  • Strong communication and collaboration skills, with the ability to explain technical decisions and align priorities with product and engineering teams.
  • Comfortable using AI-assisted development tools during technical work while maintaining full ownership and understanding of the code and solutions produced.

Responsibilities

  • Lead and execute the migration of production workloads from legacy compute environments to EKS, including networking, IAM configuration, and functional parity.
  • Design, maintain, and evolve shared compute and networking components, including private connectivity, load balancers, and internal platform APIs.
  • Operate critical production infrastructure, maintaining high availability, resilience, scalability, and performance while participating in on-call rotations.
  • Evolve GitOps and infrastructure-as-code practices, including the maintenance and improvement of deployment and rollout tooling.
  • Investigate and resolve complex infrastructure and production issues while contributing to long-term reliability improvements.
  • Serve as a technical point of contact for internal engineering teams that depend on shared compute infrastructure.
  • Collaborate across platform and product engineering teams to negotiate technical priorities, improve self-service capabilities, and reduce infrastructure complexity and technical debt.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now