Senior Software Engineer (SRE), Cloud Compute
J
JobgetherCloud Infrastructure
BrazilFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSKubernetesDistributed SystemsNetworking
Requirements
- Proven experience as a backend, platform, infrastructure, or Site Reliability Engineer, preferably at mid-to-senior level.
- Strong hands-on experience operating and evolving Kubernetes infrastructure, particularly Amazon EKS, in production environments.
- Practical experience with AWS services such as VPC, IAM, load balancers, Auto Scaling, and Elastic Beanstalk.
- Solid understanding of distributed systems and networking concepts, including TCP/IP, DNS, proxies, and inter-service communication.
- Experience working with production systems and reliability practices, including availability, resilience, incident response, and on-call operations.
- Experience migrating workloads between compute platforms or modernizing legacy infrastructure is highly valuable.
- Familiarity with GitOps and Infrastructure as Code, particularly ArgoCD, Pulumi, or equivalent technologies.
- Ability to work comfortably with systems undergoing modernization while balancing technical debt, reliability, and delivery needs.
- Strong communication and collaboration skills, with the ability to explain technical decisions and align priorities with product and engineering teams.
- Comfortable using AI-assisted development tools during technical work while maintaining full ownership and understanding of the code and solutions produced.
Responsibilities
- Lead and execute the migration of production workloads from legacy compute environments to EKS, including networking, IAM configuration, and functional parity.
- Design, maintain, and evolve shared compute and networking components, including private connectivity, load balancers, and internal platform APIs.
- Operate critical production infrastructure, maintaining high availability, resilience, scalability, and performance while participating in on-call rotations.
- Evolve GitOps and infrastructure-as-code practices, including the maintenance and improvement of deployment and rollout tooling.
- Investigate and resolve complex infrastructure and production issues while contributing to long-term reliability improvements.
- Serve as a technical point of contact for internal engineering teams that depend on shared compute infrastructure.
- Collaborate across platform and product engineering teams to negotiate technical priorities, improve self-service capabilities, and reduce infrastructure complexity and technical debt.
View Full Description & ApplyYou'll be redirected to the employer's site