Senior Site Reliability Engineer - AI Platform
New
J
JobgetherAI Platform
Work from anywhere in BrazilFull-TimeSenior
SalaryCompetitive salary
Apply NowOpens the employer's application page
Job Details
- Languages
- Portuguese, English
- Experience
- 7 years
- Required Skills
- KubernetesCI/CDSoftware EngineeringDistributed Systems
Requirements
- At least 7 years of professional experience as an SRE or Software Engineer.
- Strong interest in applied AI, including an understanding of AI models, costs, observability, tokens, and practical use cases.
- Solid experience with Infrastructure as Code and modern platform technologies such as Kubernetes, containers, Vault, and API Ingress.
- Strong software engineering fundamentals, including version control, automated testing, deployment automation, code review, and technical design documentation.
- Deep knowledge of modern CI/CD practices and continuous integration and deployment workflows.
- Proven experience designing and building software and systems architectures completely from scratch.
- Experience working with high-scale systems, performance optimization, and distributed system architectures.
- Strong knowledge of database modeling, structuring, and evolution.
- Hands-on experience working with agile methodologies such as Kanban or Scrum.
- Professional proficiency in both Portuguese and English.
Responsibilities
- Design and build foundational platform architectures from scratch to support scalable AI adoption and hyper-automation.
- Develop reliable, secure, and highly scalable infrastructure that enables teams to build and operate AI-powered automations with autonomy.
- Establish platform guardrails, protection mechanisms, and engineering standards covering reliability, security, observability, and operational performance.
- Build and maintain infrastructure using Infrastructure as Code and modern platform technologies, including Kubernetes, containers, Vault, and API Ingress.
- Develop and optimize CI/CD processes that enable consistent, automated, and reliable software delivery.
- Address high-scale performance challenges and design solutions for complex distributed systems.
- Contribute to the evolution of database architecture, including modeling, structuring, and long-term scalability.
- Apply software engineering best practices across version control, testing, deployment automation, code reviews, and technical design documentation.
- Explore and incorporate applied AI capabilities by considering model behavior, costs, token usage, observability, and relevant use cases.
- Help shape the technical direction of the platform while mentoring, technically leading, or influencing other SRE and engineering professionals when needed.
View Full Description & ApplyYou'll be redirected to the employer's site