Senior Site Reliability Engineer - AI Platform

New
J
JobgetherAI Platform
Work from anywhere in BrazilFull-TimeSenior
SalaryCompetitive salary
Apply NowOpens the employer's application page

Job Details

Languages
Portuguese, English
Experience
7 years
Required Skills
KubernetesCI/CDSoftware EngineeringDistributed Systems

Requirements

  • At least 7 years of professional experience as an SRE or Software Engineer.
  • Strong interest in applied AI, including an understanding of AI models, costs, observability, tokens, and practical use cases.
  • Solid experience with Infrastructure as Code and modern platform technologies such as Kubernetes, containers, Vault, and API Ingress.
  • Strong software engineering fundamentals, including version control, automated testing, deployment automation, code review, and technical design documentation.
  • Deep knowledge of modern CI/CD practices and continuous integration and deployment workflows.
  • Proven experience designing and building software and systems architectures completely from scratch.
  • Experience working with high-scale systems, performance optimization, and distributed system architectures.
  • Strong knowledge of database modeling, structuring, and evolution.
  • Hands-on experience working with agile methodologies such as Kanban or Scrum.
  • Professional proficiency in both Portuguese and English.

Responsibilities

  • Design and build foundational platform architectures from scratch to support scalable AI adoption and hyper-automation.
  • Develop reliable, secure, and highly scalable infrastructure that enables teams to build and operate AI-powered automations with autonomy.
  • Establish platform guardrails, protection mechanisms, and engineering standards covering reliability, security, observability, and operational performance.
  • Build and maintain infrastructure using Infrastructure as Code and modern platform technologies, including Kubernetes, containers, Vault, and API Ingress.
  • Develop and optimize CI/CD processes that enable consistent, automated, and reliable software delivery.
  • Address high-scale performance challenges and design solutions for complex distributed systems.
  • Contribute to the evolution of database architecture, including modeling, structuring, and long-term scalability.
  • Apply software engineering best practices across version control, testing, deployment automation, code reviews, and technical design documentation.
  • Explore and incorporate applied AI capabilities by considering model behavior, costs, token usage, observability, and relevant use cases.
  • Help shape the technical direction of the platform while mentoring, technically leading, or influencing other SRE and engineering professionals when needed.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive salary
Apply Now