Senior Platform Architect

New
J
JobgetherCloud Infrastructure
Remote opportunity in LATAM, ESTFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSGCPKubernetesAzureCI/CDDevOpsTerraform

Requirements

  • 5+ years of experience in platform engineering, SRE, or infrastructure.
  • Significant hands-on experience operating production systems at scale.
  • Strong SRE/DevOps foundation including SLOs, post-mortems, and incident management.
  • Deep expertise with Terraform, including complex state and reusable modules.
  • Strong production experience with GitOps tools such as ArgoCD or Flux.
  • Advanced Kubernetes knowledge, including cluster operations and troubleshooting.
  • Strong experience with at least one major public cloud (AWS, Azure, or GCP).
  • Experience building CI/CD pipelines using GitHub Actions, Cloud Build, or GitLab CI.
  • Automation-first mindset with experience designing systems to reduce manual toil.
  • Experience directing and reviewing output from agentic coding and AI development tools.
  • Excellent communication skills for technical and leadership audiences.

Responsibilities

  • Design and operate scalable backend and AI infrastructure supporting real-time and batch workloads.
  • Establish architectural patterns prioritizing performance, reliability, scalability, and multi-tenant operation.
  • Maintain deployment workflows covering versioning, staged rollouts, automated releases, and rollback strategies.
  • Build and operate production LLM and agentic systems, integrating model providers, APIs, and gateways.
  • Manage reliability for AI workloads including rate limits, quotas, guardrails, and operational resilience.
  • Extend infrastructure-as-code practices using Terraform and reusable modules.
  • Maintain GitOps-based deployment workflows using tools such as ArgoCD or Flux.
  • Operate Kubernetes workloads, focusing on scaling, isolation, and infrastructure capacity management.
  • Improve observability through metrics, logging, tracing, SLOs, and incident response.
  • Define operational standards to support long-term platform growth.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now