Senior Platform Architect
New
J
JobgetherCloud Infrastructure
Remote opportunity in LATAM, ESTFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSGCPKubernetesAzureCI/CDDevOpsTerraform
Requirements
- 5+ years of experience in platform engineering, SRE, or infrastructure.
- Significant hands-on experience operating production systems at scale.
- Strong SRE/DevOps foundation including SLOs, post-mortems, and incident management.
- Deep expertise with Terraform, including complex state and reusable modules.
- Strong production experience with GitOps tools such as ArgoCD or Flux.
- Advanced Kubernetes knowledge, including cluster operations and troubleshooting.
- Strong experience with at least one major public cloud (AWS, Azure, or GCP).
- Experience building CI/CD pipelines using GitHub Actions, Cloud Build, or GitLab CI.
- Automation-first mindset with experience designing systems to reduce manual toil.
- Experience directing and reviewing output from agentic coding and AI development tools.
- Excellent communication skills for technical and leadership audiences.
Responsibilities
- Design and operate scalable backend and AI infrastructure supporting real-time and batch workloads.
- Establish architectural patterns prioritizing performance, reliability, scalability, and multi-tenant operation.
- Maintain deployment workflows covering versioning, staged rollouts, automated releases, and rollback strategies.
- Build and operate production LLM and agentic systems, integrating model providers, APIs, and gateways.
- Manage reliability for AI workloads including rate limits, quotas, guardrails, and operational resilience.
- Extend infrastructure-as-code practices using Terraform and reusable modules.
- Maintain GitOps-based deployment workflows using tools such as ArgoCD or Flux.
- Operate Kubernetes workloads, focusing on scaling, isolation, and infrastructure capacity management.
- Improve observability through metrics, logging, tracing, SLOs, and incident response.
- Define operational standards to support long-term platform growth.
View Full Description & ApplyYou'll be redirected to the employer's site