Staff Platform Engineer
New
J
JobgetherHealth and Fitness
Fully remote US opportunityFull-TimeStaff
Salary155,000 - 190,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- AWSPythonTypeScriptCI/CDTerraformGitHub ActionsDatadog
Requirements
- 8+ years of experience in DevOps, platform engineering, Site Reliability Engineering, or a closely related discipline.
- Deep expertise with AWS, Terraform, ECS/Fargate, and CI/CD pipelines.
- Strong TypeScript skills with the ability to develop production-quality software and internal tooling.
- Solid understanding of AWS IAM, SSO, and observability technologies.
- Strong architectural thinking and the ability to make sound technical decisions across reliability, scalability, security, and cost.
- Ability to communicate technical concepts clearly and influence teams without direct authority.
- Strong collaboration, mentorship, and cross-functional communication skills.
- Comfortable balancing hands-on engineering work with architecture, technical planning, and mentoring.
- Experience building or improving internal developer platforms and developer experience (preferred).
- Proficiency in Python (preferred).
- Experience building LLM-powered applications or AI agents (preferred).
- Experience implementing infrastructure and monitoring practices using Datadog-as-code (preferred).
Responsibilities
- Eliminate operational toil through automation and self-service by building reusable modules and tooling that allow engineering teams to provision services, manage deployments, and resolve issues independently.
- Maintain platform health and delivery outcomes, including ownership of incident response and participation in on-call operations.
- Advance AI-native development practices by applying modern AI-powered workflows and automation to routine platform maintenance and engineering tasks.
- Optimize cloud infrastructure costs across AWS and relevant third-party providers.
- Architect and advise engineering teams on infrastructure for novel systems, including machine learning workflows, real-time event processing, and video streaming.
- Own and continuously evolve the Terraform estate at scale.
- Build and maintain internal developer platform capabilities and delivery automation using GitHub Actions and ECS.
- Own and operate core AWS platform services, including VPC, Transit Gateway, Route 53, CloudFront, IAM, SSO, and Secrets Manager.
- Drive reliability and operational excellence through Datadog monitoring and incident response.
- Provide technical leadership through RFCs, architecture and design reviews, code reviews, and cross-team technical influence.
View Full Description & ApplyYou'll be redirected to the employer's site