Staff DevOps Automation Engineer
New
J
JobgetherTechnology
United StatesFull-TimeStaff
SalaryAnnual salary range of $180,000–$220,000.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSDockerPythonKubernetesAzureGoTerraformGitHub Actions
Requirements
- 5+ years of experience in software, platform, or infrastructure engineering, including hands-on experience supporting data and/or AI/ML infrastructure.
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- Deep hands-on expertise with at least one major cloud platform, particularly Azure or AWS, along with working familiarity with multi-cloud environments and cloud-based data and AI services.
- Strong experience with Infrastructure as Code tools such as Terraform or Pulumi, CI/CD platforms such as GitHub Actions or GitLab CI, and containerization and orchestration technologies including Docker and Kubernetes.
- Proficiency in at least one programming or scripting language such as Python, Go, or Bash for automation and infrastructure tooling.
- Strong understanding of cloud security and networking fundamentals, with demonstrated experience provisioning secure resources and operating production systems.
- Experience with monitoring, observability, incident response, reliability engineering, and production operations.
- Proven ability to establish technical direction at platform scale, influence architecture and design decisions across teams, and develop engineering standards.
- Strong mentoring, communication, documentation, and stakeholder collaboration skills, with the ability to operate effectively in ambiguous and high-impact environments.
- Willingness and ability to travel up to 10%.
Responsibilities
- Design, build, automate, and operate secure and scalable cloud infrastructure across Azure using Infrastructure as Code technologies such as Terraform and Pulumi.
- Embed governance, compliance, security, data quality, and auditability into infrastructure and platform solutions by design.
- Develop and maintain CI/CD pipelines, self-service capabilities, and resource provisioning frameworks that enable engineering and data teams to access compute, data, and ML resources efficiently.
- Establish and manage platform operations, including monitoring, observability, alerting, incident response, reliability practices, support models, and operational processes.
- Evaluate emerging technologies and cloud services, assess their practical value, and implement modern tooling that improves platform capabilities, scalability, and user experience.
- Partner with platform leadership, engineering teams, and internal stakeholders to translate business and technical requirements into reliable, well-engineered solutions.
- Drive cross-team architecture and design decisions, establish technical standards, and contribute to platform-wide engineering practices.
- Mentor engineers and promote technical excellence, automation, continuous learning, and collaboration across the organization.
- Support infrastructure initiatives for data and AI workloads, including increasingly LLM-driven and agentic applications.
View Full Description & ApplyYou'll be redirected to the employer's site