Lead Platform & Infrastructure Engineer
J
JobgetherSoftware & IT
Based in the United StatesFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- DockerKubernetesCI/CDLinuxTerraformAnsibleHelm
Requirements
- 8+ years of experience building, deploying, and operating enterprise infrastructure and platform solutions.
- Leadership experience in Platform Engineering, Infrastructure Engineering, DevOps, or Site Reliability Engineering.
- Deep hands-on expertise with Docker, Kubernetes, K3s, Helm, Linux, and networking.
- Strong experience with Infrastructure-as-Code and automation tools like Terraform, Ansible, and Packer.
- Proven experience designing and operating secure, highly available infrastructure, including secrets management, observability, and disaster recovery.
- Experience with CI/CD platforms such as GitHub Actions, GitLab CI, or Gitea Actions.
- Strong understanding of infrastructure security, access controls, and network architecture.
- Demonstrated ability to troubleshoot complex production environments.
- Strong technical leadership and cross-functional collaboration skills.
- Experience supporting regulated industries, customer-managed environments, or AI/ML platform operations is preferred.
Responsibilities
- Design, implement, and maintain scalable platform infrastructure supporting cloud, single-tenant, and customer-managed deployment environments.
- Lead infrastructure architecture and operational practices for secure, highly available production platforms.
- Build and manage Infrastructure-as-Code solutions, automated provisioning, configuration management, deployment automation, and environment lifecycle processes.
- Develop and operate containerized platforms using Docker, Kubernetes, K3s, and Helm.
- Establish and maintain CI/CD pipelines, release automation, and modern DevOps practices.
- Partner with security teams to implement encryption, network segmentation, and infrastructure-hardening standards.
- Drive platform observability through monitoring, logging, alerting, and capacity planning.
- Support production deployments, incident response, and on-call activities to maintain service reliability.
- Contribute to disaster recovery and infrastructure continuity strategies.
- Collaborate with engineering, security, data, and AI/ML teams to support evolving platform requirements.
View Full Description & ApplyYou'll be redirected to the employer's site