Senior DevOps Platform Engineer
New
S
Salt AIAI life sciences
Listing locations: USA; Fully remote role, U.S. time zonesFull-TimeSenior
Salary140,000 - 180,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience operating production cloud infrastructure (GCP and/or AWS)
- Required Skills
- AWSDockerPythonBashGCPKubernetesCI/CDTerraformCloudFormation
Requirements
- Have 5+ years of experience operating production cloud infrastructure using GCP and/or AWS.
- Have experience with CI/CD systems and pipeline design, such as GitLab CI or GitHub Actions.
- Have production experience with Kubernetes and container orchestration.
- Have hands-on experience with infrastructure-as-code tools such as Terraform or CloudFormation.
- Be proficient with Docker and core containerization concepts.
- Have experience implementing and operating observability stacks covering metrics, logs, and traces.
- Understand web application deployment, scaling, and reliability fundamentals.
- Have scripting skills in Python, Bash, or similar languages for automation.
- Have experience integrating automated tests into deployment workflows.
- Bring problem-solving and debugging skills, especially in distributed systems.
- Database administration, backup, and recovery experience is a plus.
- Familiarity with cloud-native and multi-tenant application security best practices is a plus.
- Experience with performance testing, capacity planning, cost optimization, networking, load balancing, or traffic routing is a plus.
- Experience with AI/ML, HPC, or other data- and compute-intensive workloads is a plus.
Responsibilities
- Build and maintain high-speed, reliable CI/CD pipelines.
- Own deployment consistency across SaaS, VPC, and on-prem environments using GCP, AWS, and Azure.
- Define and enforce infrastructure and deployment standards to prevent environment drift.
- Implement and evolve monitoring, logging, and alerting for platform health, performance, and SLOs.
- Automate infrastructure provisioning and configuration using infrastructure-as-code.
- Design and maintain application security, backup, and disaster recovery procedures.
- Partner with developers to improve application performance, reliability, and operability.
- Participate in on-call rotations, incident response, root cause analysis, and remediation.
- Maintain documentation for infrastructure, deployment processes, and runbooks.
- Enable automated testing in pipelines and stable test environments for QA.
View Full Description & ApplyYou'll be redirected to the employer's site