Senior DevOps Platform Engineer

New
S
Salt AIAI life sciences
Listing locations: USA; Fully remote role, U.S. time zonesFull-TimeSenior
Salary140,000 - 180,000 USD per year
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience operating production cloud infrastructure (GCP and/or AWS)
Required Skills
AWSDockerPythonBashGCPKubernetesCI/CDTerraformCloudFormation

Requirements

  • Have 5+ years of experience operating production cloud infrastructure using GCP and/or AWS.
  • Have experience with CI/CD systems and pipeline design, such as GitLab CI or GitHub Actions.
  • Have production experience with Kubernetes and container orchestration.
  • Have hands-on experience with infrastructure-as-code tools such as Terraform or CloudFormation.
  • Be proficient with Docker and core containerization concepts.
  • Have experience implementing and operating observability stacks covering metrics, logs, and traces.
  • Understand web application deployment, scaling, and reliability fundamentals.
  • Have scripting skills in Python, Bash, or similar languages for automation.
  • Have experience integrating automated tests into deployment workflows.
  • Bring problem-solving and debugging skills, especially in distributed systems.
  • Database administration, backup, and recovery experience is a plus.
  • Familiarity with cloud-native and multi-tenant application security best practices is a plus.
  • Experience with performance testing, capacity planning, cost optimization, networking, load balancing, or traffic routing is a plus.
  • Experience with AI/ML, HPC, or other data- and compute-intensive workloads is a plus.

Responsibilities

  • Build and maintain high-speed, reliable CI/CD pipelines.
  • Own deployment consistency across SaaS, VPC, and on-prem environments using GCP, AWS, and Azure.
  • Define and enforce infrastructure and deployment standards to prevent environment drift.
  • Implement and evolve monitoring, logging, and alerting for platform health, performance, and SLOs.
  • Automate infrastructure provisioning and configuration using infrastructure-as-code.
  • Design and maintain application security, backup, and disaster recovery procedures.
  • Partner with developers to improve application performance, reliability, and operability.
  • Participate in on-call rotations, incident response, root cause analysis, and remediation.
  • Maintain documentation for infrastructure, deployment processes, and runbooks.
  • Enable automated testing in pipelines and stable test environments for QA.
View Full Description & ApplyYou'll be redirected to the employer's site
140,000 - 180,000 USD per year
Apply Now