Site Reliability Engineer

New
C
CloudbedsHospitality software
Location: United States of AmericaFull-TimeSenior
Salary120,000 - 150,000 USD per year
Apply NowOpens the employer's application page

Job Details

Languages
Good written and verbal communication in English.
Experience
5+ years of experience as a DevOps or SRE working within the AWS ecosystem; 5+ years of experience with Kubernetes (EKS) and Helm charts.
Required Skills
AWSGrafanaPrometheusTerraformGitHub ActionsDatadogHelm

Requirements

  • Have 5+ years of experience as a DevOps or SRE working within the AWS ecosystem.
  • Have 5+ years of experience with Kubernetes (EKS) and Helm charts.
  • Have experience designing, building, and supporting CI/CD pipelines with ArgoCD and GitHub Actions.
  • Have experience with infrastructure-as-code methodologies using Terraform.
  • Have experience with observability and monitoring using Grafana, Prometheus, DataDog, and Cloudwatch.
  • Have experience with incident management, full stack troubleshooting, performance analysis, and root cause analysis.
  • Have experience with web application systems such as Nginx, ingress controllers, load balancing, and content delivery networks.
  • Have experience with MySQL, PostgreSQL, or Aurora databases and middleware such as Redis, Memcached, and SQS.
  • Have networking skills with VPC, Security Groups, and Network ACLs.
  • Be able to work remotely and manage your own time in a global team.
  • Have good written and verbal communication in English.
  • Have a Bachelor's degree in Computer Science or equivalent experience.

Responsibilities

  • Design and implement reliable, scalable AWS architecture.
  • Maintain and support highly loaded Kubernetes (EKS) clusters and infrastructure components.
  • Support the CI/CD process with ArgoCD and GitOps.
  • Automate platform deployments with Terraform infrastructure-as-code.
  • Develop and improve product observability and monitoring systems using Grafana, Prometheus, DataDog, and Cloudwatch.
  • Participate in incident management and root cause analysis.
  • Optimize system performance and troubleshoot issues.
  • Collaborate with development teams on monitoring best practices and reliability targets.
  • Collaborate with security teams to implement and maintain security best practices.
  • Provide guidance to other engineering teams during infrastructure support rotation.
View Full Description & ApplyYou'll be redirected to the employer's site
120,000 - 150,000 USD per year
Apply Now