Sr. Site Reliability Engineer

New
I
IO Connect ServicesCloud Services
Remote, MexicoFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSPythonGCPJavaKubernetesMicrosoft AzureCI/CDTerraformAnsible

Requirements

  • Bachelor’s degree in computer science or related discipline.
  • Knowledge of IaC technologies such as Terraform, Ansible, Puppet, or Chef.
  • Experience with cluster creation and management through Kubernetes.
  • Experience with cloud services (AWS, Azure, Google Cloud), including VM and Virtual Network configuration.
  • Understanding of design patterns: IaaS, PaaS, and SaaS.
  • Experience with CI/CD pipelines.
  • Scripting proficiency with PowerShell.
  • Understanding of IP addressing and network masking.
  • Programming ability in high-level languages like Python, Java, C/C++, Ruby, or JavaScript.
  • Experience with distributed storage (NFS, HDFS, Ceph, Amazon S3) and resource management frameworks (Mesos, Kubernetes, Yarn).

Responsibilities

  • Design, build, maintain, and scale production services and server farms across multiple data centers.
  • Design and enhance software architecture to improve scalability, reliability, capacity, and performance.
  • Write automation code for provisioning and operating infrastructure.
  • Collaborate with development and QA teams to implement infrastructure and deployment pipelines.
  • Troubleshoot incidents, identify root causes, and write postmortem reviews.
  • Monitor systems to identify trends and respond to automated system alerts.
  • Author and update documentation of specifications, systems, and procedures.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now