Site Reliability Engineer

New
A
ArangoDBDatabase Infrastructure
(Remote)- IndiaFull-Time
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
AWSDockerPythonGCPKubernetesGoGrafanaPrometheusCI/CDLinux

Requirements

  • SRE or DevOps Engineer background in cloud-native environments.
  • Strong proficiency in AWS and GCP platforms.
  • Advanced knowledge of Linux internals including processes and environment variables.
  • Hands-on experience with containerization and orchestration (Docker, Kubernetes at scale).
  • Experience with CI/CD pipelines (Jenkins, CircleCI).
  • Experience with monitoring, alerting, and logging tools (Prometheus, Grafana, ELK stack).
  • Proficiency in Git version control.
  • Programming proficiency in Golang or Python.
  • Strong understanding of core networking and security best practices.
  • Systematic troubleshooting ability for complex infrastructure issues.
  • Self-organized and autonomous remote work style.

Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management.
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems.
  • Develop strategies for disaster recovery, high availability, and fault tolerance.
  • Identify system bottlenecks, troubleshoot, and resolve issues across the stack.
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health.
  • Participate in on-call rotations to support critical production systems and respond to incidents.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now