Site Reliability Engineer
New
A
ArangoDBDatabase Infrastructure
(Remote)- IndiaFull-Time
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- AWSDockerPythonGCPKubernetesGoGrafanaPrometheusCI/CDLinux
Requirements
- SRE or DevOps Engineer background in cloud-native environments.
- Strong proficiency in AWS and GCP platforms.
- Advanced knowledge of Linux internals including processes and environment variables.
- Hands-on experience with containerization and orchestration (Docker, Kubernetes at scale).
- Experience with CI/CD pipelines (Jenkins, CircleCI).
- Experience with monitoring, alerting, and logging tools (Prometheus, Grafana, ELK stack).
- Proficiency in Git version control.
- Programming proficiency in Golang or Python.
- Strong understanding of core networking and security best practices.
- Systematic troubleshooting ability for complex infrastructure issues.
- Self-organized and autonomous remote work style.
Responsibilities
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
- Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
- Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management.
- Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems.
- Develop strategies for disaster recovery, high availability, and fault tolerance.
- Identify system bottlenecks, troubleshoot, and resolve issues across the stack.
- Implement monitoring, logging, and alerting systems to ensure visibility into system health.
- Participate in on-call rotations to support critical production systems and respond to incidents.
View Full Description & ApplyYou'll be redirected to the employer's site