Site Reliability Engineer (SRE) - Engineering Productivity
New
J
JobgetherSoftware Engineering
IndiaFull-TimeMiddle
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PostgreSQLPythonGoLinux
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent professional experience.
- Approximately 5+ years of relevant experience.
- Working knowledge of Go, Python, and/or shell scripting.
- Strong Linux or UNIX administration and debugging capabilities.
- Hands-on experience operating software systems or complex infrastructure at scale.
- Experience with server provisioning, storage, and networking.
- Practical experience with infrastructure-as-code and automated management.
- Strong software troubleshooting and analytical skills.
- Experience with databases (e.g., MariaDB, PostgreSQL, MongoDB) is desirable.
- Experience with Docker, virtualization, or container technologies is a plus.
- Experience with monitoring/observability tools like Prometheus, Loki, or Grafana is beneficial.
Responsibilities
- Design, build, deploy, and operate critical production systems with a focus on scalability, reliability, and security.
- Develop automation to reduce operational toil and improve engineering workflows.
- Monitor infrastructure proactively, improve alerting, and implement automated incident responses.
- Create and maintain incident response procedures and operational runbooks.
- Build and deploy new systems using staged rollouts to minimize risk.
- Investigate and resolve infrastructure issues while supporting software engineering teams.
- Collaborate with vendors to diagnose platform-related issues.
- Write post-incident reviews and implement corrective measures.
- Partner with product development teams to identify and resolve infrastructure bottlenecks.
View Full Description & ApplyYou'll be redirected to the employer's site