Senior Site Reliability Engineer - Data Persistence
New
W
Wikimedia FoundationNon-profit / Technology
Please note that we are currently able to hire in the following: US States: Arizona, California, Colorado, Connecticut, District of Columbia*, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Missouri, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Oregon, Pennsylvania, Puerto Rico*, Rhode Island, Tennessee, Texas, Utah, Vermont, Virginia, Washington, West Virginia, Wisconsin and Wyoming (*US Territory or Federal District). Countries: Brazil, Canada, Colombia, Germany, Ghana, India, Indonesia, Italy, Kenya*, Mexico, Morocco, Netherlands, Poland, Singapore*, South Africa, Spain, Switzerland and the United Kingdom., Ability to work across multiple time zonesFull-TimeSenior
Salary116,633 - 181,243 USD per year
Apply NowOpens the employer's application page
Job Details
- Languages
- Strong English language skills
- Experience
- 6+ years
- Required Skills
- PythonKubernetesLinuxDevOpsAnsibleDistributed Systems
Requirements
- 6+ years of experience in an SRE, Operations, or DevOps role.
- Proficiency in shell and at least one scripting language (Python, Go, Bash, Ruby).
- Experience with configuration management tools (Puppet, Ansible).
- Deep experience with distributed caching systems and performance optimization.
- Strong Linux system-level troubleshooting and package management skills (Debian).
- History of identifying process gaps and implementing automation.
- Ability to lead and participate in incident response and root cause analysis.
- Strong English verbal and written communication skills.
- Ability to travel 1-2 times per year for in-person events.
Responsibilities
- Perform day-to-day operational/DevOps tasks on public-facing infrastructure, including deployment, maintenance, configuration, and troubleshooting.
- Implement and utilize configuration management and deployment tools such as Puppet and Kubernetes.
- Lead continuous improvement efforts by automating service installation, configuration, and maintenance.
- Collaborate with product teams on architectural design of new services to ensure scalability.
- Participate in a 24/7 on-call rotation involving incident response, diagnosis, and system outage follow-up.
- Mentor peers in areas of technical and operational strength.
View Full Description & ApplyYou'll be redirected to the employer's site