Senior Site Reliability Engineer - Data Persistence
New
W
Wikimedia FoundationNon-profit Technology
US States: Arizona, California, Colorado, Connecticut, District of Columbia*, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Missouri, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Oregon, Pennsylvania, Puerto Rico*, Rhode Island, Tennessee, Texas, Utah, Vermont, Virginia, Washington, West Virginia, Wisconsin and Wyoming (*US Territory or Federal District). Countries: Brazil, Canada, Colombia, Germany, Ghana, India, Indonesia, Italy, Kenya*, Mexico, Morocco, Netherlands, Poland, Singapore*, South Africa, Spain, Switzerland and the United Kingdom.Full-TimeSenior
Salary116,633 - 181,243 USD per year
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 6+ years
- Required Skills
- PythonKubernetesLinuxAnsibleDistributed Systems
Requirements
- 6+ years experience in an SRE/Operations/DevOps role as part of a team
- Experience with shell and scripting languages (Python, Go, Bash, Ruby)
- Experience with configuration management tools (Puppet, Ansible)
- Experience with distributed caching systems and their underlying algorithms
- Experience with package management on Linux systems (Debian)
- Strong Linux system-level troubleshooting skills
- Proven history of automating tasks and processes
- Strong English language skills (verbal and written)
- Experience leading and participating in incident response and post-incident reviews
Responsibilities
- Perform day-to-day operational/DevOps tasks on Wikimedia’s public facing infrastructure (deployment, maintenance, configuration, troubleshooting)
- Implement and utilize configuration management and deployment tools (Puppet, Kubernetes)
- Lead continuous improvement, by automating the installation, configuration and maintenance of services on our platform
- Work closely with product teams to design scalable service architectures
- Participate in a 24/7 on-call rotation including incident response and root cause analysis
- Collaborate with a global, cross-functional team in an asynchronous environment
- Mentor peers in technical and operational areas
View Full Description & ApplyYou'll be redirected to the employer's site