Senior Site Reliability Engineer - Infrastructure Foundations
New
W
Wikimedia FoundationTechnology, Non-profit
US States: Arizona, California, Colorado, Connecticut, District of Columbia*, Florida, Georgia, Idaho, Illinois, Indiana, Iowa, Maryland, Massachusetts, Michigan, Minnesota, Missouri, New Jersey, New Mexico, New York, North Carolina, Ohio, Oklahoma, Oregon, Pennsylvania, Puerto Rico*, Rhode Island, Tennessee, Texas, Utah, Vermont, Virginia, Washington, West Virginia, Wisconsin and Wyoming. Countries: Brazil, Canada, Colombia, Germany, Ghana, India, Indonesia, Italy, Kenya*, Mexico, Morocco, Netherlands, Poland, Singapore*, South Africa, Spain, Switzerland and the United Kingdom., Working across multiple time zonesFull-TimeSenior
Salary113,082 - 175,725 USD per year
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Experience
- 6+ years
- Required Skills
- PythonKubernetesLinuxDevOpsAnsible
Requirements
- 6+ years of experience in an SRE/Operations/DevOps role.
- Experience with shell and scripting languages (specifically Python, Go, Bash, Ruby).
- Proficiency with configuration management tools such as Puppet or Ansible.
- Strong Linux system-level troubleshooting skills.
- Experience with package management on Linux (specifically Debian).
- Experience designing and managing infrastructure security for large fleets.
- History of automating tasks, processes, and identifying automation opportunities.
- Experience leading and participating in incident response and post-incident reviews.
- Strong English language skills (verbal and written).
- Ability and willingness to travel 1-2 times a year for in-person events.
Responsibilities
- Perform day-to-day operational/DevOps tasks on public-facing infrastructure including deployment, maintenance, configuration, and troubleshooting.
- Implement and utilize configuration management and deployment tools such as Puppet and Kubernetes.
- Lead continuous improvement by automating the installation, configuration, and maintenance of services.
- Work closely with product teams to assist in architectural design of new services and ensure scalability.
- Participate in a 24/7 on-call rotation involving incident response, diagnosis, and follow-up on system outages or alerts.
- Collaborate with a global, cross-functional team in an asynchronous communication environment.
- Mentor peers in areas of technical and operational strength.
View Full Description & ApplyYou'll be redirected to the employer's site