Senior Site Reliability / Gitops Engineer
New
J
JobgetherInformation Technology
CanadaFull-TimeSenior
SalaryCompensation shaped according to geographic location, experience, and performance, with regular compensation reviews.
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Required Skills
- PythonCloud ComputingElasticSearchKubernetesGrafanaPrometheusLinux
Requirements
- Strong understanding of modern hosting architectures and an automation-first approach based on Infrastructure as Code.
- A product-oriented mindset with an interest in building reusable infrastructure products.
- Professional experience with Python development, including work on large or complex projects.
- Hands-on experience with Kubernetes or other container orchestration technologies.
- Proven experience managing and deploying cloud infrastructure through code and automation.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from distributed storage such as Ceph to database-backed systems.
- Hands-on experience administering enterprise Linux servers.
- Strong knowledge of cloud computing concepts, architectures, and technologies.
- Bachelor’s degree or higher, preferably in Computer Science, Engineering, or a related technical discipline.
- Strong English communication skills across email, chat, video calls, voice communication, and in-person collaboration.
- Experience working effectively within globally distributed teams.
Responsibilities
- Drive the development of automation and GitOps practices within the team while acting as an embedded technical lead.
- Collaborate closely with the infrastructure architecture function to align technical solutions with broader architecture objectives.
- Design and architect infrastructure services that can be delivered as reusable products across the organization.
- Develop and strengthen Infrastructure as Code practices by continuously improving automation, processes, consistency, and reusability.
- Automate software operations across private and public clouds while accounting for the complexity and operational requirements of distributed systems.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliability and continuity.
- Troubleshoot complex infrastructure issues, support capacity planning, investigate performance challenges, and improve system resilience.
- Implement and maintain observability, monitoring, and alerting solutions using technologies such as Prometheus, Grafana, and Elasticsearch.
- Collaborate with globally distributed engineering, operations, and support teams to resolve technical challenges and deliver reliable services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive or manual operational tasks.
View Full Description & ApplyYou'll be redirected to the employer's site