Senior Software Engineering Manager, Managed Gateways SREs
New
J
JobgetherSecurity & IT
CanadaFull-TimeManager
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- KubernetesGoGrafanaPrometheusDevOps
Requirements
- Proven experience leading Site Reliability Engineering, DevOps, or Infrastructure Engineering teams in fast-paced, high-growth organizations.
- Strong expertise designing, deploying, and operating highly available distributed systems on AWS, Azure, GCP, or similar cloud platforms.
- Extensive hands-on experience with Kubernetes and container orchestration in production environments.
- Proficiency in Go (Golang) or another modern programming language used for infrastructure automation and platform development.
- Solid understanding of observability practices, including experience with tools such as Prometheus, Grafana, OpenTelemetry, or comparable platforms.
- Demonstrated experience managing production incidents, performing root cause analysis, and implementing long-term reliability improvements.
- Familiarity with API gateways, service mesh technologies, or network proxy solutions is considered an asset.
- Strong leadership, communication, collaboration, and stakeholder management skills, with the ability to balance strategic vision and technical execution.
- Cloud or Kubernetes certifications (such as AWS Certified DevOps Engineer, CKA, or CKAD), open-source contributions, or experience in developer infrastructure organizations are advantageous.
Responsibilities
- Build and lead a high-performing Site Reliability Engineering team, establishing engineering standards, best practices, and a culture of operational excellence.
- Contribute directly to critical infrastructure implementations during the team's growth phase, combining leadership with hands-on technical execution.
- Design, deploy, and optimize highly available, scalable cloud-native infrastructure using Kubernetes and modern cloud technologies.
- Drive reliability initiatives by managing monitoring, alerting, incident response, root cause analysis, and continuous service improvements.
- Develop automation, self-service tooling, and operational processes that improve deployment efficiency and reduce manual effort.
- Define, monitor, and improve service level objectives (SLOs), service level indicators (SLIs), and overall platform reliability.
- Collaborate with engineering, product, and support teams to ensure new features are operationally ready and aligned with long-term platform goals.
- Mentor engineers, support career development, and foster a collaborative, high-performance team environment.
View Full Description & ApplyYou'll be redirected to the employer's site