Senior Software Engineer - Reliability, Infrastructure, and Tooling
New
J
JobgetherSoftware Engineering
CanadaFull-TimeSenior
Salary$135,000–$300,000 USD compensation range
Apply NowOpens the employer's application page
Job Details
- Required Skills
- KafkaKubernetesClickhouseLinuxDistributed SystemsNetworking
Requirements
- Strong professional experience building and operating non-trivial production applications, particularly involving high concurrency or distributed workloads.
- Significant experience with Kubernetes or an equivalent large-scale container orchestration platform.
- Strong understanding of Linux internals and networking with the ability to troubleshoot across system layers.
- Proven experience using observability, monitoring, and logging tooling to diagnose production issues.
- Experience operating large-scale, globally distributed systems.
- Experience responding to and managing complex production incidents.
- Experience operating open-source infrastructure technologies such as Kafka or ClickHouse.
- Strong systems-thinking skills to reason about infrastructure dependencies and control mechanisms.
- Strong communication and collaboration skills for working with partner engineering teams.
- A pragmatic engineering approach balancing delivery with long-term maintainability and operational cost.
Responsibilities
- Ramp up on a complex global architecture involving distributed databases, messaging systems, and networking infrastructure.
- Design and ship reliability-focused engineering work such as load balancing, instrumentation, and scalability improvements.
- Build and evolve internal infrastructure and developer tooling to enable independent operation by product teams.
- Partner with product development teams to co-design systems for security, maintainability, and operational readiness.
- Develop observability capabilities to make system behavior measurable and actionable.
- Participate in a shared on-call rotation and contribute to effective incident response and prevention.
- Investigate complex system-level problems across distributed infrastructure and production environments.
- Improve configuration management practices to reduce technical debt and complexity.
- Automate repetitive operational processes to improve engineering efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site