Senior Software Engineer - Reliability, Infrastructure, and Tooling

New
J
JobgetherSoftware Engineering
CanadaFull-TimeSenior
Salary$135,000–$300,000 USD compensation range
Apply NowOpens the employer's application page

Job Details

Required Skills
KafkaKubernetesClickhouseLinuxDistributed SystemsNetworking

Requirements

  • Strong professional experience building and operating non-trivial production applications, particularly involving high concurrency or distributed workloads.
  • Significant experience with Kubernetes or an equivalent large-scale container orchestration platform.
  • Strong understanding of Linux internals and networking with the ability to troubleshoot across system layers.
  • Proven experience using observability, monitoring, and logging tooling to diagnose production issues.
  • Experience operating large-scale, globally distributed systems.
  • Experience responding to and managing complex production incidents.
  • Experience operating open-source infrastructure technologies such as Kafka or ClickHouse.
  • Strong systems-thinking skills to reason about infrastructure dependencies and control mechanisms.
  • Strong communication and collaboration skills for working with partner engineering teams.
  • A pragmatic engineering approach balancing delivery with long-term maintainability and operational cost.

Responsibilities

  • Ramp up on a complex global architecture involving distributed databases, messaging systems, and networking infrastructure.
  • Design and ship reliability-focused engineering work such as load balancing, instrumentation, and scalability improvements.
  • Build and evolve internal infrastructure and developer tooling to enable independent operation by product teams.
  • Partner with product development teams to co-design systems for security, maintainability, and operational readiness.
  • Develop observability capabilities to make system behavior measurable and actionable.
  • Participate in a shared on-call rotation and contribute to effective incident response and prevention.
  • Investigate complex system-level problems across distributed infrastructure and production environments.
  • Improve configuration management practices to reduce technical debt and complexity.
  • Automate repetitive operational processes to improve engineering efficiency.
View Full Description & ApplyYou'll be redirected to the employer's site
$135,000–$300,000 USD compensation range
Apply Now