Senior Site Reliability Engineer (Compute Node Team)

New
J
JobgetherCloud Infrastructure
UKFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
KubernetesLinuxDistributed Systems

Requirements

  • Deep expertise in Linux systems, including user space, kernel space, and core kernel subsystems such as scheduling, memory management, filesystems, cgroups, and namespaces.
  • Strong understanding of system architecture, boundaries, and performance trade-offs across different infrastructure layers.
  • Hands-on experience with virtualization technologies, particularly QEMU/KVM, including VM lifecycle management and performance optimization.
  • Practical experience with container technologies, namespaces, and resource isolation mechanisms.
  • Strong debugging and problem-solving skills with a structured, hypothesis-driven approach to incident investigation.
  • Solid understanding of SRE principles, including reliability engineering, system design, and operational ownership.
  • Experience building and operating observability solutions rather than only consuming monitoring tools.
  • Ability to translate system behavior into actionable reliability improvements.
  • Experience with Kubernetes internals, node-level components, or large-scale compute platforms is a plus.
  • Familiarity with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash analysis is beneficial.
  • Experience with hardware-level debugging, GPU infrastructure, NVLink, InfiniBand, or open-source infrastructure projects is considered an advantage.

Responsibilities

  • Ensure the reliability, availability, and performance of compute nodes running virtual machines across cloud environments.
  • Analyze and troubleshoot complex Linux systems across both user space and kernel space.
  • Investigate and resolve production issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance.
  • Work hands-on with virtualization technologies, including QEMU/KVM and Linux-native virtualization solutions.
  • Design and improve observability capabilities at the infrastructure layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
  • Lead incident response activities, perform root-cause analysis, and drive post-incident improvements.
  • Collaborate with platform, kernel, hypervisor, GPU, and infrastructure teams to improve system architecture and operational reliability.
  • Develop solutions that enhance scalability, performance, and maintainability of compute platforms.
  • Contribute to engineering practices that promote automation, reliability, and continuous improvement.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now