Senior Site Reliability Engineer (Compute Node Team)
New
J
JobgetherCloud Infrastructure
UKFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- KubernetesLinuxDistributed Systems
Requirements
- Deep expertise in Linux systems, including user space, kernel space, and core kernel subsystems such as scheduling, memory management, filesystems, cgroups, and namespaces.
- Strong understanding of system architecture, boundaries, and performance trade-offs across different infrastructure layers.
- Hands-on experience with virtualization technologies, particularly QEMU/KVM, including VM lifecycle management and performance optimization.
- Practical experience with container technologies, namespaces, and resource isolation mechanisms.
- Strong debugging and problem-solving skills with a structured, hypothesis-driven approach to incident investigation.
- Solid understanding of SRE principles, including reliability engineering, system design, and operational ownership.
- Experience building and operating observability solutions rather than only consuming monitoring tools.
- Ability to translate system behavior into actionable reliability improvements.
- Experience with Kubernetes internals, node-level components, or large-scale compute platforms is a plus.
- Familiarity with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash analysis is beneficial.
- Experience with hardware-level debugging, GPU infrastructure, NVLink, InfiniBand, or open-source infrastructure projects is considered an advantage.
Responsibilities
- Ensure the reliability, availability, and performance of compute nodes running virtual machines across cloud environments.
- Analyze and troubleshoot complex Linux systems across both user space and kernel space.
- Investigate and resolve production issues involving CPU, memory, NUMA, cgroups, scheduling, and system performance.
- Work hands-on with virtualization technologies, including QEMU/KVM and Linux-native virtualization solutions.
- Design and improve observability capabilities at the infrastructure layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
- Lead incident response activities, perform root-cause analysis, and drive post-incident improvements.
- Collaborate with platform, kernel, hypervisor, GPU, and infrastructure teams to improve system architecture and operational reliability.
- Develop solutions that enhance scalability, performance, and maintainability of compute platforms.
- Contribute to engineering practices that promote automation, reliability, and continuous improvement.
View Full Description & ApplyYou'll be redirected to the employer's site