Staff Network Engineer

R
Radian ArcAI Infrastructure
Based in Germany; 100% remote work within EuropeContractStaff
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English
Required Skills
PythonBashLinuxNetworking

Requirements

  • Extensive hands-on experience designing and operating large-scale datacenter networks in production.
  • Expert knowledge of routing protocols including BGP, OSPF, and ECMP.
  • Deep expertise in AI networking fabrics and practical RoCE/RDMA implementation.
  • Strong understanding of NCCL communication patterns and their impact on network topology.
  • Experience tuning large-scale RoCE fabrics, including PFC and ECN congestion management.
  • Proven experience with high-speed Ethernet and NVIDIA/Mellanox networking hardware.
  • Proficiency in troubleshooting across hardware, firmware, and Linux kernel networking.
  • Strong Python and/or Bash skills for infrastructure automation and tool development.
  • Experience designing network observability systems and analyzing operational data.
  • Ability to own long-term architecture while remaining hands-on with execution.
  • Demonstrated ability to lead technical initiatives and influence teams without formal authority.
  • Professional working proficiency in English.

Responsibilities

  • Design, deploy, and operate high-performance GPU networking fabrics using RoCE, RDMA, Spectrum-X, leaf-spine, and fat-tree architectures.
  • Optimize east-west traffic, congestion control, and latency for distributed AI training and inference workloads.
  • Own Layer 2/3 architecture including BGP, ECMP, EVPN/VXLAN, and Linux-based networking.
  • Design and operate secure edge connectivity, including WAF, TLS termination, and DDoS mitigation.
  • Design inter-datacenter private connectivity, including dark-fiber rings, WAN links, and DWDM transport.
  • Lead infrastructure initiatives from architecture and lab validation through to production rollout.
  • Establish automation for provisioning, configuration management, and lifecycle management.
  • Act as the senior escalation point for complex network incidents and root-cause analysis.
  • Define and monitor SLAs and SLOs for network performance and reliability.
  • Partner with compute, storage, SRE, and platform teams to set organizational networking standards.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now