Staff Network Engineer
R
Radian ArcAI Infrastructure
Based in Germany; 100% remote work within EuropeContractStaff
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English
- Required Skills
- PythonBashLinuxNetworking
Requirements
- Extensive hands-on experience designing and operating large-scale datacenter networks in production.
- Expert knowledge of routing protocols including BGP, OSPF, and ECMP.
- Deep expertise in AI networking fabrics and practical RoCE/RDMA implementation.
- Strong understanding of NCCL communication patterns and their impact on network topology.
- Experience tuning large-scale RoCE fabrics, including PFC and ECN congestion management.
- Proven experience with high-speed Ethernet and NVIDIA/Mellanox networking hardware.
- Proficiency in troubleshooting across hardware, firmware, and Linux kernel networking.
- Strong Python and/or Bash skills for infrastructure automation and tool development.
- Experience designing network observability systems and analyzing operational data.
- Ability to own long-term architecture while remaining hands-on with execution.
- Demonstrated ability to lead technical initiatives and influence teams without formal authority.
- Professional working proficiency in English.
Responsibilities
- Design, deploy, and operate high-performance GPU networking fabrics using RoCE, RDMA, Spectrum-X, leaf-spine, and fat-tree architectures.
- Optimize east-west traffic, congestion control, and latency for distributed AI training and inference workloads.
- Own Layer 2/3 architecture including BGP, ECMP, EVPN/VXLAN, and Linux-based networking.
- Design and operate secure edge connectivity, including WAF, TLS termination, and DDoS mitigation.
- Design inter-datacenter private connectivity, including dark-fiber rings, WAN links, and DWDM transport.
- Lead infrastructure initiatives from architecture and lab validation through to production rollout.
- Establish automation for provisioning, configuration management, and lifecycle management.
- Act as the senior escalation point for complex network incidents and root-cause analysis.
- Define and monitor SLAs and SLOs for network performance and reliability.
- Partner with compute, storage, SRE, and platform teams to set organizational networking standards.
View Full Description & ApplyYou'll be redirected to the employer's site