US Infrastructure & Operations Technical Lead

New
R
RadiantInfrastructure Operations
Location: US East Coast. A flexible remote-first working environment, US Eastern TimeFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
8+ years infrastructure engineering, SRE, platform ops, or large-scale production infrastructure experience
Required Skills
PythonBashLinuxTerraformAnsibleNetworking

Requirements

  • 8+ years infrastructure engineering, SRE, platform ops, or large-scale production infrastructure experience
  • 2+ years in technical leadership or engineering management with direct reports
  • Strong experience operating production infrastructure at scale (on-prem, private cloud, or hybrid)
  • Deep Linux expertise: performance tuning, debugging, kernel/system behaviour, production troubleshooting
  • Hands-on across the stack: bare metal, OS, network, storage, and platform services
  • Strong infrastructure fundamentals: compute, storage, networking in real-world production environments
  • Incident-heavy environment experience (24x7 ops, on-call, major incident response, postmortems)
  • Strong networking skills including TCP/IP, routing, switching, DNS, latency, and packet-level debugging
  • Bare-metal operations experience (Redfish, IPMI, lifecycle management, hardware troubleshooting)
  • Strong automation and configuration management skills (Ansible preferred)
  • Strong scripting skills (Python, Bash or similar)
  • Willingness to travel within the US and Europe as required

Responsibilities

  • Lead a small but high-impact US Infrastructure Operations team, owning both people leadership and technical execution
  • Ensure 99.9%+ platform uptime across US-region services
  • Act as the senior US operational owner for production infrastructure, accountable for reliability, incident outcomes, and day-to-day operational execution
  • Own US-side incident leadership, driving fast and effective resolution of production-impacting infrastructure issues
  • Participate in on-call rotation and lead from the front during major incidents
  • Build and improve Infrastructure as Code workflows (Terraform, Ansible or equivalent)
  • Contribute directly to scaling decisions, capacity planning, and reliability improvements
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now