US Infrastructure & Operations Technical Lead
New
R
RadiantInfrastructure Operations
Location: US East Coast. A flexible remote-first working environment, US Eastern TimeFull-TimeLead
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years infrastructure engineering, SRE, platform ops, or large-scale production infrastructure experience
- Required Skills
- PythonBashLinuxTerraformAnsibleNetworking
Requirements
- 8+ years infrastructure engineering, SRE, platform ops, or large-scale production infrastructure experience
- 2+ years in technical leadership or engineering management with direct reports
- Strong experience operating production infrastructure at scale (on-prem, private cloud, or hybrid)
- Deep Linux expertise: performance tuning, debugging, kernel/system behaviour, production troubleshooting
- Hands-on across the stack: bare metal, OS, network, storage, and platform services
- Strong infrastructure fundamentals: compute, storage, networking in real-world production environments
- Incident-heavy environment experience (24x7 ops, on-call, major incident response, postmortems)
- Strong networking skills including TCP/IP, routing, switching, DNS, latency, and packet-level debugging
- Bare-metal operations experience (Redfish, IPMI, lifecycle management, hardware troubleshooting)
- Strong automation and configuration management skills (Ansible preferred)
- Strong scripting skills (Python, Bash or similar)
- Willingness to travel within the US and Europe as required
Responsibilities
- Lead a small but high-impact US Infrastructure Operations team, owning both people leadership and technical execution
- Ensure 99.9%+ platform uptime across US-region services
- Act as the senior US operational owner for production infrastructure, accountable for reliability, incident outcomes, and day-to-day operational execution
- Own US-side incident leadership, driving fast and effective resolution of production-impacting infrastructure issues
- Participate in on-call rotation and lead from the front during major incidents
- Build and improve Infrastructure as Code workflows (Terraform, Ansible or equivalent)
- Contribute directly to scaling decisions, capacity planning, and reliability improvements
View Full Description & ApplyYou'll be redirected to the employer's site