Lead Engineer - Agentic AI Systems
New
J
JobgetherAI infrastructure
US; based in United StatesFull-TimeLead
SalaryBase salary ranging from $122,000 to $214,000 per year; Target annual bonus of 20% of annualized base salary
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of experience in systems engineering, cloud architecture, and infrastructure leadership.
- Required Skills
- AWSPythonGCPKubernetesAzure
Requirements
- Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
- 5+ years of experience in systems engineering, cloud architecture, and infrastructure leadership.
- Experience architecting multi-region cloud infrastructures, hybrid edge-cloud networks, and large-scale Kubernetes or GKE deployment fleets.
- Track record leading complex engineering initiatives, establishing technical vision, and influencing technology investments and vendor decisions.
- Knowledge of GPU, NPU, and TPU architectures, low-latency networking, distributed storage, and inference acceleration runtimes.
- Proficiency with Python, Kubernetes, and containerization technologies.
- Experience with major cloud platforms such as GCP, Azure, or AWS.
- Understanding of scalable AI infrastructure, model serving, MLOps/LLMOps, and operational requirements of AI and agentic systems.
- Architectural, analytical, and problem-solving abilities to balance performance, reliability, security, scalability, and cost.
- Leadership and communication skills, including technical direction, engineer mentoring, and collaboration across infrastructure, security, IT, AI, and vendor teams.
- A Master’s degree in Computer Science or a related discipline is preferred.
- Experience in the food service or a related industry is preferred.
Responsibilities
- Define technical strategy, hardware roadmaps, and deployment architectures for distributed AI compute, cloud platforms, and store-level edge environments.
- Establish standards for infrastructure reliability, disaster recovery, hardware security, and compute cost governance.
- Lead architectural investigations, capacity planning, and benchmarking for AI workloads, multimodal models, and low-latency edge inference.
- Design scalable model-serving, MLOps, and LLMOps infrastructure for enterprise AI, machine learning, and agentic solutions.
- Partner with Enterprise IT, Security, stakeholders, and hardware vendors on network topologies, edge hardware specifications, secure API gateways, and deployment architectures.
- Evaluate cloud providers, GPU vendors, and edge hardware manufacturers, and support vendor selection and technical requirements and service-level agreements.
- Drive infrastructure right-sizing, security improvements, performance optimization, and cost efficiency.
- Provide technical oversight and architectural review across infrastructure and AI engineering initiatives.
- Establish frameworks, engineering best practices, and technology standards for AI platforms.
- Lead and mentor cross-functional engineering teams and guide delivery across concurrent strategic initiatives.
View Full Description & ApplyYou'll be redirected to the employer's site