AI Systems Performance Specialist
B
Bright Vision TechnologiesAI infrastructure
100% Remote (Continental United States)Full-TimeSenior
Salary130,000 - 180,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ Years
- Required Skills
- PythonCloud ComputingPyTorchC++Deep LearningLLMDistributed Systems
Requirements
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing.
- Expert-level programming skills in Python and C++.
- Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries.
- Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference.
- Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools.
- Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture.
- Excellent analytical, troubleshooting, communication, and technical leadership skills.
Responsibilities
- Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency.
- Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads.
- Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
- Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage.
- Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks.
- Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability.
- Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines.
- Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities.
- Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices.
- Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering.
View Full Description & ApplyYou'll be redirected to the employer's site