Machine Learning Infrastructure Engineer
B
Bright Vision TechnologiesTechnology Consulting
100% Remote (U.S.)Full-TimeSenior
Salary105,000 - 143,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 6+ years
- Required Skills
- PythonKubernetesC++GoRustDistributed Systems
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Deep experience operating high-throughput, low-latency services in production.
- Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
- Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
- Experience with observability stacks including metrics, tracing, and structured logging.
- Solid grounding in performance engineering and capacity planning.
- Strong communication and incident response skills.
Responsibilities
- Design and build reliable inference platforms for serving large machine learning models in production.
- Operate serving systems with a focus on request routing, batching, and caching.
- Optimize GPU utilization and implement autoscaling mechanisms for diverse model workloads.
- Develop and maintain end-to-end observability across AI production environments.
- Conduct performance engineering and capacity planning for model deployment.
View Full Description & ApplyYou'll be redirected to the employer's site