AI Platform Engineer
New
B
Bright Vision TechnologiesAI Platform Engineering
100% Remote (Continental United States)Full-TimeSenior
Salary130,000 - 180,000 USD per year
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ Years
- Required Skills
- AWSDockerPythonGCPKubernetesC++AzureGoRustLLM
Requirements
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
- Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
- Extensive experience with Large Language Model (LLM) serving, model inference optimization, and production AI infrastructure.
- Hands-on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks.
- Strong expertise in Kubernetes, container orchestration, Docker, and cloud-native application architectures.
- Experience optimizing GPU workloads using CUDA, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure.
- Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking, scalability, observability, and security best practices.
Responsibilities
- Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
- Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
- Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
- Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
- Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
- Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
- Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
- Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to deploy and support production AI models.
View Full Description & ApplyYou'll be redirected to the employer's site