Senior Machine Learning Engineer
New
A
AirArtificial Intelligence
Arlington, Virginia, United States; Pittsburgh, Pennsylvania, United States; RemoteFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- AWSPythonGCPKubernetesAzureDeep LearningDistributed Systems
Requirements
- 5+ years of experience building production machine learning systems, ML infrastructure, or distributed systems.
- Deep experience designing, building, and operating production ML infrastructure or ML platforms.
- Experience building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models.
- Strong understanding of the modern LLM lifecycle, including data preparation, training, evaluation, model artifacts, deployment, inference, and monitoring.
- Experience with LLM fine-tuning and post-training workflows (e.g., supervised fine-tuning, LoRA/QLoRA).
- Experience building reproducible ML pipelines involving dataset versioning, experiment tracking, and model versioning.
- Experience building and operating production GPU infrastructure across AWS, GCP, or Azure.
- Strong programming experience in Python and experience building production-quality software.
- Deep experience with containers, Kubernetes, and cloud platforms.
- Strong understanding of distributed systems and computationally intensive ML workloads at scale.
- Experience designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.
Responsibilities
- Design and build LLMOps infrastructure supporting the development, evaluation, deployment, and continuous improvement of production language models.
- Build scalable training and fine-tuning infrastructure for commercial and open-weight language models.
- Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation.
- Build infrastructure for distributed training and GPU-accelerated ML workloads.
- Develop experiment management infrastructure that enables engineers to compare models, datasets, hyperparameters, prompts, and training techniques.
- Build automated evaluation pipelines that determine whether new models or model versions are ready for production deployment.
- Build and operate scalable model-serving and inference infrastructure for open-weight and fine-tuned models.
- Build observability for model training and inference, including metrics, tracing, logging, resource utilization, and cost.
- Optimize training and inference workloads for GPU utilization, throughput, latency, reliability, and infrastructure cost.
View Full Description & ApplyYou'll be redirected to the employer's site