Senior Machine Learning Engineer

New
A
AirArtificial Intelligence
Arlington, Virginia, United States; Pittsburgh, Pennsylvania, United States; RemoteFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
AWSPythonGCPKubernetesAzureDeep LearningDistributed Systems

Requirements

  • 5+ years of experience building production machine learning systems, ML infrastructure, or distributed systems.
  • Deep experience designing, building, and operating production ML infrastructure or ML platforms.
  • Experience building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models.
  • Strong understanding of the modern LLM lifecycle, including data preparation, training, evaluation, model artifacts, deployment, inference, and monitoring.
  • Experience with LLM fine-tuning and post-training workflows (e.g., supervised fine-tuning, LoRA/QLoRA).
  • Experience building reproducible ML pipelines involving dataset versioning, experiment tracking, and model versioning.
  • Experience building and operating production GPU infrastructure across AWS, GCP, or Azure.
  • Strong programming experience in Python and experience building production-quality software.
  • Deep experience with containers, Kubernetes, and cloud platforms.
  • Strong understanding of distributed systems and computationally intensive ML workloads at scale.
  • Experience designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.

Responsibilities

  • Design and build LLMOps infrastructure supporting the development, evaluation, deployment, and continuous improvement of production language models.
  • Build scalable training and fine-tuning infrastructure for commercial and open-weight language models.
  • Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation.
  • Build infrastructure for distributed training and GPU-accelerated ML workloads.
  • Develop experiment management infrastructure that enables engineers to compare models, datasets, hyperparameters, prompts, and training techniques.
  • Build automated evaluation pipelines that determine whether new models or model versions are ready for production deployment.
  • Build and operate scalable model-serving and inference infrastructure for open-weight and fine-tuned models.
  • Build observability for model training and inference, including metrics, tracing, logging, resource utilization, and cost.
  • Optimize training and inference workloads for GPU utilization, throughput, latency, reliability, and infrastructure cost.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now