Senior Machine Learning Systems Engineer
New
J
JobgetherMachine Learning
Remote work from within the United States.Full-TimeSenior
SalaryBase salary range of $216,700–$303,400 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- PythonGCPKubernetesPyTorchSparkTensorflowTerraformMLOps
Requirements
- 5+ years of professional experience working with machine learning infrastructure.
- Hands-on experience optimizing ML workloads including GPU profiling and resource efficiency.
- Deep experience with cloud technologies, particularly GCP BigQuery and Google Cloud Storage.
- Practical experience with infrastructure-as-code tools such as Terraform.
- Experience administering MLOps technologies for experiment tracking and model registries such as MLflow or Weights & Biases.
- Strong programming skills in Python and frameworks including PyTorch and/or TensorFlow.
- Deep experience with distributed training and compute frameworks, including Ray and Kubernetes.
- Strong understanding of scalable data processing and distributed systems.
- Experience with graph databases like Neo4j, JanusGraph, or TigerGraph is a plus.
- Experience with graph neural networks and frameworks like PyTorch Geometric or Deep Graph Library is a plus.
Responsibilities
- Design and implement end-to-end model lifecycle and MLOps patterns covering data preparation, model management, experiment tracking, and deployment.
- Develop and support graph machine learning platforms and codebases to enable scalable model development.
- Partner with machine learning engineers to improve training performance, reduce training times, and optimize GPU costs.
- Optimize large-scale batch data processing using cloud data warehouses and technologies like Apache Beam, Apache Spark, and Ray Data.
- Architect pipelines for creating and maintaining massive graph datasets containing billions of nodes and edges.
- Administer and integrate tooling for experiment tracking, model serving, and model registries.
- Build infrastructure prioritizing scalability, reliability, and ease of use.
- Collaborate with platform users to remove technical friction from ML workflows.
View Full Description & ApplyYou'll be redirected to the employer's site