Staff Machine Learning Systems Engineer
New
J
JobgetherMachine Learning
Based in the United StatesFull-TimeStaff
SalaryBase salary range of $230,000–$322,000 USD
Apply NowOpens the employer's application page
Job Details
- Experience
- 8+ years
- Required Skills
- PythonGCPKubernetesPyTorchSparkTensorflowTerraformMLOps
Requirements
- 8+ years of experience in machine learning infrastructure, including training and deployment environments.
- Hands-on experience optimizing machine learning systems (memory profiling, GPU profiling, performance tuning).
- Deep experience with cloud technologies, specifically Google Cloud Platform (BigQuery, Google Cloud Storage).
- Experience with infrastructure-as-code tools such as Terraform.
- Experience administering and integrating MLOps platforms like MLflow or Weights & Biases.
- Strong proficiency in Python.
- Familiarity with machine learning frameworks such as PyTorch and TensorFlow.
- Deep experience with distributed machine learning and data processing frameworks, specifically Ray and Kubernetes.
- Strong understanding of high-performance platform architecture and the ML development lifecycle.
- Experience with graph databases (e.g., Neo4j, JanusGraph, TigerGraph) is highly valued.
- Experience with graph neural networks and frameworks (e.g., PyTorch Geometric, Deep Graph Library) is highly valued.
Responsibilities
- Design and implement end-to-end model lifecycle patterns and MLOps capabilities covering data preparation, model management, experiment tracking, and deployment.
- Lead zero-to-one development of graph machine learning infrastructure and codebases to enable scalable model development.
- Collaborate with machine learning engineers to optimize model performance, training duration, memory utilization, and GPU costs.
- Optimize batch data processing and pipelines using Apache Beam, Apache Spark, and Ray Data.
- Architect pipelines to build and maintain graph structures with billions of nodes and edges.
- Build and evolve cloud-based infrastructure for machine learning platforms focused on scalability and reliability.
- Administer and integrate MLOps tools for experiment tracking and model serving.
- Develop solutions to improve the machine learning development lifecycle and platform usability.
View Full Description & ApplyYou'll be redirected to the employer's site