Principal Staff Engineer, AI Platform Research, Data Science

New
J
JobgetherCybersecurity AI
USFull-TimeStaff
Salary$195,000–$290,000 per year for U.S. candidates.
Apply NowOpens the employer's application page

Job Details

Experience
5+ years
Required Skills
DockerPythonKubernetesGoRustSparkMLOpsLangChainDistributed Systems

Requirements

  • Bachelor's, Master's, or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience.
  • 5+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at significant scale.
  • Previous experience operating at a Staff-level engineering capacity, with demonstrated technical leadership, architecture ownership, and mentorship responsibilities.
  • Strong hands-on expertise in LLM engineering, including fine-tuning, prompt engineering, deployment, Retrieval-Augmented Generation, and agentic workflow development.
  • Proven experience designing and delivering large-scale distributed systems, including expertise in sharding, partitioning, concurrency, fault tolerance, and performance optimization.
  • Expert-level proficiency in at least one programming language such as Python, Go, Rust, or JVM-based technologies.
  • Strong experience with AI/ML and MLOps technologies such as MLflow, SageMaker, Vertex AI, LangChain, or LlamaIndex.
  • Experience with distributed processing frameworks such as Spark, Dask, or Flink.
  • Strong knowledge of cloud environments including AWS, GCP, or OCI and associated data services.
  • Experience with Docker and Kubernetes in production environments.
  • Familiarity with streaming technologies such as Kafka or Pulsar.
  • Experience with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, or Kubeflow.

Responsibilities

  • Architect, build, and optimize highly scalable data platforms and pipelines supporting LLMs, neural networks, Retrieval-Augmented Generation (RAG), and AI agentic systems at Exabyte scale.
  • Design and deploy agentic workflows and agent-harnessing capabilities that enable autonomous, data-driven security solutions.
  • Remain hands-on in software development, writing elegant, production-ready code with a strong focus on performance, maintainability, testing, reliability, and rapid delivery.
  • Design fault-tolerant, cost-effective distributed systems using advanced approaches to sharding, partitioning, concurrency, and large-scale data processing.
  • Provide technical leadership across data modeling, normalization, semantic cataloging, and platform architecture for AI/ML workloads.
  • Establish and advance MLOps and DataOps standards for LLM-powered systems, including monitoring, observability, automation, and zero-touch recovery.
  • Own the complete lifecycle of critical AI and data services, from architecture and development through testing, deployment, production monitoring, and continuous improvement.
  • Partner with Data Scientists, Product Managers, and engineering teams to transform research prototypes into scalable, secure, production-grade services.
  • Drive adoption of modern AI platform technologies and continuously evaluate emerging tools, frameworks, and development practices.
  • Mentor engineers through technical workshops, architecture discussions, design reviews, and hands-on guidance.
View Full Description & ApplyYou'll be redirected to the employer's site
$195,000–$290,000 per year for U.S. candidates.
Apply Now