Principal Staff Engineer, AI Platform Research, Data Science
New
J
JobgetherCybersecurity AI
USFull-TimeStaff
Salary$195,000–$290,000 per year for U.S. candidates.
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years
- Required Skills
- DockerPythonKubernetesGoRustSparkMLOpsLangChainDistributed Systems
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience.
- 5+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at significant scale.
- Previous experience operating at a Staff-level engineering capacity, with demonstrated technical leadership, architecture ownership, and mentorship responsibilities.
- Strong hands-on expertise in LLM engineering, including fine-tuning, prompt engineering, deployment, Retrieval-Augmented Generation, and agentic workflow development.
- Proven experience designing and delivering large-scale distributed systems, including expertise in sharding, partitioning, concurrency, fault tolerance, and performance optimization.
- Expert-level proficiency in at least one programming language such as Python, Go, Rust, or JVM-based technologies.
- Strong experience with AI/ML and MLOps technologies such as MLflow, SageMaker, Vertex AI, LangChain, or LlamaIndex.
- Experience with distributed processing frameworks such as Spark, Dask, or Flink.
- Strong knowledge of cloud environments including AWS, GCP, or OCI and associated data services.
- Experience with Docker and Kubernetes in production environments.
- Familiarity with streaming technologies such as Kafka or Pulsar.
- Experience with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, or Kubeflow.
Responsibilities
- Architect, build, and optimize highly scalable data platforms and pipelines supporting LLMs, neural networks, Retrieval-Augmented Generation (RAG), and AI agentic systems at Exabyte scale.
- Design and deploy agentic workflows and agent-harnessing capabilities that enable autonomous, data-driven security solutions.
- Remain hands-on in software development, writing elegant, production-ready code with a strong focus on performance, maintainability, testing, reliability, and rapid delivery.
- Design fault-tolerant, cost-effective distributed systems using advanced approaches to sharding, partitioning, concurrency, and large-scale data processing.
- Provide technical leadership across data modeling, normalization, semantic cataloging, and platform architecture for AI/ML workloads.
- Establish and advance MLOps and DataOps standards for LLM-powered systems, including monitoring, observability, automation, and zero-touch recovery.
- Own the complete lifecycle of critical AI and data services, from architecture and development through testing, deployment, production monitoring, and continuous improvement.
- Partner with Data Scientists, Product Managers, and engineering teams to transform research prototypes into scalable, secure, production-grade services.
- Drive adoption of modern AI platform technologies and continuously evaluate emerging tools, frameworks, and development practices.
- Mentor engineers through technical workshops, architecture discussions, design reviews, and hands-on guidance.
View Full Description & ApplyYou'll be redirected to the employer's site