Principal Data Engineer, LLM/AI Platforms

New
J
JobgetherData Engineering, AI
Based in United StatesFull-TimePrincipal
SalaryCAD $210,000–$320,000 annual base salary for Canadian-based employment, plus variable/incentive compensation, equity, and benefits.
Apply NowOpens the employer's application page

Job Details

Experience
10+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms
Required Skills
AWSDockerPythonGCPJVMKafkaKubernetesSpark

Requirements

  • Master’s degree or PhD in Computer Science, Data Engineering, or related STEM discipline, or equivalent practical experience.
  • 10+ years of progressive experience in Data Engineering or Platform Engineering.
  • At least 3 years architecting and building AI/ML or Data Science platforms at massive scale.
  • 3+ years of experience in a Principal or Staff-level engineering capacity with technical leadership and mentorship experience.
  • Hands-on expertise with LLM engineering including fine-tuning, prompt engineering, RAG, and agentic workflow development.
  • Proven experience designing and delivering large-scale distributed systems, including sharding, partitioning, concurrency, and fault-tolerant architectures.
  • Expert-level proficiency in Python or JVM-based technologies.
  • Deep experience with distributed data processing frameworks such as Spark, Dask, or Flink.
  • Strong knowledge of cloud platforms (AWS, GCP, or OCI) and their associated data services.
  • Expertise with containerization and orchestration technologies including Docker and Kubernetes.
  • Experience with messaging and streaming technologies such as Kafka or Pulsar.
  • Familiarity with data warehousing/orchestration (Snowflake, BigQuery, Airflow, Kubeflow) and MLOps technologies (MLflow, SageMaker, Vertex AI).

Responsibilities

  • Architect, implement, and optimize data platforms and pipelines designed for LLMs, RAG, and advanced AI agentic systems at Exabyte scale.
  • Drive the adoption and deployment of agentic workflows and agent-harnessing techniques to support autonomous, data-driven capabilities.
  • Design highly scalable, fault-tolerant, secure, and cost-effective data solutions that enable rapid iteration.
  • Develop production-ready code with strong attention to performance, maintainability, testing, and operational reliability.
  • Provide technical leadership in data modeling, normalization, semantic cataloging, and data architecture for AI and machine learning workloads.
  • Establish MLOps and DataOps best practices including monitoring, observability, and automated recovery.
  • Own the end-to-end lifecycle of critical data services from development through deployment and continuous optimization.
  • Collaborate with data scientists and product managers to transform research prototypes into robust, production-ready services.
  • Lead technical workshops, design reviews, and mentor engineers across teams.
  • Champion DevSecOps practices and engineering standards across large-scale distributed data environments.
View Full Description & ApplyYou'll be redirected to the employer's site
CAD $210,000–$320,000 annual base salary for Canadian-based employment, plus variable/incentive compensation, equity, and benefits.
Apply Now