Principal Data Engineer, LLM/AI Platforms
New
J
JobgetherData Engineering, AI
Based in United StatesFull-TimePrincipal
SalaryCAD $210,000–$320,000 annual base salary for Canadian-based employment, plus variable/incentive compensation, equity, and benefits.
Apply NowOpens the employer's application page
Job Details
- Experience
- 10+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms
- Required Skills
- AWSDockerPythonGCPJVMKafkaKubernetesSpark
Requirements
- Master’s degree or PhD in Computer Science, Data Engineering, or related STEM discipline, or equivalent practical experience.
- 10+ years of progressive experience in Data Engineering or Platform Engineering.
- At least 3 years architecting and building AI/ML or Data Science platforms at massive scale.
- 3+ years of experience in a Principal or Staff-level engineering capacity with technical leadership and mentorship experience.
- Hands-on expertise with LLM engineering including fine-tuning, prompt engineering, RAG, and agentic workflow development.
- Proven experience designing and delivering large-scale distributed systems, including sharding, partitioning, concurrency, and fault-tolerant architectures.
- Expert-level proficiency in Python or JVM-based technologies.
- Deep experience with distributed data processing frameworks such as Spark, Dask, or Flink.
- Strong knowledge of cloud platforms (AWS, GCP, or OCI) and their associated data services.
- Expertise with containerization and orchestration technologies including Docker and Kubernetes.
- Experience with messaging and streaming technologies such as Kafka or Pulsar.
- Familiarity with data warehousing/orchestration (Snowflake, BigQuery, Airflow, Kubeflow) and MLOps technologies (MLflow, SageMaker, Vertex AI).
Responsibilities
- Architect, implement, and optimize data platforms and pipelines designed for LLMs, RAG, and advanced AI agentic systems at Exabyte scale.
- Drive the adoption and deployment of agentic workflows and agent-harnessing techniques to support autonomous, data-driven capabilities.
- Design highly scalable, fault-tolerant, secure, and cost-effective data solutions that enable rapid iteration.
- Develop production-ready code with strong attention to performance, maintainability, testing, and operational reliability.
- Provide technical leadership in data modeling, normalization, semantic cataloging, and data architecture for AI and machine learning workloads.
- Establish MLOps and DataOps best practices including monitoring, observability, and automated recovery.
- Own the end-to-end lifecycle of critical data services from development through deployment and continuous optimization.
- Collaborate with data scientists and product managers to transform research prototypes into robust, production-ready services.
- Lead technical workshops, design reviews, and mentor engineers across teams.
- Champion DevSecOps practices and engineering standards across large-scale distributed data environments.
View Full Description & ApplyYou'll be redirected to the employer's site