Senior Data Engineer

New
A
AlpacaFinancial Services
Remote - North America - LATAMFull-TimeSenior
SalaryCompetitive Salary & Stock Options
Apply NowOpens the employer's application page

Job Details

Experience
5+ years of experience in Data Engineering, including 2+ years building and operating scalable, low-latency data platforms handling > 100M events/day
Required Skills
PythonSQLGCPKafkaKubernetesTerraformAnsible

Requirements

  • 5+ years of experience in Data Engineering, including 2+ years building and operating scalable, low-latency data platforms handling > 100M events/day.
  • Strong hands-on experience running data infrastructure on Kubernetes, with cloud-native tooling like Docker and Helm.
  • Production experience with IaC: Terraform, Ansible, and ArgoCD (or equivalents).
  • Deep knowledge of distributed systems (storage, transactions, and query processing) with hands-on experience operating open-source query engines like Trino or Presto.
  • Strong experience with object storage and open table formats, specifically Apache Iceberg.
  • Experience with streaming and CDC systems: Kafka, Redpanda, and Debezium.
  • Hands-on experience with orchestration frameworks (Airflow) and ELT tools (Airbyte).
  • Strong working knowledge of Python and SQL for building pipelines and platform tooling.
  • Experience with Google Cloud Platform and its data services (GCS, Cloud Build, Cloud SQL, Dataproc, etc); or related experience with other cloud services.
  • Ability to thrive in a fast-paced startup environment and adapt infrastructure to rapidly changing needs.

Responsibilities

  • Design, build, and evolve the core data platform infrastructure e.g., distributed query engines, orchestration, warehousing, cataloging, and more.
  • Own our lakehouse infrastructure as code, managing deployments through Terraform and Ansible on Kubernetes.
  • Build and maintain low-latency streaming and CDC ingestion pipelines, as well as batch ingestion paths landing in Iceberg.
  • Develop and scale our BI landscape so downstream teams and agents get performant, self-serve access to lakehouse data.
  • Enforce platform reliability best practices, including monitoring and alerting, on-call rotations, incident response, maintenance windows, runbooks, and SLAs.
  • Partner with DevOps, Analytics Engineering, and other stakeholders to close infrastructure gaps and support new data requirements.
View Full Description & ApplyYou'll be redirected to the employer's site
Competitive Salary & Stock Options
Apply Now