Senior Data/ML Engineer - Onboarding

New
S
SardineFinancial Crime/Fraud
Remote - United States or CanadaFull-TimeSenior
SalaryUS: $150K – $205K • Offers Equity; Canada: CA$170K – CA$255K • Offers Equity
Apply NowOpens the employer's application page

Job Details

Experience
8+ years
Required Skills
PythonSQLGCPSparkBigQuery

Requirements

  • 8+ years building production data and ML systems with ownership of both pipelines and models.
  • Deep proficiency in Python and SQL.
  • Fluency in distributed processing frameworks such as Spark, Beam, or Flink.
  • Strong understanding of streaming semantics like windowing, watermarks, and exactly-once processing.
  • Hands-on experience with cloud data stacks, specifically GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or AWS equivalents.
  • Experience with Docker, Kubernetes, Terraform, and CI/CD.
  • Expertise in feature stores, training/serving skew, gradient-boosted tree models, and model explainability.
  • Experience with high-volume, low-latency serving environments.
  • Domain experience in fraud, risk, payments, lending, or identity/KYC.
  • Knowledge of data governance in regulated environments including PII, encryption, and auditability.
  • Strong written communication skills for documentation and cross-functional collaboration.
  • Ability to work with high ambiguity and bias toward action.

Responsibilities

  • Own the data ingestion layer for device telemetry, transaction events, and KYC signals using streaming (Pub/Sub, Beam, Flink) and batch pipelines (Airflow, Spark).
  • Build and evolve the feature platform, computing features for both streaming and batch, served to rules engines and models under a sub-second budget.
  • Establish feature correctness through streaming-versus-batch reconciliation, recomputation tests, and drift monitoring.
  • Productionize fraud and identity ML models on Vertex AI and Kubeflow, including automated retraining and champion/challenger promotion.
  • Engineer complex risk signals for KYC, AML, and identity verification, transforming multi-vendor data into model-ready features.
  • Integrate third-party enrichment providers and manage consortium network failover, caching, and cost.
  • Manage the warehouse and modeling layer in BigQuery, including partitioning, staging, and migration strategies.
  • Design entity resolution and graph data structures to link customer identifiers across clients.
  • Enforce data security and residency requirements, including PII handling and field-level encryption.
View Full Description & ApplyYou'll be redirected to the employer's site
US: $150K – $205K • Offers Equity; Canada: CA$170K – CA$255K • Offers Equity
Apply Now