Senior Data Engineer (Web Scraping)

New
J
JobgetherData engineering
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Required Skills
PythonHTMLSeleniumRESTful APIsPlaywright

Requirements

  • Demonstrated professional experience building and operating production web-scraping systems at scale.
  • Ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and production support.
  • Strong production-level Python engineering skills, including experience developing maintainable applications rather than standalone scripts.
  • Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium.
  • Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior.
  • Experience addressing pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies.
  • Understanding of data pipelines, data quality, and validation, storage, and downstream consumption of collected data.
  • Experience deploying, monitoring, and supporting production workloads in a cloud environment.
  • AWS experience is desirable.
  • Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg, is a plus.
  • Experience with PySpark or other distributed data-processing technologies is beneficial.
  • Familiarity with Docker, Terraform or other infrastructure-as-code tools, and Grafana or comparable observability platforms is advantageous.

Responsibilities

  • Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines.
  • Design scalable scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling.
  • Build, maintain, and improve production scrapers for new and existing data sources.
  • Investigate websites and select acquisition methods, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling.
  • Evaluate build-versus-buy options for scraping infrastructure and external services.
  • Account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations in web-data acquisition.
  • Diagnose scraping challenges involving website changes, dynamic content, authentication, sessions, rate limits, and concurrency.
  • Integrate scraping workloads into scalable data-platform and lakehouse architectures.
  • Monitor, debug, maintain, and continuously improve production workloads.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now