Senior Data Engineer (Web Scraping)
New
J
JobgetherData engineering
Based in IndiaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Required Skills
- PythonHTMLSeleniumRESTful APIsPlaywright
Requirements
- Demonstrated professional experience building and operating production web-scraping systems at scale.
- Ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and production support.
- Strong production-level Python engineering skills, including experience developing maintainable applications rather than standalone scripts.
- Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium.
- Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior.
- Experience addressing pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies.
- Understanding of data pipelines, data quality, and validation, storage, and downstream consumption of collected data.
- Experience deploying, monitoring, and supporting production workloads in a cloud environment.
- AWS experience is desirable.
- Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg, is a plus.
- Experience with PySpark or other distributed data-processing technologies is beneficial.
- Familiarity with Docker, Terraform or other infrastructure-as-code tools, and Grafana or comparable observability platforms is advantageous.
Responsibilities
- Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines.
- Design scalable scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling.
- Build, maintain, and improve production scrapers for new and existing data sources.
- Investigate websites and select acquisition methods, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling.
- Evaluate build-versus-buy options for scraping infrastructure and external services.
- Account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations in web-data acquisition.
- Diagnose scraping challenges involving website changes, dynamic content, authentication, sessions, rate limits, and concurrency.
- Integrate scraping workloads into scalable data-platform and lakehouse architectures.
- Monitor, debug, maintain, and continuously improve production workloads.
View Full Description & ApplyYou'll be redirected to the employer's site