Senior Data Product Engineer

New
J
JobgetherCybersecurity Data
Fully remote across LATAM, flexible working hoursFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page

Job Details

Languages
English proficiency at B2 level or higher
Required Skills
AWSGraphQLPostgreSQLPythonAirflowSeleniumPlaywright

Requirements

  • Strong hands-on experience designing and operating advanced web-scraping systems at scale.
  • Demonstrated experience working with anti-scraping and bot-protection mechanisms, including proxy pools, headless browsers, rate limiting, IP blocking, fingerprinting, or similar techniques.
  • Strong professional experience with Python and data processing, including production-grade pipeline development.
  • Extensive experience with web-scraping and DOM-parsing technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
  • Hands-on AWS experience, particularly with S3 and boto3; familiarity with EKS, IAM/IRSA, Parameter Store, and ECR is highly valuable.
  • Experience building reliable, schedulable data pipelines with Airflow or an equivalent orchestration platform.
  • Strong knowledge of PostgreSQL, SQL, relational database fundamentals, schema design, and data-access patterns.
  • Ability to take end-to-end ownership of technical solutions, make sound architectural decisions, and work effectively with a high degree of autonomy.
  • Experience participating in code reviews and maintaining strong software engineering standards.
  • English proficiency at B2 level or higher.

Responsibilities

  • Architect, build, and scale distributed web-scraping systems capable of reliably collecting data from thousands of online sources.
  • Develop robust approaches for handling anti-scraping and bot-protection mechanisms, including proxy rotation, CAPTCHAs, rate limiting, IP blocking, browser fingerprinting, and headless browsing.
  • Build and maintain data extraction workflows using technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, and CSS selectors.
  • Refactor and productionize existing Python data-collection pipelines, strengthening reliability, observability, error handling, retry mechanisms, monitoring, and alerting.
  • Build schedulable, containerized ingestion workflows using Airflow or comparable orchestration technologies.
  • Design PostgreSQL schemas, views, partitioning strategies, and efficient data-access patterns for processed web data.
  • Develop and maintain cloud infrastructure and data services using AWS technologies including S3, EKS, IAM/IRSA, Parameter Store, and ECR.
  • Build and maintain the data-access layer using GraphQL and Hasura, while collaborating with Data Science and SaaS teams on reliable API contracts.
View Full Description & ApplyYou'll be redirected to the employer's site
View details
Apply Now