Senior Data Product Engineer
New
J
JobgetherCybersecurity Data
Fully remote across LATAM, flexible working hoursFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Languages
- English proficiency at B2 level or higher
- Required Skills
- AWSGraphQLPostgreSQLPythonAirflowSeleniumPlaywright
Requirements
- Strong hands-on experience designing and operating advanced web-scraping systems at scale.
- Demonstrated experience working with anti-scraping and bot-protection mechanisms, including proxy pools, headless browsers, rate limiting, IP blocking, fingerprinting, or similar techniques.
- Strong professional experience with Python and data processing, including production-grade pipeline development.
- Extensive experience with web-scraping and DOM-parsing technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
- Hands-on AWS experience, particularly with S3 and boto3; familiarity with EKS, IAM/IRSA, Parameter Store, and ECR is highly valuable.
- Experience building reliable, schedulable data pipelines with Airflow or an equivalent orchestration platform.
- Strong knowledge of PostgreSQL, SQL, relational database fundamentals, schema design, and data-access patterns.
- Ability to take end-to-end ownership of technical solutions, make sound architectural decisions, and work effectively with a high degree of autonomy.
- Experience participating in code reviews and maintaining strong software engineering standards.
- English proficiency at B2 level or higher.
Responsibilities
- Architect, build, and scale distributed web-scraping systems capable of reliably collecting data from thousands of online sources.
- Develop robust approaches for handling anti-scraping and bot-protection mechanisms, including proxy rotation, CAPTCHAs, rate limiting, IP blocking, browser fingerprinting, and headless browsing.
- Build and maintain data extraction workflows using technologies such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, and CSS selectors.
- Refactor and productionize existing Python data-collection pipelines, strengthening reliability, observability, error handling, retry mechanisms, monitoring, and alerting.
- Build schedulable, containerized ingestion workflows using Airflow or comparable orchestration technologies.
- Design PostgreSQL schemas, views, partitioning strategies, and efficient data-access patterns for processed web data.
- Develop and maintain cloud infrastructure and data services using AWS technologies including S3, EKS, IAM/IRSA, Parameter Store, and ECR.
- Build and maintain the data-access layer using GraphQL and Hasura, while collaborating with Data Science and SaaS teams on reliable API contracts.
View Full Description & ApplyYou'll be redirected to the employer's site