Senior Python Data Engineer (OCR & Document Processing)
New
I
InetumInsurance
RomaniaFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5-10 years
- Required Skills
- AWSPythonSQLGitData engineeringCI/CD
Requirements
- 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields.
- Proven experience building scalable data ingestion and processing pipelines.
- Experience working with large volumes of unstructured and semi-structured data.
- Experience designing cloud-based data solutions.
- Strong programming skills in Python.
- Strong SQL knowledge.
- Hands-on experience with AWS services, including S3, Step Functions, and CloudWatch.
- Experience processing unstructured documents such as PDF, Word, Excel, PowerPoint, and Email content.
- Experience building connectors and integrations with enterprise content repositories (e.g., SharePoint).
- Experience with OCR and document extraction tools (AWS Textract or equivalent).
- Experience designing and implementing data ingestion and transformation pipelines.
- Familiarity with vector databases and Retrieval-Augmented Generation (RAG) concepts.
- Experience with software development best practices: Git, CI/CD, and automated testing.
Responsibilities
- Design and implement scalable data ingestion pipelines for processing high volumes of unstructured documents, including PDFs, scans, emails, and Office files.
- Integrate, configure, and optimize OCR and document extraction technologies to maximize text extraction accuracy and document understanding.
- Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
- Develop connectors and integrations for document sources such as SharePoint, email systems, and enterprise repositories.
- Design and maintain vector database schemas and retrieval mechanisms to support Retrieval-Augmented Generation (RAG) solutions and AI applications.
- Ensure document processing pipelines meet enterprise security, compliance, performance, and availability requirements.
- Implement monitoring, validation, and quality-control mechanisms to identify and manage low-confidence OCR and extraction results.
- Optimize data processing workflows for scalability, reliability, and low-latency operations.
- Collaborate with AI Engineers, Backend Engineers, and Platform teams to deliver end-to-end AI-powered document processing solutions.
- Develop and maintain cloud-native data ingestion solutions on public cloud platforms.
View Full Description & ApplyYou'll be redirected to the employer's site