Technical SME Data Engineer
J
JobgetherFederal regulatory analytics
Listing location: US; Workplace type: Remote; based in the United StatesFull-TimeSenior
Salary not disclosed
Apply NowOpens the employer's application page
Job Details
- Experience
- 5+ years of professional experience working with decision analysis and machine learning algorithms in cloud-native environments. 5+ years of experience using Amazon Web Services EMR Spark is preferred.
- Required Skills
- Machine LearningSparkGitHub
Requirements
- Have 5+ years of professional experience working with decision analysis and machine learning algorithms in cloud-native environments.
- Bring hands-on experience with Decision Trees, Random Forests, Gradient Boosted Trees, Linear Regression, Collaborative Filtering, and K-Means.
- Understand Apache Spark architecture and internals, including core APIs, SparkSQL, high-level data access tools, and Spark streaming.
- Have experience cleaning data and developing comparative trend analyses using large-scale datasets.
- Be able to translate organizational goals into working machine learning models, statistical models, or pattern recognition solutions.
- Have experience working with open-source and community solutions and managing source code using GitHub.
- Hold a bachelor’s degree in Mathematics, Data Science, or a similar field.
- Be able to communicate technical findings to program, policy, and technical stakeholders.
- Experience using Amazon Web Services EMR Spark for 5+ years is preferred.
- Experience with Zeppelin or similar data science interpreters is preferred.
- Be familiar with federal AI governance requirements, including model documentation and bias testing.
- Prior experience supporting federal financial regulators is preferred; experience in a federal FISMA Moderate or comparable security and compliance environment is beneficial.
- A master’s degree in Mathematics, Data Science, or a related discipline is strongly preferred.
Responsibilities
- Design, build, and implement decision analysis and machine learning models, including Decision Trees, Random Forests, Gradient Boosted Trees, Linear Regression, Collaborative Filtering, and K-Means.
- Develop large-scale data processing pipelines using Apache Spark core APIs, SparkSQL, and streaming capabilities.
- Clean, transform, and validate data, and perform comparative trend analysis across large datasets.
- Translate organizational goals and business requirements into machine learning models, statistical models, and pattern recognition solutions.
- Implement analytical solutions in cloud-native environments and support scalable data processing.
- Ensure AI and machine learning initiatives comply with applicable federal governance policies, including documentation and bias testing.
- Maintain source code through GitHub and contribute to open-source and community-driven solutions.
- Collaborate with program, policy, and technical stakeholders to communicate analytical findings, model results, and recommendations.
- Contribute subject matter expertise to technical discussions, solution development, and project delivery.
View Full Description & ApplyYou'll be redirected to the employer's site