We are seeking an ML Data Engineer, based in Cairo, with strong experience in building scalable data pipelines, feature stores, and governed ML data workflows. The ideal candidate has 7+ years in data engineering roles with a focus on machine learning enablement using Cloudera, PySpark, and MLFlow. You will lead the creation of reusable, traceable, and production-ready features to accelerate ML development and deployment.
What You Will Do:
Design and implement scalable data pipelines using PySpark, Hive, and Cloudera platforms.
Build and maintain centralized feature stores for ML models ensuring training-serving consistency.
Collaborate with data scientists to transform raw datasets into production-grade features.
Integrate data validation, versioning, and schema governance using tools like Great Expectations.
Monitor feature drift, lineage, and availability to ensure model accuracy and trustworthiness.
Enable fast experimentation through curated, documented, and metadata-rich feature sets.
Maintain feature version control and auditability using DVC, Delta Lake, and MLFlow integration.
Streamline ML data pipelines to reduce latency and support both batch and streaming use cases.
Build PySpark pipelines across batch and real-time systems integrated with Hive, HDFS, and Azure Data Lake.
Develop and operate centralized feature repositories using Feast, Hopsworks, or custom implementations.
Implement feature validation checks and quality gates using Great Expectations or PyDeequ.
Maintain metadata tracking, audit logs, and access control aligned with data governance frameworks.
Connect feature pipelines with MLFlow to enable experiment tracking and downstream reuse.
Optimize pipeline performance for low-latency serving in high-throughput environments.
Support schema evolution, feature versioning, and reusability across model teams.
What We Are Looking For:
7–8+ years in data engineering or ML data engineering roles, focused on machine learning enablement.
Hands-on experience with feature stores, ETL workflows, and data validation frameworks.
Strong background in Cloudera, Hadoop, Spark, PySpark, and HDFS environments.
Proficient in Python, SQL, and modern data versioning tools like DVC and Delta Lake.
Familiarity with MLFlow, metadata tracking, and reusable data architecture best practices.
Experience working with governance protocols: audit trails, access control, and schema enforcement.
Bachelor’s or Master’s in Data Engineering, Computer Science, or a related discipline.
How to Apply:
If you are passionate about building governed, reusable, and scalable data systems that empower machine learning teams, we’d love to connect. Please submit your latest resume and a technical portfolio showcasing your feature store work, PySpark pipelines, and metadata governance practices.