Telecom Churn Prediction
Production-grade telecom churn prediction system: a config-driven scikit-learn pipeline with SHAP explainability, MLflow experiment tracking and FastAPI serving, trained on…
Data Engineer · AI Engineer · Data Scientist
I’m Muhammad Farooq (also known as Muhammad Farooq Shafi), a data engineer and machine learning practitioner focused on building systems that hold up under real scrutiny — not just notebooks that ran once. My work spans the full data lifecycle: ETL pipelines with automated data-quality gates, dbt-modeled analytics warehouses backed by dozens of automated tests, and ML systems evaluated honestly against real baselines rather than a single flattering metric. I’ve built and deployed projects across intrusion detection, recommendation systems, NLP triage, demand forecasting, and computer vision — each one served through a real API or interface, versioned, and tested, not left as a one-off script. I write regularly about the engineering discipline behind trustworthy data work: honest evaluation, reproducible pipelines, and deployment done right.
Production-grade telecom churn prediction system: a config-driven scikit-learn pipeline with SHAP explainability, MLflow experiment tracking and FastAPI serving, trained on…
An end-to-end marketing analytics data pipeline that ingests, transforms and models campaign data for reporting.
A knowledge-base search system built with retrieval-augmented generation for accurate, grounded question answering.
A time-series forecasting system for energy demand, built for planning and load-balancing use cases.
Graph-based analysis of organizational networks to surface influence, collaboration and communication patterns.
A classical machine learning system for detecting fraudulent transactions with a focus on precision/recall trade-offs.
A live snapshot of my GitHub activity — commits, streaks and languages.
View GitHub profile →Machine Learning Engineer (Internship) · CodeAlpha
Developed and implemented machine learning models for real-world applications using Python, TensorFlow and scikit-learn. Performed data preprocessing, feature selection and model tuning, and collaborated with cross-functional teams to integrate ML solutions into existing products.
Data Science Intern (Internship) · Arch Technologies
Assisted in data collection, cleaning and preprocessing to support machine learning projects. Developed and tested predictive models using Python and libraries including scikit-learn and pandas, and created data visualizations to communicate insights.
Data Science Intern (Internship) · Oasis Infobyte
Supported data preprocessing, feature engineering and exploratory data analysis for ongoing machine learning projects, and assisted in building and evaluating predictive models with scikit-learn and pandas.
Data Science Job Simulation (Forage) · British Airways
Scraped and analysed customer review data to uncover findings, and built a predictive model to understand the factors that influence buying behaviour.
Data Science Job Simulation (Forage) · BCG X
Completed a customer churn analysis simulation, conducting data analysis in Python (pandas, NumPy), engineering and optimizing a random forest model that achieved a 50% recall rate in predicting churn, and delivering an executive summary of findings.
Data Science Job Simulation (Forage) · Commonwealth Bank
Built data engineering pipelines to aggregate and extract insights from datasets, applied anonymization for data-privacy compliance, and proposed data analysis approaches for well-structured, efficient databases.
Data Analytics Job Simulation (Forage) · Deloitte
Completed a job simulation involving data analysis and forensic technology, created a data dashboard using Tableau, and used Excel to classify data and draw business conclusions.
Data Analytics Job Simulation (Forage) · Quantium
Developed expertise in data preparation and customer analytics using transaction datasets, identified benchmark stores for uplift testing, and created reports for informed, evidence-based strategic decisions.
Data Science Job Simulation (Forage) · Lloyds Banking Group
Completed a Forage job simulation for Lloyds Banking Group's data science team.
Bachelor's degree, Artificial Intelligence
Jul 2023 – May 2027
Focused on the design, development and deployment of AI-driven systems. Coursework and projects covered machine learning, deep learning, computer vision, natural language processing and robotics, with hands-on expertise in Python, TensorFlow and PyTorch.
Associate's degree, Data Science
Apr 2024 – Mar 2026
Developed a strong foundation in data analysis, statistics, programming and machine learning.
Associate's Degree, Business Administration and Management
Jun 2025 – Jun 2026
Participated in business case studies, management discussions and leadership development activities focused on strategy, marketing and organizational management.
IBM · Credential ID MFXW3LL9UI6G
Google · Jun 2025
Coursera · Jun 2025
Coursera
IBM · Jun 2025 · Credential ID FVMJDQSAUDL6
HP LIFE · Feb 2025
HP LIFE · Mar 2025
Udemy
HP LIFE · Feb 2025
HP LIFE · Feb 2025
Udemy · Nov 2024
Udemy · Nov 2024
Udemy
Design and build reliable, monitored pipelines that move data from source systems into warehouses and marts, ready for analytics.
Model and deploy star-schema warehouses on Snowflake, Databricks or DuckDB with automated data-quality testing.
Train, explain and serve ML models with MLflow tracking and FastAPI/Docker deployment.
Turn raw data into decision-ready dashboards with Streamlit, Power BI or Tableau.
A wrong number on a dashboard raises a question: where did it come from? Here’s what data lineage in data engineering actually tracks, and how dbt provides it almost for free.
Read more →
Reprocessing a full table works until it doesn’t. Here’s change data capture vs batch ETL, when each is the right fit, and a real project example of batch done well.
Read more →
A running total or rank within a group used to mean a slow, hard-to-read subquery. Here’s window functions vs subqueries in SQL, and when each actually fits.
Read more →I work on data pipelines, warehouse modeling, ETL/ELT, and applied ML/AI projects — from a single dashboard to a full production pipeline.
Yes — update this answer with your current availability.
Python, SQL, dbt, Airflow, Docker, and cloud platforms including AWS, Azure and GCP, alongside scikit-learn/TensorFlow/PyTorch for ML work.
Send a message and I'll get back to you soon — or reach out directly: