This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.
Data Science · ADHD Prediction · Healthcare Analytics · Machine Learning · Classification · WiDS Datathon 2025 · Data Preprocessing · Exploratory Data Analysis · Feature Engineering · Logistic Regression · Random Forest · Gradient Boosting
Megaminds has supported academic requirements in data science / artificial intelligence and machine learning, data science and related disciplines.