Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.
Read Model Answer →