Academic Model Answers
Library for UK Postgraduates

Browse tutor-verified model answers across MBA, Law, Finance, Research Methods and more. Use as study references for your own work.

247 model answers 30+ subjects covered 50+ UK universities
Find your assignment

Search the Library

Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.

Filtering by “Missing Values” Clear filters

Available Model Answers (4)

Real-time Database Sync
Machine Learning on Big Data 3,000 words

Machine Learning on Big Data using PySpark

This group assessment for the Machine Learning on Big Data module requires students to apply machine learning techniques to a large real-world dataset using PySpark DataFrame and Spark ML. The coursework is designed to develop practical understanding of big data processing, machine learning implementation, model optimisation and evaluation. Students work in groups of three or four and complete a machine learning project using a suitable large dataset obtained from a benchmarking source such as Kaggle, their workplace or another valid resource. The recommended dataset size is between 300 MB and 1 GB. The main task requires students to select a dataset containing multiple classes rather than a binary classification problem and develop an appropriate machine learning solution. The project begins with loading and preprocessing the data using PySpark DataFrame. Students are expected to address issues such as missing values, data normalisation, feature engineering and class imbalance. For text datasets, additional preprocessing techniques such as stemming, lemmatisation and TF-IDF may be applied. Students must then select and implement an appropriate machine learning method using PySpark MLlib or the Spark ML package. The brief permits approaches including multiclass classification, clustering, ensemble learning and text mining. Model parameters must be optimised using techniques such as grid search or random search. Students are required to evaluate their trained models using suitable metrics, including accuracy, F1-score, precision, recall and classification matrices where appropriate. The brief encourages students to maintain model accuracy and robustness above 90% for the highest possible mark and to discuss steps taken to address bias and variance. The results must be visualised or printed clearly, with appropriate interpretation and analysis of the findings. Students must also consider Legal, Social, Ethical and Professional (LSEP) issues throughout the project. Each student selects one LSEP principle and discusses relevant concerns such as dataset bias, privacy or ethical implications, together with suitable mitigation strategies. The final report should be approximately 3,000 words with a tolerance of ±10% and submitted as a single HTML report using the template provided on the module Moodle site. The report should consolidate the individual contributions of all group members into one comprehensive and user-friendly analytics report. The main assessment is weighted 60% for the report and 40% for the group presentation. The presentation is conducted online through Microsoft Teams, and all group members must participate. The assessment evaluates understanding of Spark, preprocessing, modelling, optimisation, evaluation and the ability to explain and interpret the implemented solution.

Read Model Answer →
Big Data Analytics / Machine Learning 3,000 words

Machine Learning on Big Data Using PySpark: Large-Scale Data Analysis and Predictive Modelling

This group-based Machine Learning on Big Data project requires students to apply machine learning techniques to a large real-world dataset using PySpark DataFrames and Spark machine-learning libraries. Students select a substantial dataset, ideally between approximately 300 MB and 1 GB, from sources such as Kaggle, workplace data or other valid repositories, and develop an end-to-end big-data analytics workflow. CN7030 CRWK 26T1 The project begins with data loading and preprocessing using PySpark. Students are expected to handle missing values, perform data normalisation and feature engineering, identify class imbalance and propose appropriate mitigation strategies. Where text datasets are selected, additional preprocessing may include stemming, lemmatization and TF-IDF representation. CN7030 CRWK 26T1 The modelling stage requires implementation of an appropriate machine-learning approach using PySpark MLlib or Spark ML. The brief expects a multiclass rather than binary classification problem and allows techniques including multiclass classification, ensemble learning, clustering and text mining. Students must justify their model choice and consider model robustness, bias and variance when attempting to improve predictive performance. CN7030 CRWK 26T1 Students then perform hyperparameter tuning using techniques such as grid search or random search and evaluate the resulting model with appropriate measures. Relevant evaluation outputs may include accuracy, F1-score, precision, recall and a confusion matrix. Results should also be visualised or clearly presented and interpreted to identify meaningful patterns and performance characteristics. CN7030 CRWK 26T1 The project additionally requires consideration of Legal, Social, Ethical and Professional (LSEP) issues. Students discuss potential ethical concerns associated with their dataset, including bias and privacy risks, and propose suitable mitigation strategies. The final work is consolidated into a single user-friendly HTML analytics report that clearly presents the group's preprocessing, modelling, optimisation, evaluation and interpretation. CN7030 CRWK 26T1 CN7030 CRWK 26T1 Overview word count: approximately 335 words. If you are also uploading the presentation separately to the Reference Library, that should be a second entry under “Presentations and Academic Posters”, because the presentation forms a distinct 40% component and assesses understanding of Spark, preprocessing, modelling, optimisation, evaluation and responses to examiner questions. CN7030 CRWK 26T1

Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning 2,500 words

Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science

This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.

Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning 2,500 words

Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science

This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.

Read Model Answer →