Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.
Principles of Data Science
3,000 words
Principles of Data Science – Predictive Modelling and Data Analysis
This individual assessment for the Principles of Data Science module requires students to select, apply and critically evaluate data science methods, tools and techniques using one of three provided datasets and its associated scenario. The main assessment takes the form of a 3,000-word report in which students explore their chosen dataset, identify an appropriate predictive modelling approach, build and evaluate models, interpret the findings and critically reflect on the overall process and outcomes. The assessment addresses the principles and foundations of data science, statistical methods, data preparation, visualisation, predictive modelling, decision making and the critical evaluation of data science techniques and tools. Students begin by exploring the selected dataset to understand its structure, characteristics and limitations. Although the supplied datasets have already been cleaned, students may undertake additional data preparation or transformation where necessary. Any preprocessing decisions must be justified in relation to the requirements of the selected analytical methods. Feature selection should also be considered as part of preparing the data for model development. The assessment requires students to identify suitable forms of analysis for the selected scenario and justify their choice of methods. At least two different techniques must be used to develop models with predictive capacity for the response variable in the chosen dataset. The models must be trained and tested consistently, using the same training and test datasets so that their performance can be compared fairly. Where appropriate, students should also provide insight into feature importance and explain the contribution of relevant variables to predictive performance. Model performance must be evaluated using suitable metrics, followed by a clear description of the findings and recommendations appropriate for the intended audience. The report should document the complete analytical workflow, including data exploration, preprocessing, feature selection, model development, testing and evaluation. Students are expected to explain and justify the decisions made throughout the process rather than simply presenting code or model results. The assessment also requires students to demonstrate practical proficiency in data science tools and techniques. The brief expects the use of R for completing the assignment and requires evidence of important elements of the code, although the complete code does not need to be submitted. Data visualisation must be used to support the written discussion and communicate relevant findings effectively. The assessment is evaluated across theoretical knowledge and method selection, data exploration and processing, technical application and model evaluation, communication of findings, and overall presentation and referencing. The assessment therefore combines technical implementation with critical analysis, requiring students to explain why particular methods were selected, evaluate their effectiveness and consider the limitations and implications of the resulting findings. A separate second assessment component accompanies the written report. This component requires a presentation of the key findings from the written work using a maximum of five slides and a presentation duration of no more than seven minutes. It should summarise the dataset, methods, key findings and project outcomes while providing critical reflective commentary on lessons learned, factors affecting success and potential real-world applications.
Read Model Answer →
Principles of Data Science
2,000 words
Principles of Data Science – Data Analysis Portfolio
This portfolio assignment for the Principles of Data Science module at Coventry University requires students to analyse the Global Life-Work Balance Index 2025 dataset using statistical and data science techniques in R. The dataset ranks 60 countries according to life-work balance using factors including statutory annual leave, paid maternity leave, sick leave, healthcare, public safety, public happiness, LGBTQ inclusivity and average working hours per employee. The assignment has a 2,000-word equivalent limit, excluding the reference list and output. The portfolio consists of two main tasks. Task 1 is a group task involving multivariate data analysis. Students must use R to perform Principal Component Analysis (PCA) and Cluster Analysis on the dataset. For PCA, students analyse quantitative variables, produce and interpret relevant visualisations such as screeplots, biplots and loadings plots, and investigate the effects of Region and Healthcare System. The PCA analysis also requires comparison of the overall dataset with countries from Europe. The cluster analysis component requires students to cluster both countries and variables using different distance metrics and hierarchical clustering methods. Students compare methods such as Manhattan and Euclidean distances and single linkage and Ward’s method, present comparisons in compact tables, and interpret relevant dendrograms. They must then compare the conclusions obtained from PCA and Cluster Analysis, identifying common insights and apparent conflicts and discussing the extent to which the results are explainable rather than simply interpretable. Task 2 is an individual task focusing on Exploratory Data Analysis and Linear Models. Students create a scatter matrix using ggpairs(), investigate strongly correlated variables, and identify quantitative variables that may help predict Region for European and Asian countries. They then develop and critically assess linear regression models for predicting Score, including models based on employment variables and broader quantitative predictors. Model comparison and selection use concepts including AIC, while diagnostic plots are used to identify countries requiring further investigation. The individual task also requires students to use European Life-Work Balance Index 2023 data to make predictions for European countries not included in the 2025 dataset and to construct a Residuals versus Fitted Values plot. Finally, students must combine the conclusions from their individual linear modelling work with the PCA and Cluster Analysis findings to identify specific discoveries about the variables and countries in the dataset. R code, output and relevant plots must be included directly within the reports. The assignment encourages use of the R tidyverse and requires appropriate referencing of sources. The brief specifies APA-style referencing for the individual and group work. It also states that generative AI may be used for inspiration but not for generating answers or analysing the datasets, and any permitted AI use must be acknowledged and documented.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.
Read Model Answer →
Artificial Intelligence
2,000 words
End-to-End Applied AI Development — Comparative Machine Learning and Neural Network Modelling on a Public Dataset
This assessment runs a complete applied AI development cycle end to end: problem definition, dataset selection, preprocessing, model building, optimisation, evaluation and critical reflection. Students identify a real-world problem themselves, formulate a research question from it, and source a suitable dataset from a recognised public repository such as UCI, Kaggle, Data.gov or OpenML. Dataset choice carries more weight than students expect. It must be genuinely suitable for supervised learning, complex enough to make preprocessing and feature engineering meaningful, and — critically — structured so that a traditional machine learning approach and a deep learning approach can be sensibly compared on it. A dataset too small or too clean makes the neural network component pointless; one too large or too noisy makes the whole pipeline unfinishable within the page limit. The source must be referenced and the choice explicitly justified against the research problem. The modelling requirement is fixed: at least two supervised machine learning models, plus one artificial neural network built in a mainstream deep learning framework, all trained and tested. The comparison between them is the analytical core of the work. Reporting that the neural network scored higher is not an answer; explaining why, in terms of the data's structure and each model's inductive assumptions, is. Marks are distributed across problem framing, the traditional models, the deep learning model, evaluation and critical analysis including responsible AI considerations, and academic communication. That responsible AI component is easy to overlook and is not decorative — it asks what the model's limitations mean for anyone who might rely on it. Presentation requirements are specific. The report is page-limited rather than purely word-limited, and every plot must be described in the text while also being legible enough to communicate on its own — a common failure is dense default library output pasted in without axis labels or scale. The implementation is documented in a notebook combining markdown and code cells so the development process is visible, not just the final result, and submissions typically include the cleaned dataset alongside the code. The strongest submissions treat the notebook and the report as one argument. Weaker ones produce a working notebook and then write a report that describes it, rather than a report that uses it as evidence. Our support on assessments of this type is guidance-based. Typical areas of help include: advising on whether a candidate dataset can actually support the required model comparison, explaining how to justify preprocessing decisions, clarifying which evaluation metrics suit which problem type and why accuracy alone is often misleading, showing how to structure a critical limitations and responsible AI discussion, checking Harvard referencing, and reviewing a student's own draft against the published marking criteria.
Read Model Answer →