Academic Model Answers
Library for UK Postgraduates

Browse tutor-verified model answers across MBA, Law, Finance, Research Methods and more. Use as study references for your own work.

201 model answers 30+ subjects covered 50+ UK universities
Find your assignment

Search the Library

Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.

Filtering by “Exploratory data analysis” Clear filters

Available Model Answers (13)

Real-time Database Sync
Time Series Modelling

Time Series Modelling Case Study – Oil Price Forecasting

This individual Time Series Modelling Case Study focuses on analysing and forecasting oil price data using established time-series techniques and an alternative modelling approach. The assessment requires students to work with daily oil price information covering the period from 2024 to 2026 and investigate the underlying patterns, stationarity and forecasting behaviour of the data. The first part of the assignment involves exploratory data analysis and time-series modelling using an ARMA-based approach. Students are required to create appropriate visualisations of the data, perform exploratory analysis and conduct tests for non-stationarity, including relevant stationarity diagnostics such as ADF, ACF and PACF analysis and differencing where required. An ARMA model must then be defined, with suitable model parameters identified using the AIC likelihood approach. The assessment requires the student to examine possible combinations of model parameters, assess model residuals, evaluate model performance using appropriate metrics such as RMSE, and produce forecasts extending 24 months into the future. Confidence intervals must also be included with the forecasts. The second part requires students to investigate an alternative modelling solution for the same oil-price time-series data. Possible approaches discussed in the assignment include models such as LSTM and Prophet. Students are expected to conduct a literature review relating to the selected alternative model, build and apply the model, tune relevant hyperparameters where appropriate, generate 24-month forecasts, create suitable visualisations and calculate appropriate evaluation metrics. The final part of the assessment requires a 6–8 page report describing the modelling process, forecasts, analysis and inferences. The report should explain the reasoning behind the analytical and modelling choices rather than simply presenting numerical results. Students are expected to critically discuss why particular approaches were selected, how the modelling decisions may have influenced the results, how forecasts compare with subsequently collected real data where available, and what improvements could be made in future work. The assessment evaluates both the technical implementation and the quality of the written analysis. The code component assesses completion of the modelling and forecasting tasks, stationarity testing, the alternative solution, code quality and annotation. The report component assesses discussion of the analysis and inferences, comparison of the modelling approaches, clarity of interpretation, report structure, appropriate use of figures and suitable academic references. The submission consists of a report in PDF or Word format, with the code submitted separately or through an appropriate repository.

Read Model Answer →
Data Science / Time Series Analysis / Machine Learning

Time Series Modelling Case Study: Oil Price Forecasting with ARMA and Alternative Models

This Time Series Modelling Case Study requires students to analyse real-world oil-price time-series data and develop forecasting models capable of predicting future values. The assessment combines traditional statistical time-series techniques with an alternative forecasting approach, requiring students to demonstrate practical modelling skills, critical research engagement and evidence-based interpretation of forecasting results. The coursework is completed individually and contributes 40% of the assessment. assing,,, (1) The assessment is divided into three main parts. Part 1 focuses on developing an ARMA-based forecasting model using daily oil-price data covering approximately 2024 to 2026. Students begin with exploratory data analysis and initial visualisation before testing whether the time series is stationary. Where necessary, appropriate transformations or differencing must be applied to obtain stationarity. Students then define an ARMA model and identify suitable p, d and q parameters using an AIC-based model-selection procedure across the parameter ranges specified in the brief. assing,,, (1) Model adequacy must be assessed using diagnostic analysis. Students inspect residuals, generate additional ACF plots, examine residual distributions and evaluate prediction performance using appropriate metrics such as RMSE. The selected model is then used to forecast oil prices 24 months into the future, with appropriate confidence intervals added to communicate forecast uncertainty. assing,,, (1) Part 2 requires students to research and implement an alternative forecasting approach. Suggested examples include LSTM and Prophet, although another appropriate model may be proposed. Students conduct a literature review supporting the alternative method, build and where relevant hyperparameter-tune the model, generate another 24-month forecast, visualise predictions and confidence intervals, and calculate suitable evaluation metrics. This component is intended to demonstrate independent research and the ability to propose an alternative solution rather than relying only on the conventional ARMA approach. assing,,, (1) Part 3 consists of a 6–8 page technical report explaining the modelling process, forecasting results and resulting inferences. The report should provide a critical analysis rather than simply reproducing numerical outputs. Students are expected to explain why results occurred, justify modelling choices, evaluate how those choices influenced performance, compare forecasts with subsequently observed real data where possible, and construct a coherent narrative supported by plots, images, summary statistics and academic literature. Future improvements to the modelling approach should also be critically discussed. assing,,, (1) Submission consists of both the report and working code. The code may be submitted directly or through an accessible Colab or GitHub repository and must reproduce all models, figures and numerical results presented in the report. The assessment allocates 60% of the marks to code and 40% to the report. Within the coding component, modelling and forecasting completion accounts for 40 marks and code quality and annotation for 20 marks. The report is assessed on analysis and inference, methodological justification, comparison of the two modelling approaches, presentation quality, figures and use of appropriate references. assing,,, (1) Key technical expectations include appropriate testing for stationarity, use of methods such as ADF, ACF, PACF and differencing, systematic model selection, forecasting, evaluation and clear comparison between the traditional ARMA model and the chosen alternative approach. Higher-quality work is expected to interpret what the forecasts mean, identify potential improvements and demonstrate sound technical communication rather than merely reporting model outputs. assing,,, (1) Important for the public Reference Library: the brief explicitly states that students must not use generative AI to write the report, and the rubric indicates that AI text-generation use may result in zero marks for the whole assignment. Therefore, the public entry should remain a high-level description of the assessment rather than material intended for direct submission. assing,,, (1)

Read Model Answer →
Machine Learning / Cloud Computing / Artificial Intelligence 4,004 words

Cloud-Based Machine Learning for Financial Fraud Detection

This Machine Learning on Cloud assessment requires students to design, implement and critically evaluate a cloud-oriented machine-learning solution for financial fraud detection. Working as a group, students address a scenario in which a financial-services organisation requires an automated model capable of detecting fraudulent transactions, reducing financial losses and improving customer security. The project combines machine-learning development with critical evaluation of cloud infrastructure, data preparation, model performance and responsible AI considerations. NUL - LD7187 -Assessment Brief … The project begins with a cloud feasibility study comparing at least two major machine-learning platforms such as Microsoft Azure, Amazon Web Services and Google Cloud Platform. Students evaluate factors including performance, scalability, cost, compliance, system integration and vendor lock-in before providing a justified recommendation. The project then moves into Exploratory Data Analysis, where patterns, anomalies and correlations within the fraud dataset are investigated using visualisations such as heatmaps, histograms and boxplots. NUL - LD7187 -Assessment Brief … A substantial part of the assessment focuses on data preprocessing and class imbalance. Students are expected to clean and transform the dataset, apply scaling and encoding, perform feature engineering and investigate techniques such as SMOTE, undersampling and cost-sensitive learning. These decisions must be justified in terms of their potential effect on predictive performance. NUL - LD7187 -Assessment Brief … Students must then select and train at least two machine-learning models. Suggested algorithms include Logistic Regression, Random Forest, XGBoost and Neural Networks. Appropriate cross-validation and hyperparameter-tuning procedures should be applied, followed by systematic evaluation using precision, recall, F1-score, AUC and precision-recall curves. Supporting visualisations should include confusion matrices, ROC curves and feature-importance analysis. NUL - LD7187 -Assessment Brief … The final component addresses professionalism and ethics in cloud-based AI, including bias, fairness, transparency, data privacy and environmental sustainability. Overall, the project integrates cloud-platform selection, exploratory analytics, preprocessing, imbalanced-data handling, predictive modelling, model evaluation and ethical AI into an applied financial fraud-detection solution. NUL - LD7187 -Assessment Brief … Note: the uploaded brief does not explicitly name a referencing system. If your portal requires a selection, I would use Not specified rather than assume Harvard.

Read Model Answer →
Principles of Data Science 2,000 words

Principles of Data Science – Data Analysis Portfolio

This portfolio assignment for the Principles of Data Science module at Coventry University requires students to analyse the Global Life-Work Balance Index 2025 dataset using statistical and data science techniques in R. The dataset ranks 60 countries according to life-work balance using factors including statutory annual leave, paid maternity leave, sick leave, healthcare, public safety, public happiness, LGBTQ inclusivity and average working hours per employee. The assignment has a 2,000-word equivalent limit, excluding the reference list and output. The portfolio consists of two main tasks. Task 1 is a group task involving multivariate data analysis. Students must use R to perform Principal Component Analysis (PCA) and Cluster Analysis on the dataset. For PCA, students analyse quantitative variables, produce and interpret relevant visualisations such as screeplots, biplots and loadings plots, and investigate the effects of Region and Healthcare System. The PCA analysis also requires comparison of the overall dataset with countries from Europe. The cluster analysis component requires students to cluster both countries and variables using different distance metrics and hierarchical clustering methods. Students compare methods such as Manhattan and Euclidean distances and single linkage and Ward’s method, present comparisons in compact tables, and interpret relevant dendrograms. They must then compare the conclusions obtained from PCA and Cluster Analysis, identifying common insights and apparent conflicts and discussing the extent to which the results are explainable rather than simply interpretable. Task 2 is an individual task focusing on Exploratory Data Analysis and Linear Models. Students create a scatter matrix using ggpairs(), investigate strongly correlated variables, and identify quantitative variables that may help predict Region for European and Asian countries. They then develop and critically assess linear regression models for predicting Score, including models based on employment variables and broader quantitative predictors. Model comparison and selection use concepts including AIC, while diagnostic plots are used to identify countries requiring further investigation. The individual task also requires students to use European Life-Work Balance Index 2023 data to make predictions for European countries not included in the 2025 dataset and to construct a Residuals versus Fitted Values plot. Finally, students must combine the conclusions from their individual linear modelling work with the PCA and Cluster Analysis findings to identify specific discoveries about the variables and countries in the dataset. R code, output and relevant plots must be included directly within the reports. The assignment encourages use of the R tidyverse and requires appropriate referencing of sources. The brief specifies APA-style referencing for the individual and group work. It also states that generative AI may be used for inspiration but not for generating answers or analysing the datasets, and any permitted AI use must be acknowledged and documented.

Read Model Answer →
Data Management / Business Analytics 2,500 words

Data and Decision Making (BS776) — Business Report: Two-Source Data Analysis in Python for Evidence-Based Decision-Making

This Level 7 report applies data management theory to a self-selected industry problem and carries it through to a working Python analysis and a defensible business recommendation. The brief is deliberately open on sector — finance, healthcare, transport, cyber security, business intelligence and others are all permitted — but firm on one point: the chosen topic must carry a genuine business implication rather than being a purely technical or clinical analysis. The work therefore begins by framing a specific data-driven decision the organisation needs to make, and returns to that decision at every stage. Two distinct data sources are then identified from approved open repositories and critically evaluated side by side. The evaluation covers the data types each holds, how the data was collected and what bias that introduces, how each is stored and managed, and where the weaknesses lie — proposing concrete data management solutions for the problems identified rather than simply cataloguing them. The analytical core examines, transforms and explores both datasets using univariate and multivariate techniques. All work is carried out in Python within Google Colab, with full screenshots of the code and outputs placed in the appendices and the live Colab link shared for verification. Charts and tables sit in the main body where they support interpretation, each labelled and referenced back to its data source, and each appendix is cited from the narrative so the reader can move between argument and evidence. Data cleaning and transformation steps are shown and justified, not glossed. Findings are reported at length and converted into a clear recommendation covering both the immediate decision and the current and future direction of data management for the business. The limitations section is written honestly — sample coverage, data recency, the assumptions the transformation forced, and what the proposed solution cannot address. Running alongside this, the module's weekly consolidation discussions are evidenced. Five or more critical responses across units two to nine are screenshotted, dated, individually labelled as appendices, and each supported by academic and practice references. Crucially, these are not left sitting in the appendix: they are cited and used within the main body to support the critical discussion, which is where the marks for that component sit. The report follows the prescribed structure — title page, executive summary, contents, introduction, main section with subsections per task, findings, recommendations, limitations, conclusion, Harvard reference list and full appendices — submitted as a single file.

Read Model Answer →
Machine Learning and Deep Learning 2,000 words

Development and Evaluation of Deep Learning Models for Healthcare Classification

This individual technical assessment focuses on the design, development, analysis and evaluation of a deep learning solution for a healthcare-related classification problem. Students select one of two provided scenarios: Polycystic Ovary Syndrome (PCOS) detection using ultrasound images or heartbeat classification using electrocardiogram (ECG) signals. The objective is to develop an appropriate deep learning approach and demonstrate critical understanding of the complete machine learning workflow, from initial data exploration through to model evaluation and reflection. Students may either design and train a deep learning model from scratch or customise and fine-tune an existing pre-trained architecture. The complete work is presented through a single Jupyter Notebook integrating Python code, technical discussion, results and visualisations. The notebook must clearly define the selected healthcare problem, explain its significance, justify methodological and architectural choices, and critically evaluate the resulting solution. The first stage involves exploratory data analysis and preprocessing, including investigation of class distributions, data imbalance and relevant patterns. Students prepare the data through techniques such as normalisation, augmentation, train-validation-test splitting and appropriate handling of class imbalance. This is followed by model design, training, validation and hyperparameter tuning, with the architecture selected according to the characteristics of the data and classification task. Model performance must then be evaluated using appropriate classification measures, including precision, recall, F1-score, ROC curves and area under the curve (AUC). The developed model should also be compared against suitable benchmark approaches, which may include traditional machine learning algorithms or alternative deep learning architectures. This comparison should identify the relative strengths and limitations of the proposed solution. The final component requires clear visual presentation and critical reflection on the complete modelling process, including limitations, challenges and opportunities for improvement. Importantly, grading prioritises methodological rigour, analytical depth and critical evaluation rather than simply achieving the highest predictive accuracy. Overview word count: approximately 330 words. AI restriction: this brief only permits automated AI tools for spelling and grammar checking. It explicitly prohibits tools such as ChatGPT, Gemini or Copilot from authoring assessment text or code; any permitted AI use must also be acknowledged.

Read Model Answer →
Machine Learning / Artificial Intelligence and Data Science 1,000 words

End-to-End Machine Learning Model Development, Tuning and Evaluation

This Level 7 Machine Learning and Intelligent Agents assessment requires students to develop and document an end-to-end machine-learning solution, covering the complete workflow from data preparation through model training, tuning, testing and evaluation. Students select an appropriate dataset or scenario, formulate a research question and determine whether the problem is most appropriately addressed through supervised learning, unsupervised learning or reinforcement learning. Suitable machine-learning techniques must then be implemented to create a model that can be systematically trained and tested. The assignment requires students to follow a structured machine-learning development process and document the complete development journey. The report should explain the selected scenario, data collection or dataset, Exploratory Data Analysis (EDA), rationale for selecting particular machine-learning methods, model training, fine-tuning and evaluation. Model performance must be assessed using appropriate established metrics, with relevant published research used to justify methodological decisions and support the interpretation of results. The technical implementation should demonstrate the ability to identify the performance of machine-learning algorithms, implement machine-learning approaches using one or more object-oriented programming languages, and determine which algorithms are most appropriate for a particular analytical brief. These requirements directly correspond to the module learning outcomes relating to machine-learning performance, implementation and algorithm selection. Students are advised to document their work within a Jupyter Notebook, combining Markdown explanations with executable code. The notebook may be submitted directly or converted to PDF. Alternatively, students may prepare the 1,000-word report in Microsoft Word, provided that the Python code is included within the submitted document. Assessment is divided into three principal areas: Introduction (20 marks), Machine Learning Process (40 marks), and Evaluation of Model Performance (40 marks). Higher-level work is expected to demonstrate strong understanding of machine-learning concepts, a functioning and thoroughly tested implementation, appropriate selection of algorithms and critical evaluation of the developed solution. Overall, the assessment integrates research-question formulation, data exploration, algorithm selection, programming, model optimisation and evidence-based evaluation within a reproducible machine-learning workflow. All academic sources must be presented using Harvard referencing.

Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning 2,500 words

Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science

This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.

Read Model Answer →
Machine Learning / Cloud Computing 3,984 words

Machine Learning on Cloud: Fraud Detection Model Design, Training and Evaluation

This postgraduate group project focuses on the design, development and critical evaluation of a cloud-based machine learning solution for financial fraud detection. The scenario involves a financial services organisation seeking to detect fraudulent transactions in order to reduce financial losses and improve customer security. Students are required to analyse an appropriate dataset, develop a machine learning solution, evaluate its effectiveness and critically consider the suitability of cloud technologies for deployment. The project begins with a cloud feasibility study, requiring critical comparison of at least two major machine learning platforms such as Microsoft Azure, Amazon Web Services and Google Cloud Platform. Evaluation criteria include performance, scalability, cost, compliance, integration and vendor lock-in, followed by a justified recommendation. Students then conduct exploratory data analysis to identify patterns, anomalies and correlations using appropriate visualisations such as heatmaps, histograms and boxplots. A substantial component addresses data preprocessing and class imbalance. Students are expected to clean and transform the data, apply scaling and encoding, perform feature engineering and investigate approaches such as SMOTE, undersampling and cost-sensitive learning. Each preprocessing choice must be justified in terms of its potential impact on model performance. Students must select and train at least two machine learning models, with suggested approaches including Logistic Regression, Random Forest, XGBoost and Neural Networks. Model development incorporates cross-validation and hyperparameter tuning. Evaluation uses fraud-relevant measures including Precision, Recall, F1 score, AUC and precision-recall curves, supported by confusion matrices, ROC curves and feature-importance visualisations. The project concludes with critical consideration of professional and ethical issues in cloud-based AI, including bias, fairness, transparency, data privacy and sustainability. The overall assessment therefore integrates cloud-platform evaluation, machine learning development, imbalanced classification, model evaluation and responsible AI practice. Overview word count: approximately 340 words.

Read Model Answer →
Big Data Analytics / Data Analytics 4,000 words

Big Data Analytics Using Python and Business Intelligence with Tableau

This individual Big Data Analytics assessment requires students to critically analyse data using programming languages, statistical techniques, data visualisation methods and business intelligence software. The assessment combines practical data analytics using Python with interactive business intelligence and dashboard development using Tableau, requiring evidence of technical implementation, research, critical appraisal and justification of the selected analytical approaches. The first section, worth 70%, is based on a dataset containing accidental drug-related deaths recorded in Connecticut between 2012 and 2024. The dataset contains 12,964 observations and 49 variables. Students are required to conduct exploratory data analysis using Python, develop three to four research questions, formulate null and alternative hypotheses, and apply appropriate statistical methods. The analytical process also requires an evaluation of alternative technologies and methodological approaches, supported by relevant research. Students must justify their selected methodology and present a clear workflow diagram. The solution-development stage involves data preprocessing, descriptive statistical analysis, visualisation, answering the research questions and conducting statistical significance testing to determine whether the null hypothesis should be accepted or rejected. Evidence of coding and a link to working code are also required. The final Python component requires evaluation of the findings, consideration of limitations and recommendations for future development using emerging technologies. The second section, worth 30%, focuses on Business Intelligence using Tableau and uses a historical Olympic Games dataset containing 271,116 rows and 15 columns. Students analyse relationships between medals and host cities, athlete age and medal type, season and medal counts, and sex and medal type. The section culminates in an interactive Tableau dashboard containing at least four interconnected sheets, where changes to relevant parameters are reflected across the dashboard. Overall, the assessment develops practical competence in Python-based analytics, statistical reasoning, research-question development, hypothesis testing, data visualisation and interactive business intelligence dashboard design. Overview word count: approximately 350 words. AI note: the brief permits AI for limited support such as grammar, structure, organisation of ideas and suggestions, but states that the main content, analysis and conclusions must remain the student's own work. AI use must also be declared.

Read Model Answer →
Machine Learning / Artificial Intelligence and Data Science 1,011 words

End-to-End Machine Learning Model Development, Testing and Evaluation

This Level 7 Machine Learning and Intelligent Agents assignment requires students to develop and document an end-to-end machine-learning solution, covering the complete process from data preparation through model training, tuning, testing and evaluation. Students independently select a suitable dataset or scenario, formulate an appropriate research question and determine whether the problem should be addressed using supervised learning, unsupervised learning or reinforcement learning. Appropriate machine-learning algorithms must then be implemented to create a model capable of being trained and objectively tested. The assessment encourages the use of a structured machine-learning development methodology. Students are expected to explain the selected scenario and data source, perform suitable data preparation and Exploratory Data Analysis (EDA), and provide a reasoned justification for the machine-learning methods selected. The development process should demonstrate how the chosen algorithms are trained and fine-tuned before their performance is evaluated using established and relevant metrics. Published academic research should be incorporated to justify methodological choices and support the interpretation of results. The technical work is normally documented within a Jupyter Notebook, combining Markdown explanations with executable code cells. Alternatively, the report may be produced in Microsoft Word provided that the Python implementation is included. The assignment therefore assesses both conceptual understanding and practical programming competence. Students must demonstrate an ability to identify the performance of machine-learning algorithms, implement machine-learning techniques using an object-oriented programming language, and evaluate which algorithms are appropriate for a particular analytical brief. Assessment places particular emphasis on three areas: the Introduction, the Machine Learning Process, and the Evaluation of Model Performance. The marking criteria reward strong understanding of machine-learning concepts, a functioning and thoroughly tested implementation, appropriate selection of algorithms, and critical evaluation of the final solution. At the highest achievement level, implementations are expected to work without exception, satisfy the required functionality, demonstrate comprehensive testing and extend beyond the basic requirements. Overall, the assignment combines research-question formulation, data analysis, algorithm selection, machine-learning implementation, model optimisation and evidence-based evaluation within a reproducible technical workflow. All academic sources and supporting material must be presented using Harvard referencing.

Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning 2,500 words

Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science

This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.

Read Model Answer →
Data Science 1,000 words

Data Investigation Pipeline — Exploratory Analysis and Statistical Evaluation of a Chosen Dataset

This assessment simulates the opening stages of a real data investigation. Students choose their own research question and dataset, then build the full pipeline from raw data through preparation, exploration and statistical testing to visualisation — and, where the question supports it, simple modelling or forecasting. Either Python or R is acceptable; the statistical route through R typically expects an explicit hypothesis rather than a purely exploratory question. The work is structured around an established process methodology such as CRISP-DM, and the development journey is documented alongside the code rather than reported after the fact. The usual submission format is a single notebook combining markdown and code cells, so the written report and the analysis sit in one artefact, though a word-processed document containing the code is normally also accepted. The written element is short — around a thousand words — which makes selection the hardest part of the task. It must cover the scenario, the data collection, the exploratory analysis, the reasoning behind the choice of statistical tests, their results, and the visualisations. Students routinely spend that budget describing what they did and leave nothing for why. The mark distribution makes the priority explicit. Framing the problem and the data source carries the smallest share. Preparation and exploratory analysis, and evaluation of the results in context, carry the bulk in roughly equal measure. That final component is where most marks are lost: it asks for an honest assessment of accuracy, limitations and usefulness. A notebook that produces clean output and then claims more than the data supports scores below one that reports a modest result and explains precisely why it is modest. Established metrics should be used for the statistical tests, and published research cited where it informs the background or interprets the findings. Note that assessments of this type increasingly include a live demonstration in which the student explains their own project to verify authorship, so every line of the submitted work needs to be something the student can talk through unprompted. Our support on assessments of this type is guidance-based. Typical areas of help include: explaining how to scope a research question so the analysis fits the word limit, clarifying which statistical test suits which data type and why, reviewing whether a chosen visualisation communicates what it claims, showing how to write an honest limitations section, checking Harvard referencing, and reviewing a student's own draft notebook against the published marking criteria.

Read Model Answer →