Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.
Data Mining / Data Science
Customer Service Escalation Risk Analytics Using Data Mining and Machine Learning
This advanced Data Mining assessment applies the Knowledge Discovery in Databases (KDD) process to a real-world customer-service analytics scenario. Acting as a Data Scientist, students analyse a historical Customer Service Escalation Risk dataset to help an organisation reduce serious and legal customer escalations by identifying early signs of dissatisfaction, service bottlenecks and operational risk. The findings are intended to support business decisions relating to staffing, employee training, customer-journey improvement and escalation prevention. The dataset contains information covering customer demographics, account characteristics, communication channels, issue categories, operational measures such as wait times, transfers and SLA breaches, behavioural indicators including sentiment and response delays, and commercial variables such as monthly fees and contract value. The target variable contains four escalation outcomes: No escalation, Minor escalation, Serious escalation and Legal escalation. The first stage requires data exploration, visualisation and summary, including examination of variable distributions, dataset structure, descriptive characteristics and potential data-quality issues. Students then perform appropriate data cleaning, transformation, feature engineering and preprocessing. Particular attention must be given to variables that could introduce prediction leakage because of their meaning, timing or reliability. The supervised-learning component requires development and tuning of predictive models using suitable techniques such as k-nearest neighbours, Decision Trees, Support Vector Machines, ensembles or neural networks. Models must be evaluated using appropriate multiclass metrics and compared systematically, with interpretation of influential features and model behaviour. The assessment also requires unsupervised learning. After removing the escalation target, students apply and compare clustering approaches such as K-Means and hierarchical clustering. Appropriate preprocessing, encoding, normalisation or dimensionality reduction may be used, with visualisations such as PCA, t-SNE or scatterplots used to explore cluster structure and its relationship with escalation behaviour. Overall, the project assesses the student's ability to independently design a coherent KDD workflow, justify analytical decisions, compare alternative modelling approaches and communicate actionable findings to both technical and executive audiences. Overview word count: approximately 350 words. AI-use note: the brief permits AI tools only to assist with small, specific code snippets. Any AI-generated code must be clearly acknowledged and cited, while the submitted coding, analysis, interpretation and decision-making must remain the student's own work.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.
Read Model Answer →
Big Data Analytics / Data Analytics
4,000 words
Big Data Analytics Using Python and Business Intelligence with Tableau
This individual Big Data Analytics assessment requires students to critically analyse data using programming languages, statistical techniques, data visualisation methods and business intelligence software. The assessment combines practical data analytics using Python with interactive business intelligence and dashboard development using Tableau, requiring evidence of technical implementation, research, critical appraisal and justification of the selected analytical approaches. The first section, worth 70%, is based on a dataset containing accidental drug-related deaths recorded in Connecticut between 2012 and 2024. The dataset contains 12,964 observations and 49 variables. Students are required to conduct exploratory data analysis using Python, develop three to four research questions, formulate null and alternative hypotheses, and apply appropriate statistical methods. The analytical process also requires an evaluation of alternative technologies and methodological approaches, supported by relevant research. Students must justify their selected methodology and present a clear workflow diagram. The solution-development stage involves data preprocessing, descriptive statistical analysis, visualisation, answering the research questions and conducting statistical significance testing to determine whether the null hypothesis should be accepted or rejected. Evidence of coding and a link to working code are also required. The final Python component requires evaluation of the findings, consideration of limitations and recommendations for future development using emerging technologies. The second section, worth 30%, focuses on Business Intelligence using Tableau and uses a historical Olympic Games dataset containing 271,116 rows and 15 columns. Students analyse relationships between medals and host cities, athlete age and medal type, season and medal counts, and sex and medal type. The section culminates in an interactive Tableau dashboard containing at least four interconnected sheets, where changes to relevant parameters are reflected across the dashboard. Overall, the assessment develops practical competence in Python-based analytics, statistical reasoning, research-question development, hypothesis testing, data visualisation and interactive business intelligence dashboard design. Overview word count: approximately 350 words. AI note: the brief permits AI for limited support such as grammar, structure, organisation of ideas and suggestions, but states that the main content, analysis and conclusions must remain the student's own work. AI use must also be declared.
Read Model Answer →
Operations and Supply Chain Management
1,987 words
Supply Chain Analytics and Quantitative Data Analysis for Organisational Decision-Making
This postgraduate individual report focuses on the application of quantitative data analysis to Operations, Logistics and Supply Chain Management decision-making. Students are required to select an organisation from the private, public or third sector and investigate a relevant operational or supply-chain issue using quantitative data. The purpose is to demonstrate how data can be collected, prepared, analysed and interpreted to generate evidence-based insights that may support managerial decision-making. The selected dataset must relate to the organisation's operations or supply chain and may include variables such as revenues, product orders, sales, transportation costs, procurement expenditure, inventory levels or other appropriate quantitative measures. The dataset must contain at least 60 observations, and the analysis must involve at least two variables. Data may be obtained directly from organisations or from recognised secondary-data platforms and databases. The assessment consists of two equally weighted components. Part A – Motivation and Justification for the Analysis requires students to formulate relevant analytical questions and explain their practical importance by linking them to Operations and Supply Chain Management theory and business practice. Appropriate academic, industry and practitioner evidence should be used to justify the selected issue. Research questions may also be translated into testable hypotheses where appropriate. Part B – Execution of the Analysis requires students to answer the identified questions through appropriate statistical techniques. Potential methods include tables, charts, summary statistics, t-tests and regression analysis. Data may first need to be cleaned, transformed and structured before analysis. The results must then be interpreted clearly for a managerial audience such as the organisation's board, owner or CEO. The statistical analysis is expected to be conducted using Stata, with all data-cleaning, manipulation and analytical commands recorded in a reproducible do-file. The report must also demonstrate explicit links between theory and practice and contain a suitable mixture of academic and professional evidence, including at least five academic journal articles. Harvard referencing is required throughout. The resulting work demonstrates practical competence in business analytics, statistical interpretation, supply-chain decision support, reproducible analysis and evidence-based managerial communication. Overview word count: approximately 370 words.
Read Model Answer →
Data Science
1,000 words
Data Investigation Pipeline — Exploratory Analysis and Statistical Evaluation of a Chosen Dataset
This assessment simulates the opening stages of a real data investigation. Students choose their own research question and dataset, then build the full pipeline from raw data through preparation, exploration and statistical testing to visualisation — and, where the question supports it, simple modelling or forecasting. Either Python or R is acceptable; the statistical route through R typically expects an explicit hypothesis rather than a purely exploratory question. The work is structured around an established process methodology such as CRISP-DM, and the development journey is documented alongside the code rather than reported after the fact. The usual submission format is a single notebook combining markdown and code cells, so the written report and the analysis sit in one artefact, though a word-processed document containing the code is normally also accepted. The written element is short — around a thousand words — which makes selection the hardest part of the task. It must cover the scenario, the data collection, the exploratory analysis, the reasoning behind the choice of statistical tests, their results, and the visualisations. Students routinely spend that budget describing what they did and leave nothing for why. The mark distribution makes the priority explicit. Framing the problem and the data source carries the smallest share. Preparation and exploratory analysis, and evaluation of the results in context, carry the bulk in roughly equal measure. That final component is where most marks are lost: it asks for an honest assessment of accuracy, limitations and usefulness. A notebook that produces clean output and then claims more than the data supports scores below one that reports a modest result and explains precisely why it is modest. Established metrics should be used for the statistical tests, and published research cited where it informs the background or interprets the findings. Note that assessments of this type increasingly include a live demonstration in which the student explains their own project to verify authorship, so every line of the submitted work needs to be something the student can talk through unprompted. Our support on assessments of this type is guidance-based. Typical areas of help include: explaining how to scope a research question so the analysis fits the word limit, clarifying which statistical test suits which data type and why, reviewing whether a chosen visualisation communicates what it claims, showing how to write an honest limitations section, checking Harvard referencing, and reviewing a student's own draft notebook against the published marking criteria.
Read Model Answer →