Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.
Data Mining / Data Science
Customer Service Escalation Risk Analytics Using Data Mining and Machine Learning
This advanced Data Mining assessment applies the Knowledge Discovery in Databases (KDD) process to a real-world customer-service analytics scenario. Acting as a Data Scientist, students analyse a historical Customer Service Escalation Risk dataset to help an organisation reduce serious and legal customer escalations by identifying early signs of dissatisfaction, service bottlenecks and operational risk. The findings are intended to support business decisions relating to staffing, employee training, customer-journey improvement and escalation prevention. The dataset contains information covering customer demographics, account characteristics, communication channels, issue categories, operational measures such as wait times, transfers and SLA breaches, behavioural indicators including sentiment and response delays, and commercial variables such as monthly fees and contract value. The target variable contains four escalation outcomes: No escalation, Minor escalation, Serious escalation and Legal escalation. The first stage requires data exploration, visualisation and summary, including examination of variable distributions, dataset structure, descriptive characteristics and potential data-quality issues. Students then perform appropriate data cleaning, transformation, feature engineering and preprocessing. Particular attention must be given to variables that could introduce prediction leakage because of their meaning, timing or reliability. The supervised-learning component requires development and tuning of predictive models using suitable techniques such as k-nearest neighbours, Decision Trees, Support Vector Machines, ensembles or neural networks. Models must be evaluated using appropriate multiclass metrics and compared systematically, with interpretation of influential features and model behaviour. The assessment also requires unsupervised learning. After removing the escalation target, students apply and compare clustering approaches such as K-Means and hierarchical clustering. Appropriate preprocessing, encoding, normalisation or dimensionality reduction may be used, with visualisations such as PCA, t-SNE or scatterplots used to explore cluster structure and its relationship with escalation behaviour. Overall, the project assesses the student's ability to independently design a coherent KDD workflow, justify analytical decisions, compare alternative modelling approaches and communicate actionable findings to both technical and executive audiences. Overview word count: approximately 350 words. AI-use note: the brief permits AI tools only to assist with small, specific code snippets. Any AI-generated code must be clearly acknowledged and cited, while the submitted coding, analysis, interpretation and decision-making must remain the student's own work.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.
Read Model Answer →
Technical Evaluation and Professional Reflection on a Cloud-Based Machine Learning Fraud Detection Solution
This individual technical evaluative report forms the reflective and critical component of the Machine Learning on Cloud module. It builds directly upon a group project involving the development of a machine learning solution for financial fraud detection. The report requires each student to critically evaluate the technical decisions made within the group solution while also reflecting on their individual contribution, teamwork experience and professional development. The first component, Technical Evaluation, accounts for 50% of the individual assessment and has a suggested allocation of approximately 1,000 words. Students critically examine the group’s machine learning solution, including the approaches adopted for data preprocessing, model selection and evaluation. They are expected to discuss the suitability of the chosen techniques and metrics, identify limitations and explain how individual technical decisions affected the final performance and outcome of the fraud-detection solution. The second component, Reflection and Professional Issues, also accounts for 50% and has a suggested allocation of approximately 1,000 words. Students reflect critically on their personal contribution to the project and their experience of working within a team. The discussion addresses challenges encountered, lessons learned and how the experience contributed to technical and professional development. Professional and ethical considerations form an important part of the reflection. Relevant issues include algorithmic bias, fairness, sustainability and data privacy, particularly in relation to machine learning applications within financial services and cloud environments. Students are also required to attach their Seminar Activity Tracker as an appendix, providing evidence of weekly participation and knowledge development. The assessment therefore combines technical critique with reflective practice, requiring students to demonstrate that they understand not only how a machine learning solution was developed, but also why specific technical decisions were made, their consequences, the limitations of the resulting system and the wider ethical and professional implications of deploying AI on cloud infrastructure. Overview word count: approximately 315 words. The brief also states that AI may assist with areas such as grammar, structure, organisation of ideas and suggestions, but the main content, analysis and conclusions must be the student's own work, and AI use must be declared.
Read Model Answer →
Machine Learning / Cloud Computing
3,984 words
Machine Learning on Cloud: Fraud Detection Model Design, Training and Evaluation
This postgraduate group project focuses on the design, development and critical evaluation of a cloud-based machine learning solution for financial fraud detection. The scenario involves a financial services organisation seeking to detect fraudulent transactions in order to reduce financial losses and improve customer security. Students are required to analyse an appropriate dataset, develop a machine learning solution, evaluate its effectiveness and critically consider the suitability of cloud technologies for deployment. The project begins with a cloud feasibility study, requiring critical comparison of at least two major machine learning platforms such as Microsoft Azure, Amazon Web Services and Google Cloud Platform. Evaluation criteria include performance, scalability, cost, compliance, integration and vendor lock-in, followed by a justified recommendation. Students then conduct exploratory data analysis to identify patterns, anomalies and correlations using appropriate visualisations such as heatmaps, histograms and boxplots. A substantial component addresses data preprocessing and class imbalance. Students are expected to clean and transform the data, apply scaling and encoding, perform feature engineering and investigate approaches such as SMOTE, undersampling and cost-sensitive learning. Each preprocessing choice must be justified in terms of its potential impact on model performance. Students must select and train at least two machine learning models, with suggested approaches including Logistic Regression, Random Forest, XGBoost and Neural Networks. Model development incorporates cross-validation and hyperparameter tuning. Evaluation uses fraud-relevant measures including Precision, Recall, F1 score, AUC and precision-recall curves, supported by confusion matrices, ROC curves and feature-importance visualisations. The project concludes with critical consideration of professional and ethical issues in cloud-based AI, including bias, fairness, transparency, data privacy and sustainability. The overall assessment therefore integrates cloud-platform evaluation, machine learning development, imbalanced classification, model evaluation and responsible AI practice. Overview word count: approximately 340 words.
Read Model Answer →
Big Data Analytics / Data Analytics
4,000 words
Big Data Analytics Using Python and Business Intelligence with Tableau
This individual Big Data Analytics assessment requires students to critically analyse data using programming languages, statistical techniques, data visualisation methods and business intelligence software. The assessment combines practical data analytics using Python with interactive business intelligence and dashboard development using Tableau, requiring evidence of technical implementation, research, critical appraisal and justification of the selected analytical approaches. The first section, worth 70%, is based on a dataset containing accidental drug-related deaths recorded in Connecticut between 2012 and 2024. The dataset contains 12,964 observations and 49 variables. Students are required to conduct exploratory data analysis using Python, develop three to four research questions, formulate null and alternative hypotheses, and apply appropriate statistical methods. The analytical process also requires an evaluation of alternative technologies and methodological approaches, supported by relevant research. Students must justify their selected methodology and present a clear workflow diagram. The solution-development stage involves data preprocessing, descriptive statistical analysis, visualisation, answering the research questions and conducting statistical significance testing to determine whether the null hypothesis should be accepted or rejected. Evidence of coding and a link to working code are also required. The final Python component requires evaluation of the findings, consideration of limitations and recommendations for future development using emerging technologies. The second section, worth 30%, focuses on Business Intelligence using Tableau and uses a historical Olympic Games dataset containing 271,116 rows and 15 columns. Students analyse relationships between medals and host cities, athlete age and medal type, season and medal counts, and sex and medal type. The section culminates in an interactive Tableau dashboard containing at least four interconnected sheets, where changes to relevant parameters are reflected across the dashboard. Overall, the assessment develops practical competence in Python-based analytics, statistical reasoning, research-question development, hypothesis testing, data visualisation and interactive business intelligence dashboard design. Overview word count: approximately 350 words. AI note: the brief permits AI for limited support such as grammar, structure, organisation of ideas and suggestions, but states that the main content, analysis and conclusions must remain the student's own work. AI use must also be declared.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.
Read Model Answer →
Artificial Intelligence
2,000 words
End-to-End Applied AI Development — Comparative Machine Learning and Neural Network Modelling on a Public Dataset
This assessment runs a complete applied AI development cycle end to end: problem definition, dataset selection, preprocessing, model building, optimisation, evaluation and critical reflection. Students identify a real-world problem themselves, formulate a research question from it, and source a suitable dataset from a recognised public repository such as UCI, Kaggle, Data.gov or OpenML. Dataset choice carries more weight than students expect. It must be genuinely suitable for supervised learning, complex enough to make preprocessing and feature engineering meaningful, and — critically — structured so that a traditional machine learning approach and a deep learning approach can be sensibly compared on it. A dataset too small or too clean makes the neural network component pointless; one too large or too noisy makes the whole pipeline unfinishable within the page limit. The source must be referenced and the choice explicitly justified against the research problem. The modelling requirement is fixed: at least two supervised machine learning models, plus one artificial neural network built in a mainstream deep learning framework, all trained and tested. The comparison between them is the analytical core of the work. Reporting that the neural network scored higher is not an answer; explaining why, in terms of the data's structure and each model's inductive assumptions, is. Marks are distributed across problem framing, the traditional models, the deep learning model, evaluation and critical analysis including responsible AI considerations, and academic communication. That responsible AI component is easy to overlook and is not decorative — it asks what the model's limitations mean for anyone who might rely on it. Presentation requirements are specific. The report is page-limited rather than purely word-limited, and every plot must be described in the text while also being legible enough to communicate on its own — a common failure is dense default library output pasted in without axis labels or scale. The implementation is documented in a notebook combining markdown and code cells so the development process is visible, not just the final result, and submissions typically include the cleaned dataset alongside the code. The strongest submissions treat the notebook and the report as one argument. Weaker ones produce a working notebook and then write a report that describes it, rather than a report that uses it as evidence. Our support on assessments of this type is guidance-based. Typical areas of help include: advising on whether a candidate dataset can actually support the required model comparison, explaining how to justify preprocessing decisions, clarifying which evaluation metrics suit which problem type and why accuracy alone is often misleading, showing how to structure a critical limitations and responsible AI discussion, checking Harvard referencing, and reviewing a student's own draft against the published marking criteria.
Read Model Answer →