Filter by keyword, subject, or both. Updates live as new model answers are added to our portal.
Artificial Intelligence
Developing an Intelligent Chatbot and Expert System for UK Train Services
This postgraduate Advanced Artificial Intelligence group project requires students to design, implement, evaluate and demonstrate an intelligent conversational system for a UK train operating company. The chatbot combines conversational AI, expert-system concepts, predictive modelling and knowledge-based reasoning to support both railway passengers and operational staff. The coursework is worth 70% of the module and is designed to develop practical experience in applying modern AI techniques to realistic service and operational problems. The first task requires the chatbot to interact with passengers, collect journey requirements such as origin, destination, date and travel time, and identify the cheapest available train ticket. Appropriate railway ticket data sources or APIs may be used, with the selected ticket presented together with access to the relevant booking service. The second task extends the system to improve customer service through train-delay prediction. The chatbot gathers information about a passenger's current train, location, delay and destination before passing these data to one or more predictive models. Students process historical railway-performance data, train and compare suitable machine-learning models, evaluate their accuracy and integrate an appropriate model into the chatbot. The third task introduces an expert system for railway contingencies. Students extract rules and operational knowledge from provided contingency and station-disruption documents and construct a knowledge base capable of advising railway staff during events such as partial or complete line blockages. The system should gather details such as event type, location, time and severity, then provide relevant operational guidance, diversion information, alternative services and passenger advice. The overall architecture may include a user interface, NLP/NLU component, knowledge base, reasoning engine, predictive model, database and optional knowledge-acquisition component. Particular emphasis is placed on context-aware dialogue, reliable reasoning, appropriate fallback responses and effective user experience. Assessment outputs include the working chatbot, source code, a live presentation and demonstration, a detailed group technical report, and an individual contribution report. Overview word count: approximately 375 words. AI-use note: pre-trained LLMs may be used as an engine within the system, but they must not be used to generate coursework code. Any use of an LLM within the solution must be clearly justified and explained in the group report.
Read Model Answer →
Data Mining / Data Science
Customer Service Escalation Risk Analytics Using Data Mining and Machine Learning
This advanced Data Mining assessment applies the Knowledge Discovery in Databases (KDD) process to a real-world customer-service analytics scenario. Acting as a Data Scientist, students analyse a historical Customer Service Escalation Risk dataset to help an organisation reduce serious and legal customer escalations by identifying early signs of dissatisfaction, service bottlenecks and operational risk. The findings are intended to support business decisions relating to staffing, employee training, customer-journey improvement and escalation prevention. The dataset contains information covering customer demographics, account characteristics, communication channels, issue categories, operational measures such as wait times, transfers and SLA breaches, behavioural indicators including sentiment and response delays, and commercial variables such as monthly fees and contract value. The target variable contains four escalation outcomes: No escalation, Minor escalation, Serious escalation and Legal escalation. The first stage requires data exploration, visualisation and summary, including examination of variable distributions, dataset structure, descriptive characteristics and potential data-quality issues. Students then perform appropriate data cleaning, transformation, feature engineering and preprocessing. Particular attention must be given to variables that could introduce prediction leakage because of their meaning, timing or reliability. The supervised-learning component requires development and tuning of predictive models using suitable techniques such as k-nearest neighbours, Decision Trees, Support Vector Machines, ensembles or neural networks. Models must be evaluated using appropriate multiclass metrics and compared systematically, with interpretation of influential features and model behaviour. The assessment also requires unsupervised learning. After removing the escalation target, students apply and compare clustering approaches such as K-Means and hierarchical clustering. Appropriate preprocessing, encoding, normalisation or dimensionality reduction may be used, with visualisations such as PCA, t-SNE or scatterplots used to explore cluster structure and its relationship with escalation behaviour. Overall, the project assesses the student's ability to independently design a coherent KDD workflow, justify analytical decisions, compare alternative modelling approaches and communicate actionable findings to both technical and executive audiences. Overview word count: approximately 350 words. AI-use note: the brief permits AI tools only to assist with small, specific code snippets. Any AI-generated code must be clearly acknowledged and cited, while the submitted coding, analysis, interpretation and decision-making must remain the student's own work.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assessment requires students to develop a comprehensive analytical solution to a real-world healthcare prediction problem using the WiDS Datathon 2025 Health Outcomes Prediction Dataset. The dataset contains socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents, with the principal objective of developing predictive models for ADHD diagnosis. The assessment is designed to demonstrate the complete data-science lifecycle, from data preparation and exploratory analysis through predictive modelling, interpretation and evidence-based recommendations. Students begin by exploring the dataset's features, data types and distributions before addressing missing values, outliers and other inconsistencies. Appropriate feature engineering should be undertaken where necessary, followed by Exploratory Data Analysis (EDA) using relevant visualisations to identify relationships, patterns and correlations within the data. Students with limited computational resources may use a representative subset, provided that the sampling method preserves the integrity and distribution of the original dataset and is clearly justified. A major component of the assignment involves developing and comparing at least three classification models. Appropriate techniques may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Model performance should be evaluated using measures including accuracy, precision, recall, F1-score and ROC-AUC, allowing students to identify the strongest-performing model through systematic comparison. The assessment also requires model interpretation and explainability. Students should explain the results of the selected model and may apply techniques such as SHAP or LIME to investigate feature importance and individual predictions. A feature-importance visualisation must be produced, and the most influential variables should inform practical recommendations for healthcare professionals regarding the potential use of predictive modelling in supporting earlier ADHD diagnosis and intervention. Overall, the assignment integrates data cleaning, exploratory analytics, predictive modelling, model comparison, explainable AI and research-informed healthcare recommendations. Students must submit a comprehensive report of no more than 2,500 words, alongside a Jupyter Notebook containing the implementation and outputs. The report must use Harvard referencing, with appropriate academic research integrated into the analysis, recommendations and conclusion.
Read Model Answer →
Machine Learning / Data Mining / Text Mining
Machine Learning Analysis of Classification Models and Text Mining on Furniture Review Data
This technical machine-learning report demonstrates the practical application of predictive modelling and text mining using WEKA. The work is divided into two major tasks. The first evaluates and compares Support Vector Machine and Decision Tree classification models, while the second applies text-mining techniques to furniture-review data and compares multiple classifiers after preprocessing, feature selection and class balancing. The first task uses the Screenshots.arff dataset to investigate the performance of libSVM and J48 Decision Tree classifiers. A 70% training and 30% testing split is applied, and the models are manually tuned to examine how different parameter settings affect predictive performance. For libSVM, an RBF kernel is used while different gamma and cost values are tested through grid-search-style experimentation. The report identifies gamma 0.03 and cost 2 as the strongest tested combination, producing approximately 91.67% accuracy on the test split. The J48 model is also optimised by adjusting the confidence factor used for pruning. Several confidence-factor values are examined, with 0.09 producing the strongest reported result of 80% accuracy. The optimised SVM and J48 models are then compared using five-fold cross-validation, where libSVM achieves 89.75% accuracy compared with 81.25% for J48. The second task focuses on text mining of Furniture Reviews. Text preprocessing includes TF-IDF term weighting, stopword removal, stemming, conversion to lowercase and word-count generation. The resulting textual dataset is transformed into a numerical feature representation suitable for machine-learning classification. Dimensionality reduction is performed using InfoGainAttributeEval with Ranker, selecting the 900 most informative attributes. The dataset is then balanced using WEKA techniques including Resample and SpreadSubsample to reduce class bias before classification. Finally, three classifiers—Naive Bayes, libSVM and J48—are evaluated on the balanced text dataset. The reported accuracies are 90.52% for Naive Bayes, 58.62% for libSVM and 78.45% for J48. The analysis concludes that Naive Bayes performs strongest for the processed furniture-review dataset, while the wider exercise demonstrates the importance of preprocessing, parameter tuning, feature selection, class balancing and appropriate model evaluation in producing reliable classification results. Important: this upload appears to be the completed student report, not the actual assessment guideline. Because the document does not state the university, module name, academic level, academic year, required word count or prescribed referencing style, I would leave those fields as Not specified rather than guessing.
Read Model Answer →
Data Science / Artificial Intelligence and Machine Learning
2,500 words
Predicting ADHD Diagnosis Using Machine Learning and Explainable Data Science
This Data Science assignment focuses on developing a comprehensive analytical solution to a real-world healthcare prediction problem. Using the WiDS Datathon 2025 Health Outcomes Prediction Dataset, students are required to analyse complex and high-dimensional healthcare data containing socio-demographic information, diagnostic variables and functional MRI data relating to children and adolescents. The principal predictive objective is to determine ADHD diagnosis from the available features. Students may use a representative subset of the dataset where computational resources are limited, provided that the sampling approach maintains the integrity and distribution of the original data and is appropriately justified. The assessment requires a complete data-science workflow beginning with data understanding and preprocessing. Students investigate the dataset's features, data types and distributions before addressing missing values, outliers and inconsistencies. Appropriate feature engineering should then be undertaken where it can improve the predictive capability of the models. Exploratory Data Analysis is used to identify important patterns, relationships and correlations, supported by relevant visualisations that communicate meaningful insights. A major component of the work involves the development and comparison of at least three classification models for predicting ADHD diagnosis. Suitable approaches may include Logistic Regression, Random Forest, Gradient Boosting and Neural Networks. Models are evaluated using performance measures including accuracy, precision, recall, F1-score and ROC-AUC, after which the most effective model is selected based on the evidence obtained. The assessment also places substantial emphasis on model interpretation and explainability. Students must interpret the selected model and may use approaches such as SHAP or LIME to explain feature importance and individual predictions. A feature-importance visualisation is required, and the most influential variables should inform practical recommendations. The final section translates analytical findings into recommendations for healthcare professionals, considering how predictive modelling could assist early ADHD diagnosis and intervention. Research literature must be integrated into the recommendations and conclusion. The assessment therefore combines preprocessing, exploratory analysis, predictive modelling, explainable AI and evidence-based healthcare decision-making within a single applied data-science project. The required report is a maximum of 2,500 words, with code, supplementary charts and tables permitted in appendices. A Jupyter Notebook containing the implementation and outputs is also required. Harvard referencing must be used throughout.
Read Model Answer →