Data Mining / Data Science
Customer Service Escalation Risk Analytics Using Data Mining and Machine Learning
This advanced Data Mining assessment applies the Knowledge Discovery in Databases (KDD) process to a real-world customer-service analytics scenario. Acting as a Data Scientist, students analyse a historical Customer Service Escalation Risk dataset to help an organisation reduce serious and legal customer escalations by identifying early signs of dissatisfaction, service bottlenecks and operational risk. The findings are intended to support business decisions relating to staffing, employee training, customer-journey improvement and escalation prevention. The dataset contains information covering customer demographics, account characteristics, communication channels, issue categories, operational measures such as wait times, transfers and SLA breaches, behavioural indicators including sentiment and response delays, and commercial variables such as monthly fees and contract value. The target variable contains four escalation outcomes: No escalation, Minor escalation, Serious escalation and Legal escalation. The first stage requires data exploration, visualisation and summary, including examination of variable distributions, dataset structure, descriptive characteristics and potential data-quality issues. Students then perform appropriate data cleaning, transformation, feature engineering and preprocessing. Particular attention must be given to variables that could introduce prediction leakage because of their meaning, timing or reliability. The supervised-learning component requires development and tuning of predictive models using suitable techniques such as k-nearest neighbours, Decision Trees, Support Vector Machines, ensembles or neural networks. Models must be evaluated using appropriate multiclass metrics and compared systematically, with interpretation of influential features and model behaviour. The assessment also requires unsupervised learning. After removing the escalation target, students apply and compare clustering approaches such as K-Means and hierarchical clustering. Appropriate preprocessing, encoding, normalisation or dimensionality reduction may be used, with visualisations such as PCA, t-SNE or scatterplots used to explore cluster structure and its relationship with escalation behaviour. Overall, the project assesses the student's ability to independently design a coherent KDD workflow, justify analytical decisions, compare alternative modelling approaches and communicate actionable findings to both technical and executive audiences. Overview word count: approximately 350 words. AI-use note: the brief permits AI tools only to assist with small, specific code snippets. Any AI-generated code must be clearly acknowledged and cited, while the submitted coding, analysis, interpretation and decision-making must remain the student's own work.
Read Model Answer →
Machine Learning / Data Mining / Text Mining
Machine Learning Analysis of Classification Models and Text Mining on Furniture Review Data
This technical machine-learning report demonstrates the practical application of predictive modelling and text mining using WEKA. The work is divided into two major tasks. The first evaluates and compares Support Vector Machine and Decision Tree classification models, while the second applies text-mining techniques to furniture-review data and compares multiple classifiers after preprocessing, feature selection and class balancing. The first task uses the Screenshots.arff dataset to investigate the performance of libSVM and J48 Decision Tree classifiers. A 70% training and 30% testing split is applied, and the models are manually tuned to examine how different parameter settings affect predictive performance. For libSVM, an RBF kernel is used while different gamma and cost values are tested through grid-search-style experimentation. The report identifies gamma 0.03 and cost 2 as the strongest tested combination, producing approximately 91.67% accuracy on the test split. The J48 model is also optimised by adjusting the confidence factor used for pruning. Several confidence-factor values are examined, with 0.09 producing the strongest reported result of 80% accuracy. The optimised SVM and J48 models are then compared using five-fold cross-validation, where libSVM achieves 89.75% accuracy compared with 81.25% for J48. The second task focuses on text mining of Furniture Reviews. Text preprocessing includes TF-IDF term weighting, stopword removal, stemming, conversion to lowercase and word-count generation. The resulting textual dataset is transformed into a numerical feature representation suitable for machine-learning classification. Dimensionality reduction is performed using InfoGainAttributeEval with Ranker, selecting the 900 most informative attributes. The dataset is then balanced using WEKA techniques including Resample and SpreadSubsample to reduce class bias before classification. Finally, three classifiers—Naive Bayes, libSVM and J48—are evaluated on the balanced text dataset. The reported accuracies are 90.52% for Naive Bayes, 58.62% for libSVM and 78.45% for J48. The analysis concludes that Naive Bayes performs strongest for the processed furniture-review dataset, while the wider exercise demonstrates the importance of preprocessing, parameter tuning, feature selection, class balancing and appropriate model evaluation in producing reliable classification results. Important: this upload appears to be the completed student report, not the actual assessment guideline. Because the document does not state the university, module name, academic level, academic year, required word count or prescribed referencing style, I would leave those fields as Not specified rather than guessing.
Read Model Answer →