Machine Learning on Big Data using PySpark
This group assessment for the Machine Learning on Big Data module requires students to apply machine learning techniques to a large real-world dataset using PySpark DataFrame and Spark ML. The coursework is designed to develop practical understanding of big data processing, machine learning implementation, model optimisation and evaluation. Students work in groups of three or four and complete a machine learning project using a suitable large dataset obtained from a benchmarking source such as Kaggle, their workplace or another valid resource. The recommended dataset size is between 300 MB and 1 GB. The main task requires students to select a dataset containing multiple classes rather than a binary classification problem and develop an appropriate machine learning solution. The project begins with loading and preprocessing the data using PySpark DataFrame. Students are expected to address issues such as missing values, data normalisation, feature engineering and class imbalance. For text datasets, additional preprocessing techniques such as stemming, lemmatisation and TF-IDF may be applied. Students must then select and implement an appropriate machine learning method using PySpark MLlib or the Spark ML package. The brief permits approaches including multiclass classification, clustering, ensemble learning and text mining. Model parameters must be optimised using techniques such as grid search or random search. Students are required to evaluate their trained models using suitable metrics, including accuracy, F1-score, precision, recall and classification matrices where appropriate. The brief encourages students to maintain model accuracy and robustness above 90% for the highest possible mark and to discuss steps taken to address bias and variance. The results must be visualised or printed clearly, with appropriate interpretation and analysis of the findings. Students must also consider Legal, Social, Ethical and Professional (LSEP) issues throughout the project. Each student selects one LSEP principle and discusses relevant concerns such as dataset bias, privacy or ethical implications, together with suitable mitigation strategies. The final report should be approximately 3,000 words with a tolerance of ±10% and submitted as a single HTML report using the template provided on the module Moodle site. The report should consolidate the individual contributions of all group members into one comprehensive and user-friendly analytics report. The main assessment is weighted 60% for the report and 40% for the group presentation. The presentation is conducted online through Microsoft Teams, and all group members must participate. The assessment evaluates understanding of Spark, preprocessing, modelling, optimisation, evaluation and the ability to explain and interpret the implemented solution.
Read Model Answer →