Machine Learning on Big Data Using PySpark: Large-Scale Data Analysis and Predictive Modelling

University:
University of East London
Subject:
Big Data Analytics / Machine Learning
Module:
Machine Learning on Big Data
Assignment Type:
MS Technical and scientific writing
Word Count:
3,000 words
Academic Year:
2025/26

Assignment Overview

This group-based Machine Learning on Big Data project requires students to apply machine learning techniques to a large real-world dataset using PySpark DataFrames and Spark machine-learning libraries. Students select a substantial dataset, ideally between approximately 300 MB and 1 GB, from sources such as Kaggle, workplace data or other valid repositories, and develop an end-to-end big-data analytics workflow. CN7030 CRWK 26T1 The project begins with data loading and preprocessing using PySpark. Students are expected to handle missing values, perform data normalisation and feature engineering, identify class imbalance and propose appropriate mitigation strategies. Where text datasets are selected, additional preprocessing may include stemming, lemmatization and TF-IDF representation. CN7030 CRWK 26T1 The modelling stage requires implementation of an appropriate machine-learning approach using PySpark MLlib or Spark ML. The brief expects a multiclass rather than binary classification problem and allows techniques including multiclass classification, ensemble learning, clustering and text mining. Students must justify their model choice and consider model robustness, bias and variance when attempting to improve predictive performance. CN7030 CRWK 26T1 Students then perform hyperparameter tuning using techniques such as grid search or random search and evaluate the resulting model with appropriate measures. Relevant evaluation outputs may include accuracy, F1-score, precision, recall and a confusion matrix. Results should also be visualised or clearly presented and interpreted to identify meaningful patterns and performance characteristics. CN7030 CRWK 26T1 The project additionally requires consideration of Legal, Social, Ethical and Professional (LSEP) issues. Students discuss potential ethical concerns associated with their dataset, including bias and privacy risks, and propose suitable mitigation strategies. The final work is consolidated into a single user-friendly HTML analytics report that clearly presents the group's preprocessing, modelling, optimisation, evaluation and interpretation. CN7030 CRWK 26T1 CN7030 CRWK 26T1 Overview word count: approximately 335 words. If you are also uploading the presentation separately to the Reference Library, that should be a second entry under “Presentations and Academic Posters”, because the presentation forms a distinct 40% component and assesses understanding of Spark, preprocessing, modelling, optimisation, evaluation and responses to examiner questions. CN7030 CRWK 26T1

Megaminds Experience

Megaminds has supported academic requirements in big data analytics / machine learning, machine learning on big data and related disciplines.