Applying Advanced AI Methods for Analysing Text Documents
This Advanced Artificial Intelligence coursework requires students to implement and evaluate natural language processing, natural language understanding and neural-network techniques for analysing text documents. The assessment uses a supplied social-media dataset containing more than 89,000 posts linked to 947 news headlines, with each post labelled as real or fake through a combination of headline ground truth and majority-vote annotation. Each social-media post is treated as an individual text document for classification and topic-analysis purposes. CMP_6059B_7059B_2025_26_CW1-pre… The first major task focuses on identifying fake text documents. Students must preprocess the text using appropriate NLP techniques, transform documents into numerical feature representations and experiment with alternative preprocessing approaches to determine which performs best. A separate unseen test set must be reserved to evaluate generalisation, while the remaining data is used for training both shallow and deep neural-network classifiers. Students are expected to explain and justify how the data is split. CMP_6059B_7059B_2025_26_CW1-pre… Students first design a multi-layer perceptron (MLP) capable of predicting whether documents are real or fake. The architecture must be justified in terms of input dimensions, number of layers, neuron counts, activation functions and outputs. A second deep-learning neural network must then be developed for the same classification problem, with justification of the chosen network structure, layer types, activations and other configuration decisions. CMP_6059B_7059B_2025_26_CW1-pre… For both networks, students train baseline models and select three hyperparameters considered most important for improving performance. These hyperparameters must be tuned systematically, with visuals prepared to show the experimentation process and resulting performance changes. Appropriate evaluation metrics are then used to compare the trained models. The strongest MLP and deep-learning models must be saved so that they can be loaded and tested on unseen data during the final demonstration without retraining. CMP_6059B_7059B_2025_26_CW1-pre… The second task focuses on topic discovery using NLP and NLU techniques. Students perform syntactic preprocessing such as tokenisation, stop-word removal and lemmatisation or stemming, and experiment with at least two different text-representation approaches. Suggested methods include Bag of Words, TF-IDF, LDA, word vectors and word embeddings. Students must interpret the discovered topics and explain how those topics relate to document content, linked news headlines and class labels. The best topic-discovery model or models must also be saved for live analysis during the demonstration. CMP_6059B_7059B_2025_26_CW1-pre… The assessment is completed through a bench demonstration, supported by a maximum of seven PowerPoint slides. The slides should document the system design, model-improvement process, performance evaluation and discussion of results for both fake-document classification and topic discovery. Students also submit a ZIP file containing only their Python source files. The demonstration lasts up to 15 minutes, consisting of approximately 10 minutes for presentation and technical demonstration followed by 5 minutes for questions and transitions. CMP_6059B_7059B_2025_26_CW1-pre… The marking scheme allocates 45% to fake-document identification, including descriptive analysis, preprocessing, MLP design and deep-learning design; 35% to topic discovery, including preprocessing, model development and interpretation; and 20% to the structure, organisation, professionalism and Q&A quality of the demonstration. CMP_6059B_7059B_2025_26_CW1-pre… Important for the public Reference Library: the brief explicitly states that the use of Large Language Models or generative AI to produce any part of the submission is strictly prohibited, including code, data processing, testing, writing or PowerPoint content. Therefore, this entry should remain only a high-level public description of the assessment and should not be presented as material intended for direct student submission. CMP_6059B_7059B_2025_26_CW1-pre…
Read Model Answer →