Machine Learning and Big Data
2,000 words
Machine Learning for Big Data
This assignment for the Machine Learning and Big Data module requires students to produce a 2,000-word individual report demonstrating their understanding and practical application of machine learning techniques to big data. The assessment is worth 15 credits and is structured around five interconnected areas: data in big data, machine learning architecture, model deployment, model evaluation, and the machine learning lifecycle. The first part focuses on identifying and evaluating suitable datasets for a selected big data topic and determining whether the datasets are appropriate for the intended machine learning application. Students are expected to examine the characteristics of their data and apply appropriate pre-processing approaches, including consideration of attribute selection and data preparation. The second part addresses machine learning modelling architecture. Students must develop an appropriate architecture for their big data application and may compare alternative architectural approaches. The report should explain the selected machine learning techniques and demonstrate how they operate as part of the proposed system. Practical considerations such as performance, scalability, fault tolerance, technology usage and reliability should also be considered. The third part requires students to implement and deploy the proposed data and machine learning model. This includes testing, visualising and evaluating the resulting outcomes. The fourth part requires critical evaluation of the dataset selection, modelling design, implementation and application, including assessment of whether the selected machine learning techniques are appropriate for the intended purpose. The final part focuses on the complete project lifecycle. Students are expected to critically reflect on the work undertaken, identify what they have learned, evaluate the development process and explain how the machine learning application could be improved in a future implementation. The assessment develops five learning outcomes covering big data sources and applications, machine learning techniques, practical application of machine learning tools, critical evaluation of techniques and tools, and the ability to follow a complete big data analysis lifecycle. The marking criteria allocate 20% to each of these five areas. The assignment is submitted as an individual written report. The brief states that Microsoft Word should be used rather than PDF and requires students to acknowledge sources and any AI tools used in accordance with the stated AI policy.
Read Model Answer →
Artificial Intelligence
2,000 words
End-to-End Applied AI Development — Comparative Machine Learning and Neural Network Modelling on a Public Dataset
This assessment runs a complete applied AI development cycle end to end: problem definition, dataset selection, preprocessing, model building, optimisation, evaluation and critical reflection. Students identify a real-world problem themselves, formulate a research question from it, and source a suitable dataset from a recognised public repository such as UCI, Kaggle, Data.gov or OpenML. Dataset choice carries more weight than students expect. It must be genuinely suitable for supervised learning, complex enough to make preprocessing and feature engineering meaningful, and — critically — structured so that a traditional machine learning approach and a deep learning approach can be sensibly compared on it. A dataset too small or too clean makes the neural network component pointless; one too large or too noisy makes the whole pipeline unfinishable within the page limit. The source must be referenced and the choice explicitly justified against the research problem. The modelling requirement is fixed: at least two supervised machine learning models, plus one artificial neural network built in a mainstream deep learning framework, all trained and tested. The comparison between them is the analytical core of the work. Reporting that the neural network scored higher is not an answer; explaining why, in terms of the data's structure and each model's inductive assumptions, is. Marks are distributed across problem framing, the traditional models, the deep learning model, evaluation and critical analysis including responsible AI considerations, and academic communication. That responsible AI component is easy to overlook and is not decorative — it asks what the model's limitations mean for anyone who might rely on it. Presentation requirements are specific. The report is page-limited rather than purely word-limited, and every plot must be described in the text while also being legible enough to communicate on its own — a common failure is dense default library output pasted in without axis labels or scale. The implementation is documented in a notebook combining markdown and code cells so the development process is visible, not just the final result, and submissions typically include the cleaned dataset alongside the code. The strongest submissions treat the notebook and the report as one argument. Weaker ones produce a working notebook and then write a report that describes it, rather than a report that uses it as evidence. Our support on assessments of this type is guidance-based. Typical areas of help include: advising on whether a candidate dataset can actually support the required model comparison, explaining how to justify preprocessing decisions, clarifying which evaluation metrics suit which problem type and why accuracy alone is often misleading, showing how to structure a critical limitations and responsible AI discussion, checking Harvard referencing, and reviewing a student's own draft against the published marking criteria.
Read Model Answer →