Data Management / Business Analytics
2,500 words
Data and Decision Making (BS776) — Business Report: Two-Source Data Analysis in Python for Evidence-Based Decision-Making
This Level 7 report applies data management theory to a self-selected industry problem and carries it through to a working Python analysis and a defensible business recommendation. The brief is deliberately open on sector — finance, healthcare, transport, cyber security, business intelligence and others are all permitted — but firm on one point: the chosen topic must carry a genuine business implication rather than being a purely technical or clinical analysis. The work therefore begins by framing a specific data-driven decision the organisation needs to make, and returns to that decision at every stage. Two distinct data sources are then identified from approved open repositories and critically evaluated side by side. The evaluation covers the data types each holds, how the data was collected and what bias that introduces, how each is stored and managed, and where the weaknesses lie — proposing concrete data management solutions for the problems identified rather than simply cataloguing them. The analytical core examines, transforms and explores both datasets using univariate and multivariate techniques. All work is carried out in Python within Google Colab, with full screenshots of the code and outputs placed in the appendices and the live Colab link shared for verification. Charts and tables sit in the main body where they support interpretation, each labelled and referenced back to its data source, and each appendix is cited from the narrative so the reader can move between argument and evidence. Data cleaning and transformation steps are shown and justified, not glossed. Findings are reported at length and converted into a clear recommendation covering both the immediate decision and the current and future direction of data management for the business. The limitations section is written honestly — sample coverage, data recency, the assumptions the transformation forced, and what the proposed solution cannot address. Running alongside this, the module's weekly consolidation discussions are evidenced. Five or more critical responses across units two to nine are screenshotted, dated, individually labelled as appendices, and each supported by academic and practice references. Crucially, these are not left sitting in the appendix: they are cited and used within the main body to support the critical discussion, which is where the marks for that component sit. The report follows the prescribed structure — title page, executive summary, contents, introduction, main section with subsections per task, findings, recommendations, limitations, conclusion, Harvard reference list and full appendices — submitted as a single file.
Read Model Answer →
Machine Learning / Data Science
Linear Regression and Stability of the Moore–Penrose Pseudoinverse Using Python
This machine learning practical and assessment activity develops an understanding of linear regression, Ordinary Least Squares and the Moore–Penrose pseudoinverse using Python. The work progresses from generating synthetic regression datasets to implementing regression algorithms manually, applying established machine learning libraries, analysing real datasets and evaluating the stability of estimated regression coefficients. The laboratory component begins with the generation of synthetic linear regression data using NumPy, including explanatory variables, random noise and an outcome variable. Students then implement simple linear regression without relying on machine learning libraries, using the least-squares solution to estimate the intercept and slope. The resulting observations and fitted regression line are visualised using Matplotlib. The work is subsequently extended to multiple linear regression, where several independent variables are used and coefficients are first calculated manually before the same problem is solved using Scikit-learn. The laboratory also introduces application of regression to the Scikit-learn Diabetes dataset, including feature and target standardisation, model fitting, prediction, correlation analysis and interpretation of regression coefficients. It also highlights the importance of residual analysis when assessing whether a linear model is appropriate. The associated weekly challenge focuses on the stability of linear regression solutions estimated using the Moore–Penrose pseudoinverse. Using a house-price dataset containing variables such as property size, number of bedrooms, distance from the city centre and property age, students construct the design matrix, standardise features and the response variable, and calculate regression coefficients using the pseudoinverse. Students then investigate model robustness by repeatedly fitting the regression model to random subsamples of different sizes and analysing the mean and standard deviation of each coefficient. Tables, boxplots or error-bar visualisations can be used to compare coefficient variability. The final discussion considers which variables are most influential, which coefficients are most stable, how sample size affects stability and whether coefficient interpretation remains reliable across different samples. The final work is submitted as a single PDF exported from Jupyter Notebook or Google Colab, combining documented Python code, experimental results, plots and written interpretation in a professionally organised notebook. Overview word count: approximately 360 wor
Read Model Answer →
Assignment 2 - Individual project Image segmentation
Assignment tasks This assignment will focus on Image Segmentation using the ADE20K dataset. This is an individual assignment where each student will produce a report on the data analysis they will perform. You are encouraged to utilise Google Colab for the coding part of your assignment. https://herts.instructure.com/courses/129101/assignments/406384 1/86/17/26, 12:10 PM Assignment 2 - Individual project - Image segmentation - 25% You will explain and discuss the data processing, the method(s) you make use of and elaborate the outcome. You will work on the ADE20K dataset (explained below in more detail) to research viable models to train, discuss different approaches to explore and visualise the data (i.e., perform EDA), build a tool to pre-process the dataset, and customise your chosen model(s) to improve performance. You will produce a code that does semantic segmentation of the 4 classes targeted in this assignment: person, car, book, airplane. In more detail, your model(s) should identify which of these 4 classes the region of the image corresponds to, and should be applicable to any unlabelled image. To be clear: doing only binary segmentation (i.e. any class vs background) will result in a very large penalty, as you will be considered not to have done the required task. You may use more than one model, but one has to be trained partially or fully by you. Should you use more than one, you are encouraged to compare your main trained model with one or more pre-trained models.
Read Model Answer →