Team Research and Development
Team Research and Development – Research Question, Hypothesis and Statistical Data Analysis
This group assessment for 7COM1079 – Team Research and Development requires students to collaboratively select a dataset and conduct a structured research investigation using data analysis. The assessment is worth 40% of the module mark and requires group members to work together to develop and submit a final research report. All group members are expected to contribute equally to the work, with one designated group member submitting the final assessment on behalf of the group. The main objective of the assignment is to develop a meaningful research question and corresponding hypothesis based on a selected dataset. The research question defines a specific claim or issue that the group intends to investigate and provides the starting point for constructing and testing an evidence-based argument. The hypothesis provides a proposed explanation or prediction that can be examined through analysis of the available data. Together, the research question and hypothesis establish what evidence is required, how that evidence will be evaluated and how effectively the findings can support or challenge a particular position. The assignment permits three types of research questions. The first involves establishing a difference in means between two groups. The second involves establishing a difference in proportion between two groups. The third involves establishing a correlation between two measures. The selected research question should therefore be appropriate for the characteristics of the chosen dataset and should allow the group to perform meaningful statistical analysis. Following the selection of the dataset and formulation of the research question and hypothesis, students are expected to produce appropriate data visualisations and statistical analysis. The analysis should provide evidence relevant to the research question and allow the proposed hypothesis to be examined systematically. The resulting findings should be interpreted in relation to the original research question and hypothesis rather than simply presenting numerical results. The assignment provides a Microsoft Word final report template containing the required table of contents, chapter and subchapter names and explanations of the expected content. Students are instructed to download and use this template when preparing their final report. Work produced by individual students during their first assignment may be incorporated where the student examined the same dataset allocated to the group, provided the material is appropriately incorporated into the group submission. All italicised instructional text in the template must be removed before submission. The completed amended template and the group's dataset file must be submitted through Canvas. Acceptable submission formats include PDF, DOC, DOCX, CSV and XLS. The assignment has a late-submission penalty, and the submission deadline is stated as being available in the assignment specification on Canvas. Assessment is based on the criteria provided in the module rubric. The rubric is used to assess the group's work, with group members initially receiving the same mark, although peer review may be taken into account. The assignment therefore combines collaborative research, statistical reasoning, data visualisation, analytical interpretation and academic report writing. The assignment instructions explicitly state that students should not use AI for this assessment. The module team also reserves the right to arrange a viva if academic misconduct is suspected. The final report should therefore represent the group's own research, analysis, interpretation and contribution.
Read Model Answer →
Data Mining
Data Mining – Classification, Model Optimisation and Evaluation
This individual Data Mining assignment requires students to apply the techniques covered in the module using the WEKA data mining platform. The assessment is worth 40% and focuses on practical application of machine learning and data mining methods, requiring students to configure algorithms, analyse datasets, optimise model parameters, evaluate classification performance and explain their technical choices and results. The assignment also assesses the ability to critically evaluate different algorithms and models of data mining. The assignment includes several tasks covering different stages of the data mining process. Students are required to work with supplied datasets and use appropriate preprocessing and classification techniques. The datasets include a balanced screenshot dataset containing processed screenshots classified into categories such as "Okay" and "Bad", where the objective is to train a model capable of identifying inappropriate content. The data has been processed using PCA to provide a smaller four-dimensional representation while protecting privacy and reducing the size of the data. Another supplied dataset concerns furniture reviews, containing positive ("pos") and negative ("neg") written feedback, with the objective of building a model that can determine whether a furniture review belongs to either class. The assessment evaluates students' ability to understand and describe the datasets, including the number of instances, number of columns, data types and relevant statistical information. For text-based data, students must apply appropriate vectorisation and describe the resulting dataset characteristics. Students must also consider class imbalance and apply an appropriate method where necessary, explaining how the chosen approach affects the distribution of instances between the classes. A significant component of the assignment involves classification algorithms and parameter optimisation. The assessment requires students to work with algorithms including Naive Bayes, LibSVM and J48. Students must investigate appropriate parameters and perform parameter searches or fine grid searches to identify suitable configurations. They must explain the selected parameters, their impact on the model and the reasoning behind the chosen values. Model performance must be evaluated using appropriate validation techniques, including cross-validation. Students are required to compare the algorithms using results such as overall accuracy and confusion matrices. The assignment expects students to identify an appropriate or best-performing algorithm in the context of the dataset and to provide a clear explanation of the comparison rather than simply reporting numerical results. The rubric places emphasis on accurate configuration, clear explanation of parameter choices, dataset analysis, class-balance treatment, parameter optimisation, cross-validation and critical comparison of algorithm strengths and weaknesses. High-quality work should explain both the technical process and the implications of the results, with results presented clearly through appropriate tables, confusion matrices and graphical outputs where required. The submission must be a single PDF document containing the report and must not exceed 10 pages. Students are instructed to include their student ID at the beginning of the report but not their name or other identifying details so that marking remains anonymous. Screenshots are specifically required to demonstrate use of the student's ID number as the random seed; other WEKA results should be presented in the student's own tables or result formats. The brief also states that no research beyond the material covered in the module is required and therefore no citations or reference list are required. The assignment explicitly prohibits the use of Generative AI tools for creating content and prohibits using GenAI tools or proofreading services for proofreading. Students are expected to complete the practical work themselves and explain their own technical choices and results.
Read Model Answer →