Parallel Merge Sort with Load Balancing

University:
Korea University
Subject:
Computer Science / Parallel Computing
Assignment Type:
MS Technical and scientific writing

Assignment Overview

This technical research article investigates the performance limitations of conventional parallel merge sort and proposes a load-balanced alternative designed to improve processor utilisation in distributed-memory parallel computing systems. Traditional parallel merge sort progressively reduces the number of active processors during successive merge stages, causing many processors to remain idle and reducing the performance benefits of parallelisation. The proposed approach addresses this limitation by ensuring that all processors continue participating throughout the merging process. Parallel_Merge_Sort_with_Load_B… The paper first explains the conventional parallel merge-sort process, in which data are locally sorted before processors are paired for a series of merging operations. At every subsequent stage, the number of participating processors is halved until only one processor remains responsible for the final merged list. This results in poor processor utilisation and increasing workload concentration. Parallel_Merge_Sort_with_Load_B… The proposed load-balanced parallel merge sort distributes each partially sorted list across multiple processors so that every processor maintains approximately the same number of keys throughout execution. Processor groups use histograms and boundary values to determine how data should be redistributed during merging. Histogram-based partitioning reduces unnecessary data movement, while an index-swapping mechanism is introduced to avoid transferring large blocks of keys when logical processor reassignment can achieve the same result more efficiently. Parallel_Merge_Sort_with_Load_B… Parallel_Merge_Sort_with_Load_B… The algorithm was implemented in C using MPI and experimentally evaluated on a Cray T3E parallel computer and an eight-node PC cluster. Testing considered both uniform and Gaussian key distributions. Results show that performance improvements increase as processor count grows, although communication and histogram-management overhead can reduce benefits for small workloads. Parallel_Merge_Sort_with_Load_B… The proposed technique achieved a maximum merge-phase speedup of 9.6 on a 32-processor Cray T3E and 2.3 on an eight-node PC cluster when processing four million keys. The study concludes that distributing approximately equal workloads across processors can substantially improve parallel merge performance and may also be applicable to related parallel sorting algorithms. Parallel_Merge_Sort_with_Load_B… Overview word count: approximately 330 words. For your portal, I would not label this as university coursework unless you also have the actual assessment brief that uses this paper. This PDF itself only establishes a published academic paper and the authors’ Korea University affiliation.

Megaminds Experience

Megaminds has supported academic requirements in computer science / parallel computing and related disciplines.