Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

140

IPPS
2010
IEEE

142views Distributed And Parallel Com...» more IPPS 2010»

Exploiting inter-thread temporal locality for chip multithreading

15 years 2 months ago

Exploiting inter-thread temporal locality for chip multithreading

Download www.cs.virginia.edu

Multi-core organizations increasingly support multiple threads per core. Threads on a core usually share a single first-level data cache, so thread schedulers must try to minimize cache contention among threads. While this has been studied for concurrent threads with disjoint working sets, the problem has not been addressed for multi-threaded data-parallel workloads in which threads can be scheduled or constructed to improve inter-thread cache sharing. This paper proposes the symbiotic affinity scheduling (SAS) algorithm in which work is first partitioned according to the number of cores (i.e., the number of caches), and these partitions are then subdivided and scheduled among each core's available thread contexts so that threads sharing a core operate on neighboring elements to maximize cache locality. We demonstrate this concept with a series of data-parallel benchmarks. Simulations on M5 achieve an average speedup

Jiayuan Meng, Jeremy W. Sheaffer, Kevin Skadron

Real-time Traffic

Cache | Distributed And Parallel Computing | First-level Data Cache | IPPS 2010 | Threads |

claim paper

Related Content

» NoCaware cache design for multithreaded execution on tiled chip multiprocessors

» Exploiting processing locality through paging configurations in multitasked reconfigurable...

» An Evaluation of Thread Migration for Exploiting Distributed Array Locality

» Communication optimizations for global multithreaded instruction scheduling

» Parallelization performance analysis and algorithm consideration of Hough transform on chi...

» ThreeDimensional ChipMultiprocessor RunTime Thermal Management

» PseudoCircuit Accelerating Communication for OnChip Interconnection Networks

» Fast PerformanceOptimized Partial Match Address Compression for LowLatency OnChip Address ...

Post Info
More Details (n/a)

Added	13 Feb 2011
Updated	13 Feb 2011
Type	Journal
Year	2010
Where	IPPS
Authors	Jiayuan Meng, Jeremy W. Sheaffer, Kevin Skadron

Comments (0)