Academic literature on the topic 'Bank Level Parallelism (BLP)'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the lists of relevant articles, books, theses, conference reports, and other scholarly sources on the topic 'Bank Level Parallelism (BLP).'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Journal articles on the topic "Bank Level Parallelism (BLP)"

1

Shin, Wongyu, Jaemin Jang, Jungwhan Choi, Jinwoong Suh, and Lee-Sup Kim. "Bank-Group Level Parallelism." IEEE Transactions on Computers 66, no. 8 (2017): 1428–34. http://dx.doi.org/10.1109/tc.2017.2665475.

Full text
APA, Harvard, Vancouver, ISO, and other styles
2

Xue, Dongliang, Linpeng Huang, and Chentao Wu. "A pure hardware-driven scheduler for enhancing bank-level parallelism in a persistent memory controller." Future Generation Computer Systems 107 (June 2020): 383–93. http://dx.doi.org/10.1016/j.future.2020.01.047.

Full text
APA, Harvard, Vancouver, ISO, and other styles
3

Najoui, Mohamed, Mounir Bahtat, Anas Hatim, Said Belkouch, and Noureddine Chabini. "VLIW DSP-Based Low-Level Instruction Scheme of Givens QR Decomposition for Real-Time Processing." Journal of Circuits, Systems and Computers 26, no. 09 (2017): 1750129. http://dx.doi.org/10.1142/s0218126617501298.

Full text
Abstract:
QR decomposition (QRD) is one of the most widely used numerical linear algebra (NLA) kernels in several signal processing applications. Its implementation has a considerable and an important impact on the system performance. As processor architectures continue to gain ground in the high-performance computing world, QRD algorithms have to be redesigned in order to take advantage of the architectural features on these new processors. However, in some processor architectures like very large instruction word (VLIW), compiler efficiency is not enough to make an effective use of available computatio
APA, Harvard, Vancouver, ISO, and other styles
4

Khadirsharbiyani, Soheil, Jagadish Kotra, Karthik Rao, and Mahmut Taylan Kandemir. "Data Convection." ACM SIGMETRICS Performance Evaluation Review 50, no. 1 (2022): 37–38. http://dx.doi.org/10.1145/3547353.3522647.

Full text
Abstract:
Stacked DRAMs have been studied and productized in the last decade. The large available bandwidth they offer makes them an attractive choice, particularly, in high-performance computing (HPC) environments. Consequently, many prior research efforts have studied and evaluated 3D stacked DRAM-based designs. Despite offering high bandwidth, stacked DRAMs are severely constrained by the overall memory capacity offered. In this paper, we study and evaluate integrating stacked DRAM on top of a GPU in a 3D manner which in tandem with the 2.5D stacked DRAM boosts the capacity and the bandwidth without
APA, Harvard, Vancouver, ISO, and other styles
5

Ma, Jianliang, Jinglei Meng, Tianzhou Chen, and Minghui Wu. "CaLRS: A Critical-Aware Shared LLC Request Scheduling Algorithm on GPGPU." Scientific World Journal 2015 (2015): 1–10. http://dx.doi.org/10.1155/2015/848416.

Full text
Abstract:
Ultra high thread-level parallelism in modern GPUs usually introduces numerous memory requests simultaneously. So there are always plenty of memory requests waiting at each bank of the shared LLC (L2 in this paper) and global memory. For global memory, various schedulers have already been developed to adjust the request sequence. But we find few work has ever focused on the service sequence on the shared LLC. We measured that a big number of GPU applications always queue at LLC bank for services, which provide opportunity to optimize the service order on LLC. Through adjusting the GPU memory r
APA, Harvard, Vancouver, ISO, and other styles
6

Matei, Radu, and Doru Florin Chiper. "Design and Polyphase Implementation of Rotationally Invariant 2D FIR Filter Banks Based on Maximally Flat Prototype." Electronics 13, no. 14 (2024): 2829. http://dx.doi.org/10.3390/electronics13142829.

Full text
Abstract:
This paper presents a design approach for a class of rotationally invariant 2D filters of finite impulse response (FIR) type, which may form circular filter banks with imposed specifications. The design is conducted analytically in the frequency domain and starts from a maximally flat low-pass prototype based on a trapezoidal function with specified width and slope. Its trigonometric approximation is derived using the Fourier series expressed analytically, truncated to a number of terms depending on the imposed accuracy. The chosen trapezoidal function leads to significantly smaller ringing os
APA, Harvard, Vancouver, ISO, and other styles
7

GRÉWAL, G., S. COROS, and M. VENTRESCA. "A MEMETIC ALGORITHM FOR PERFORMING MEMORY ASSIGNMENT IN DUAL-BANK DSPS." International Journal of Computational Intelligence and Applications 06, no. 04 (2006): 473–97. http://dx.doi.org/10.1142/s1469026806002039.

Full text
Abstract:
To increase memory bandwidth, many programmable Digital-Signal Processors (DSPs) employ two on-chip data memories. This architectural feature supports higher memory bandwidth by allowing multiple data memory accesses to occur in parallel. Exploiting dual memory banks, however, is a challenging problem for compilers. This, in part, is due to the instruction-level parallelism, small numbers of registers, and highly specialized register capabilities of most DSPs. In this paper, we present a new methodology based on a Memetic Algorithm (MA) for assigning data to dual-bank memories. Our approach is
APA, Harvard, Vancouver, ISO, and other styles
8

Fang, Juan, Jiajia Lu, Mengxuan Wang, and Hui Zhao. "A Performance Conserving Approach for Reducing Memory Power Consumption in Multi-Core Systems." Journal of Circuits, Systems and Computers 28, no. 07 (2019): 1950113. http://dx.doi.org/10.1142/s0218126619501135.

Full text
Abstract:
With more cores integrated into a single chip and the fast growth of main memory capacity, the DRAM memory design faces ever increasing challenges. Previous studies have shown that DRAM can consume up to 40% of the system power, which makes DRAM a major factor constraining the whole system’s growth in performance. Moreover, memory accesses from different applications are usually interleaved and interfere with each other, which further exacerbates the situation in memory system management. Therefore, reducing memory power consumption has become an urgent problem to be solved in both academia an
APA, Harvard, Vancouver, ISO, and other styles
9

Fang, Juan, Mengxuan Wang, and Zelin Wei. "A memory scheduling strategy for eliminating memory access interference in heterogeneous system." Journal of Supercomputing 76, no. 4 (2020): 3129–54. http://dx.doi.org/10.1007/s11227-019-03135-7.

Full text
Abstract:
AbstractMultiple CPUs and GPUs are integrated on the same chip to share memory, and access requests between cores are interfering with each other. Memory requests from the GPU seriously interfere with the CPU memory access performance. Requests between multiple CPUs are intertwined when accessing memory, and its performance is greatly affected. The difference in access latency between GPU cores increases the average latency of memory accesses. In order to solve the problems encountered in the shared memory of heterogeneous multi-core systems, we propose a step-by-step memory scheduling strateg
APA, Harvard, Vancouver, ISO, and other styles
10

Du, Haitao, Yuhan Qin, Song Chen, and Yi Kang. "FASA-DRAM: Reducing DRAM Latency with Destructive Activation and Delayed Restoration." ACM Transactions on Architecture and Code Optimization 21, no. 2 (2024): 1–27. http://dx.doi.org/10.1145/3649455.

Full text
Abstract:
DRAM memory is a performance bottleneck for many applications, due to its high access latency. Previous work has mainly focused on data locality, introducing small but fast regions to cache frequently accessed data, thereby reducing the average latency. However, these locality-based designs have three challenges in modern multi-core systems: (1) inter-application interference leads to random memory access traffic, (2) fairness issues prevent the memory controller from over-prioritizing data locality, and (3) write-intensive applications have much lower locality and evict substantial dirty entr
APA, Harvard, Vancouver, ISO, and other styles
More sources

Dissertations / Theses on the topic "Bank Level Parallelism (BLP)"

1

Patil, Adarsh. "Heterogeneity Aware Shared DRAM Cache for Integrated Heterogeneous Architectures." Thesis, 2017. http://etd.iisc.ac.in/handle/2005/4124.

Full text
Abstract:
Integrated Heterogeneous System (IHS) processors pack throughput-oriented GPGPUs along-side latency-oriented CPUs on the same die sharing certain resources, e.g., shared last level cache, network-on-chip (NoC), and the main memory. They also share virtual and physical address spaces and unify the memory hierarchy. The IHS architecture allows for easier programmability, data management and efficiency. However, the significant disparity in the demands for memory and other shared resources between the GPU cores and CPU cores poses significant problems in exploiting the full potential of this arch
APA, Harvard, Vancouver, ISO, and other styles
2

Wang, Shao-Fu, and 王少甫. "Exploiting Bank-level Parallelism via Data Consistency Relaxation for Non-volatile Memory System." Thesis, 2015. http://ndltd.ncl.edu.tw/handle/55825295588163105110.

Full text
Abstract:
碩士<br>國立臺灣大學<br>資訊工程學研究所<br>103<br>The maturity of emerging non-volatile memory (NVM) technologies presents promising next-generation memory system design. Because of its mixed performance characteristics between DRAM and persistent store, e.g., high density, byte-addressability, and non-volatility, architects rethink the design of traditional memory hierarchy. With NVM as main memory, programmer can place non-volatile data structures on main memory and directly access them by ld/st instructions. Non-volatile data structures demand consistency and atomicity guarantees in case of sudden system
APA, Harvard, Vancouver, ISO, and other styles

Conference papers on the topic "Bank Level Parallelism (BLP)"

1

Malik, Kshitiz, Mayank Agarwal, Sam S. Stone, Kevin M. Woley, and Matthew I. Frank. "Branch-mispredict level parallelism (BLP) for control independence." In 2008 IEEE 14th International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2008. http://dx.doi.org/10.1109/hpca.2008.4658628.

Full text
APA, Harvard, Vancouver, ISO, and other styles
2

Tang, Xulong, Mahmut Kandemir, Praveen Yedlapalli, and Jagadish Kotra. "Improving bank-level parallelism for irregular applications." In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2016. http://dx.doi.org/10.1109/micro.2016.7783760.

Full text
APA, Harvard, Vancouver, ISO, and other styles
3

Ding, Wei, Diana Guttman, and Mahmut Kandemir. "Compiler Support for Optimizing Memory Bank-Level Parallelism." In 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2014. http://dx.doi.org/10.1109/micro.2014.34.

Full text
APA, Harvard, Vancouver, ISO, and other styles
4

Lee, Chang Joo, Veynu Narasiman, Onur Mutlu, and Yale N. Patt. "Improving memory bank-level parallelism in the presence of prefetching." In the 42nd Annual IEEE/ACM International Symposium. ACM Press, 2009. http://dx.doi.org/10.1145/1669112.1669155.

Full text
APA, Harvard, Vancouver, ISO, and other styles
5

Kal, Hongju, Chanyoung Yoo, and Won Woo Ro. "AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-Memory." In MICRO '23: 56th Annual IEEE/ACM International Symposium on Microarchitecture. ACM, 2023. http://dx.doi.org/10.1145/3613424.3614314.

Full text
APA, Harvard, Vancouver, ISO, and other styles
6

Poremba, Matthew, Tao Zhang, and Yuan Xie. "Fine-granularity tile-level parallelism in non-volatile memory architecture with two-dimensional bank subdivision." In DAC '16: The 53rd Annual Design Automation Conference 2016. ACM, 2016. http://dx.doi.org/10.1145/2897937.2898024.

Full text
APA, Harvard, Vancouver, ISO, and other styles
7

Kwon, Young-Cheon, Suk Han Lee, Jaehoon Lee, et al. "25.4 A 20nm 6GB Function-In-Memory DRAM, Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism, for Machine Learning Applications." In 2021 IEEE International Solid- State Circuits Conference (ISSCC). IEEE, 2021. http://dx.doi.org/10.1109/isscc42613.2021.9365862.

Full text
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!