All Stories

  1. GPU Acceleration in Acoustics and Audio Signal Processing
  2. Move the Roof: Model-driven Methodology for Designing Efficient Deep Learning Architectures
  3. Interpreting High Order Epistasis Using Sparse Transformers
  4. Transient-Execution Attacks: a Computer Architect Perspective
  5. Temperature-aware Core Management in MPSoCs: Modeling and Evaluation using MRMs
  6. A New Energy-Efficient Hybrid Wide-Operand Adder Architecture
  7. A Survey on Fully Homomorphic Encryption
  8. Adaptive Scheduling Framework for Real-Time Video Encoding on Heterogeneous Systems
  9. $2^n$ RNS Scalers for Extended 4-Moduli Sets
  10. GPU-assisted HEVC intra decoder
  11. Arithmetic-Based Binary-to-RNS Converter Modulo ${\{2^{n}{\pm}k\}}$ for $jn$ -bit Dynamic Range
  12. Reverse Converter Design via Parallel-Prefix Adders: Novel Components, Methodology, and Implementations
  13. Stretching the limits of Programmable Embedded Devices for Public-key Cryptography
  14. ROM-less RNS-to-binary converter moduli {22n − 1, 22n + 1, 2n − 3, 2n + 3}
  15. Collaborative inter-prediction on CPU+GPU systems
  16. Reconfigurable data flow engine for HEVC motion estimation
  17. On the Evaluation of Multi-core Systems with SIMD Engines for Public-Key Cryptography
  18. Performance-Aware Task Management and Frequency Scaling in Embedded Systems
  19. FEVES: Framework for Efficient Parallel Video Encoding on Heterogeneous Systems
  20. Efficient sign identification engines for integers represented in RNS extended 3-moduli set {2 n − 1, 2 n + k , 2 n + 1}
  21. Unified transform architecture for AVC, AVS, VC-1 and HEVC high-performance codecs
  22. Method for designing multi-channel RNS architectures to prevent power analysis SCA
  23. Combining flexibility with low power: Dataflow and wide-pipeline LDPC decoding engines in the Gbit/s era
  24. Cooperative CPU+GPU deblocking filter parallelization for high performance HEVC video codecs
  25. Efficient Multilevel Load Balancing on Heterogeneous CPU + GPU Systems
  26. Design and Optimization of Scientific Applications for Highly Heterogeneous and Hierarchical HPC Platforms Using Functional Computation Performance Models
  27. A Flexible Architecture for Modular Arithmetic Hardware Accelerators based on RNS
  28. An Efficient Scalable RNS Architecture for Large Dynamic Ranges
  29. Cache-aware Roofline model: Upgrading the loft
  30. Finite-Difference in Time-Domain Scalable Implementations on CUDA and OpenCL
  31. Dynamic Load Balancing for Real-Time Video Encoding on Heterogeneous CPU+GPU Systems
  32. EFFICIENT METHOD FOR DESIGNING MODULO {2 n ± k} MULTIPLIERS
  33. SchedMon: A Performance and Energy Monitoring Tool for Modern Multi-cores
  34. Monitoring Performance and Power for Application Characterization with the Cache-Aware Roofline Model
  35. Exploiting Coarse-grained Parallelism in Multi-transform Architectures for H.264/AVC High Profile Codecs
  36. Method to Design General RNS Reverse Converters for Extended Moduli Sets
  37. Open the Gates: Using High-level Synthesis towards programmable LDPC decoders on FPGAs
  38. Randomised multi-modulo residue number system architecture for double-and-add to prevent power analysis side channel attacks
  39. A Lab Project on the Design and Implementation of Programmable and Configurable Embedded Systems
  40. A comparison of computing architectures and parallelization frameworks based on a two-dimensional FDTD
  41. Exploiting task and data parallelism for advanced video coding on hybrid CPU + GPU platforms
  42. A compact and scalable RNS architecture
  43. RNS Reverse Converters for Moduli Sets With Dynamic Ranges up to $(8n+1)$ -bit
  44. An RNS-based architecture targeting hardware accelerators for modular arithmetic
  45. Accelerating the Computation of Induced Dipoles for Molecular Mechanics with Dataflow Engines
  46. The CRNS framework and its application to programmable and reconfigurable cryptography
  47. Multi-level Parallelization of Advanced Video Coding on Hybrid CPU+GPU Platforms
  48. Reconfigurable Architecture for Cryptography over Binary Finite Fields
  49. 2-Axis Magnetometers Based on Full Wheatstone Bridges Incorporating Magnetic Tunnel Junctions Connected in Series
  50. Scalable Unified Transform Architecture for Advanced Video Coding Embedded Systems
  51. Real-time implementation of remotely sensed hyperspectral image unmixing on GPUs
  52. RNS Arithmetic Units for Modulo {2^n+-k}
  53. VLSI Reverse Converter for RNS Based on the Moduli Set
  54. High Performance Unified Architecture for Forward and Inverse Quantization in H.264/AVC
  55. Fine-grain parallelism using multi-core, Cell/BE, and GPU Systems
  56. Energy efficient stream-based configurable architecture for embedded platforms
  57. Simultaneous Multi-Level Divisible Load Balancing for Heterogeneous Desktop Systems
  58. Computation of Induced Dipoles in Molecular Mechanics Simulations Using Graphics Processors
  59. Corrections to “MRC-Based RNS Reverse Converters for the Four-Moduli Sets $\{2^{n} + 1,\ 2^{n} - 1,\ 2^{n},\ 2^{2n + 1} - 1\}$ and
  60. MRC-Based RNS Reverse Converters for the Four-Moduli Sets $\{2^{n} + 1, 2^{n} - 1, 2^{n}, 2^{2n + 1} - 1\}$ and $ \{2^{n} + 1, 2^{n} - ...
  61. Configurable M-factor VLSI DVB-S2 LDPC decoder architecture with optimized memory tiling design
  62. On Realistic Divisible Load Scheduling in Highly Heterogeneous Distributed Systems
  63. Efficient implementation of multi-moduli architectures for Binary-to-RNS conversion
  64. Scheduling Divisible Loads on Heterogeneous Desktop Systems with Limited Memory
  65. Hierarchical Partitioning Algorithm for Scientific Computing on Highly Heterogeneous CPU + GPU Clusters
  66. A tutorial overview on the properties of the discrete cosine transform for encoded image and video processing
  67. Parallel Computing – Special Issue
  68. Binary-to-RNS Conversion Units for moduli {2^n ± 3}
  69. High throughput and scalable architecture for unified transform coding in embedded H.264/AVC video coding systems
  70. Real-time DVB-S2 LDPC decoding on many-core GPU accelerators
  71. Massively LDPC Decoding on Multicore Architectures
  72. Parallel LDPC Decoding
  73. Introduction
  74. A quantitative analysis of firing rate estimators: Unveiling bias sources
  75. Exploiting SIMD extensions for linear image processing with OpenCL
  76. Hardware/software co-design of H.264/AVC encoders for multi-core embedded systems
  77. H.264/AVC framework for multi-core embedded video encoders
  78. Unifying stream based and reconfigurable computing to design application accelerators
  79. An improved RNS generator 2n ± k based on threshold logic
  80. Arithmetic Units for RNS Moduli {2n-3} and {2n+3} Operations
  81. Embedded multicore architectures for LDPC decoding
  82. Elliptic Curve point multiplication on GPUs
  83. Efficient Independent Component Analysis on a GPU
  84. Challenges and trends in the development of a magnetoresistive biochip portable platform
  85. Programming Cell/BE and GPUs systems for real-time video encoding
  86. Collaborative execution environment for heterogeneous parallel systems
  87. Modeling and Evaluating Non-shared Memory CELL/BE Type Multi-core Architectures for Local Image and Video Processing
  88. Preface
  89. Euro-Par 2009 – Parallel Processing Workshops
  90. Iterative induced dipoles computation for molecular mechanics on GPUs
  91. p264
  92. Development and evaluation of scalable video motion estimators on GPU
  93. Fine-grain Parallelism Using Multi-core, Cell/BE, and GPU Systems: Accelerating the Phylogenetic Likelihood Function
  94. Parallel LDPC Decoding on GPUs Using a Stream-Based Computing Approach
  95. Modelling and programming stream-based distributed computing based on the meta-pipeline approach
  96. How GPUs can outperform ASICs for fast LDPC decoding
  97. Multi-core platforms for signal processing: source and channel coding
  98. Neural code metrics: Analysis and application to the assessment of neural models
  99. CaravelaMPI: Message Passing Interface for Parallel GPU-Based Applications
  100. Distributed Software Platform for Automation and Control of General Anaesthesia
  101. A Portable and Autonomous Magnetic Detection Platform for Biosensing
  102. BIOELECTRONIC VISION
  103. Compact and Flexible Microcoded Elliptic Curve Processor for Reconfigurable Devices
  104. Bioelectronic Vision
  105. Applying the Stream-Based Computing Model to Design Hardware Accelerators: A Case Study
  106. Parallel LDPC Decoding on the Cell/B.E. Processor
  107. On the design of distributed autonomous embedded systems for biomedical applications
  108. Efficient FPGA elliptic curve cryptographic processor over GF(2m)
  109. Design and implementation of a tool for modeling and programming deadlock free meta-pipeline applications
  110. Merged Computation for Whirlpool Hashing
  111. Edge Stream Oriented LDPC Decoding
  112. On-the-fly attestation of reconfigurable hardware
  113. Merged computation for Whirlpool hashing
  114. A Parallel Algorithm for Advanced Video Motion Estimation on Multicore Architectures
  115. Low power microarchitecture with instruction reuse
  116. Distributed Web-based Platform for Computer Architecture Simulation
  117. Heuristic Optimization Methods for Improving Performance of Recursive General Purpose Applications on GPUs
  118. Application Specific Programmable IP Core for Motion Estimation: Technology Comparison Targeting Efficient Embedded Co-Processing Units
  119. An RNS based Specific Processor for Computing the Minimum Sum-of-Absolute-Differences
  120. BRAM-LUT Tradeoff on a Polymorphic DES Design
  121. Reconfigurable architectures and processors for real-time video motion estimation
  122. QCA-LG: A tool for the automatic layout generation of QCA combinational circuits
  123. Efficient Hybrid DCT-Domain Algorithm for Video Spatial Downscaling
  124. A Run-Time Reconfigurable Processor for Video Motion Estimation
  125. Meta-Pipeline: A New Execution Mechanism for Distributed Pipeline Processing
  126. Adaptive Motion Estimation Algorithm for H.264/AVC
  127. An Efficient Expectation-Maximisation Algorithm for Spike Classification
  128. An ASIP approach for adaptive AVC Motion Estimation
  129. Efficient Method for Magnitude Comparison in RNS Based on Two Pairs of Conjugate Moduli
  130. Caravela: A Novel Stream-Based Distributed Computing Environment
  131. Additive Logistic Regression Applied to Retina Modelling
  132. Feature Selection for the Stochastic Integrate and Fire Model
  133. Design and implementation of a stream-based distributedcomputing platform using graphics processing units
  134. Data buffering optimization methods toward a uniform programming interface for gpu-based applications
  135. Embedded Systems for Portable and Mobile Video Platforms
  136. A New Hand-Held Microsystem Architecture for Biological Analysis
  137. MAESTRO2: EXPERIMENTAL EVALUATION OF COMMUNICATION PERFORMANCE IMPROVEMENT TECHNIQUES IN THE LINK LAYER
  138. Configurable Embedded Core for Controlling Electro-Mechanical Systems
  139. Improving SHA-2 Hardware Implementations
  140. Rescheduling for Optimized SHA-1 Calculation
  141. Low Power Distance Measurement Unit for Real-Time Hardware Motion Estimators
  142. On Task Scheduling Accuracy: Evaluation Methodology and Results
  143. List scheduling: extension for contention awareness and evaluation of node priorities for heterogeneous cluster architectures
  144. A programmable cellular neural network circuit
  145. Fast transcoding architectures for insertion of non-regular shaped objects in the compressed DCT-domain
  146. An FPL Bioinspired Visual Encoding System to Stimulate Cortical Neurons in Real-Time
  147. Customisable Core-Based Architectures for Real-Time Motion Estimation on FPGAs
  148. A New Efficient VLSI Architecture for Full Search Block Matching Motion Estimation
  149. Synchronous Non-local Image Processing on Orthogonal Multiprocessor Systems
  150. Exploiting Unused Time Slots in List Scheduling Considering Communication Contention
  151. A Platform Independent Parallelising Tool Based on Graph Theoretic Models
  152. Scheduling Task Graphs on Arbitrary Processor Architectures Considering Contention
  153. Customizable and Reduced Hardware Motion Estimation Processors
  154. Massive Data Classification of Neural Responses
  155. Bioinspired Stimulus Encoder for Cortical Visual Neuroprostheses
  156. On the Implementation and Evaluation of Berkeley Sockets on Maestro2 cluster computing environment
  157. Nanotechnology and the Detection of Biomolecular Recognition Using Magnetoresistive Transducers