All Stories

  1. Reaction Latency Analysis of Message Synchronization in Edge-assisted Autonomous Driving
  2. μ-ORCA: Optimizing Acceleration for Microsecond-Scale Deep Neural Network Inference on ACAP
  3. DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
  4. To Overlay or to Customize? Revisiting Architectural Choices in Heterogeneous Systems
  5. Advancing Environmental Sustainability in Data Centers via Carbon Depreciation Models
  6. A Survey on Graph Neural Network Acceleration: Algorithms, Systems, and Customized Hardware
  7. AgRefactor: Refactoring for HLS Compatibility with a Self-Evolving Agentic Workflow
  8. Physical Intelligence on the Edge: A Vision for the Decade Ahead
  9. AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
  10. Holistic Optimization Framework for FPGA Accelerators
  11. MTrain: Enable Efficient CNN Training on Heterogeneous FPGA-Based Edge Servers
  12. ART: Customizing Accelerators for DNN-Enabled Real-Time Safety-Critical Systems
  13. Assessing Quantum Layout Synthesis Tools via Known Optimal-SWAP Cost Benchmarks
  14. Invited: Coping with Interconnects
  15. Using a multilevel framework to solve quantum layout synthesis problem.
  16. Stream-HLS: Towards Automatic Dataflow Acceleration
  17. ARIES: An Agile MLIR-Based Compilation Flow for Reconfigurable Devices with AI Engines
  18. A Unified Framework for Automated Code Transformation and Pragma Insertion
  19. InTRRA: Inter-Task Resource-Repurposing Accelerator for Efficient Transformer Inference on FPGAs
  20. SAT-Accel: A Modern SAT Solver on a FPGA
  21. Compilation for Dynamically Field-Programmable Qubit Arrays with Efficient and Provably Near-Optimal Scheduling
  22. Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach
  23. FiberFlex: FPGA-based Intelligent & Distributed Fiber Sensor System for Pedestrian Recognition
  24. Amortizing Embodied Carbon Across Generations
  25. CHEF: A Framework for Deploying Heterogeneous Models on Clusters With Heterogeneous FPGAs
  26. EQ-ViT: Algorithm-Hardware Co-Design for End-to-End Acceleration of Real-Time Vision Transformer Inference on Versal ACAP Architecture
  27. Efficient Task Transfer for HLS DSE
  28. Quantum State Preparation Circuit Optimization Exploiting Don't Cares
  29. GNN-Based Performance Prediction of Quantum Optimization of Maximum Independent Set
  30. RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
  31. Reducing Smart Phone Environmental Footprints with In-Memory Processing
  32. Learning to Compare Hardware Designs for High-Level Synthesis
  33. Cross-Modality Program Representation Learning for Electronic Design Automation with High-Level Synthesis
  34. PASTA: Programming and Automation Support for Scalable Task-Parallel HLS Programs on Modern Multi-Die FPGAs
  35. CHARM 2.0: Composing Heterogeneous Accelerators for Deep Learning on Versal ACAP Architecture
  36. SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
  37. Q-Pilot: Field Programmable Qubit Array Compilation with Flying Ancillas
  38. Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and Synthesis
  39. SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry Clustering
  40. TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
  41. Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach
  42. SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration
  43. FPGA-based Accelerator for Sparse Triangular Solver
  44. Scheduling and Physical Design
  45. Challenges and Opportunities to Enable Large-Scale Computing via Heterogeneous Chiplets
  46. REFRESH FPGAs: Sustainable FPGA Chiplet Architectures
  47. AIM: Accelerating Arbitrary-Precision Integer Multiplication on Heterogeneous Reconfigurable Computing Platform Versal ACAP
  48. Efficient Hardware and Software Design for On-device Learning
  49. TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
  50. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks
  51. NeSSA: Near-Storage Data Selection for Accelerated Machine Learning Training
  52. High Performance, Low Power Matrix Multiply Design on ACAP: from Architecture, Design Challenges and DSE Perspectives
  53. Rubick: A Synthesis Framework for Spatial Architectures via Dataflow Decomposition
  54. Scalable Optimal Layout Synthesis for NISQ Quantum Processors
  55. Lightning Talk: Scaling Up Quantum Compilation – Challenges and Opportunities
  56. A Comprehensive Automated Exploration Framework for Systolic Array Designs
  57. RapidStream 2.0: Automated Parallel Implementation of Latency Insensitive FPGA Designs Through Partial Reconfiguration
  58. FPGA Acceleration of Probabilistic Sentential Decision Diagrams with High-level Synthesis
  59. FlexCNN: An End-to-end Framework for Composing CNN Accelerators on FPGA
  60. HMLib: Efficient Data Transfer for HLS Using Host Memory
  61. Callipepla: Stream Centric Instruction Set and Mixed Precision for Accelerating Conjugate Gradient Solver
  62. CHARM: C omposing H eterogeneous A ccele R ators for M atrix Multiply on Versal ACAP Architecture
  63. Sustainable AI Processing at the Edge
  64. FPGA HLS Today: Successes, Challenges, and Opportunities
  65. Enabling Weakly Supervised Temporal Action Localization From On-Device Learning of the Video Stream
  66. Qubit Mapping for Reconfigurable Atom Arrays
  67. OverGen: Improving FPGA Usability through Domain-specific Overlay Generation
  68. EF-Train: Enable Efficient On-device CNN Training on FPGA through Data Reshaping for Online Adaptation or Personalization
  69. Energy-Efficient LSTM Inference Accelerator for Real-Time Causal Prediction
  70. N-DISE
  71. AutoDSE: Enabling Software Programmers to Design Efficient FPGA Accelerators
  72. Serpens
  73. Automated accelerator optimization aided by graph neural networks
  74. Improving GNN-based accelerator design automation with meta learning
  75. H2H
  76. Automated Accelerator Optimization Aided by Graph Neural Networks
  77. SPA-GCN: Efficient and Flexible GCN Accelerator with Application for Graph Similarity Computation
  78. Sextans: A Streaming Accelerator for General-Purpose Sparse-Matrix Dense-Matrix Multiplication
  79. Accelerating SSSP for Power-Law Graphs
  80. RapidStream
  81. Algorithm-hardware Co-design of Attention Mechanism on FPGA Devices
  82. TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric Notation
  83. AutoBridge
  84. Extending High-Level Synthesis for Task-Parallel Programs
  85. HBM Connect: High-Performance HLS Interconnect for FPGA HBM
  86. MOCHA
  87. AutoDSE: Enabling Software Programmers Design Efficient FPGA Accelerators
  88. AutoSA
  89. BLINK
  90. HeteroRefactor
  91. Bonsai: High-Performance Adaptive Merge Tree Sorting
  92. Algorithm-Hardware Co-design for BQSR Acceleration in Genome Analysis ToolKit
  93. Caffeine: Toward Uniformed Representation and Acceleration for Deep Convolutional Neural Networks
  94. Overcoming Data Transfer Bottlenecks in FPGA-based DNN Accelerators via Layer Conscious Memory Management
  95. Dataflow Systolic Array Implementations of Matrix Decomposition Using High Level Synthesis
  96. LANMC
  97. HeteroCL
  98. Overcoming Data Transfer Bottlenecks in DNN Accelerators via Layer-Conscious Memory Managment
  99. HLS-based optimization and design space exploration for applications with variable loop bounds
  100. PolySA
  101. SODA
  102. TGPA
  103. Doppio: I/O-Aware Performance Analysis, Modeling and Optimization for In-memory Computing Framework
  104. ST-Accel: A High-Level Programming Platform for Streaming Applications on FPGA
  105. Latte: Locality Aware Transformation for High-Level Synthesis
  106. CPU-FPGA Co-Optimization for Big Data Applications
  107. Bandwidth Optimization Through On-Chip Memory Restructuring for HLS
  108. Caffeine
  109. Energy-Efficient CNN Implementation on a Deeply Pipelined FPGA Cluster
  110. Invited - Heterogeneous datacenters
  111. Energy Efficiency of Full Pipelining: A Case Study for Matrix Multiplication
  112. ARAPrototyper
  113. InterFS
  114. CMOST
  115. On-chip interconnection network for accelerator-rich architectures
  116. A Fully Pipelined and Dynamically Composable Architecture of CGRA
  117. Automatic memory partitioning and scheduling for throughput and power optimization