All Stories

  1. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
  2. DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O
  3. TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
  4. CMANNS: GPU-Accelerated Graph Index Construction for ANNS via Compute–Memory Disaggregation
  5. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
  6. Accelerating Stream Processing Engines via Hardware Offloading
  7. KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models
  8. Scaling Up Memory Disaggregated Applications with SMART
  9. Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared Memory
  10. Falcon: Fast OLTP Engine for Persistent Cache and Non-Volatile Memory
  11. Efficiently Answering Path Queries on Evolving Graphs