All Stories

  1. Self-Speculative Decoding for On-device MoE Acceleration
  2. Poster: Coda: Context-aware Acceleration for Distributed LLM Decoding in Edge
  3. FedNLR: Federated Learning with Neuron-wise Learning Rates