All Stories

  1. Evidence-chain-driven multimodal retrieval question answering
  2. MARS: Multimodal-Assisted Refined Semantic Alignment
  3. ACRA: An Adaptive Chain Retrieval Architecture for Multi-modal Knowledge-Augmented Visual Question Answering
  4. McGE '25: The 3rd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice
  5. Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music Generation
  6. Proceedings of the 3rd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice
  7. Decomposition and Foresight: Comparing Human and Simulated Teacher in Preference-Based Reinforcement Learning
  8. Mitigating reasoning hallucination through Multi-agent Collaborative Filtering
  9. Spatio-Temporal Deviation Calibration for Skeleton-Based Human Action Recognition
  10. MultiColor: Image Colorization by Learning from Multiple Color Spaces
  11. Charting the Uncharted: Building and Analyzing a Multifaceted Chart Question Answering Dataset for Complex Logical Reasoning Process
  12. Cross-domain document layout analysis using document style guide
  13. Atlantis: Aesthetic-oriented multiple granularities fusion network for joint multimodal aspect-based sentiment analysis
  14. K-NNDP: K-means algorithm based on nearest neighbor density peak optimization and outlier removal
  15. Simple contrastive learning in a self-supervised manner for robust visual question answering
  16. Cross-modal fine-grained alignment and fusion network for multimodal aspect-based sentiment analysis
  17. Automatic Image Aesthetic Assessment for Human-designed Digital Images
  18. Reading Scene Text with Aggregated Temporal Convolutional Encoder
  19. Progressive scene text erasing with self-supervision
  20. DDT: Dual-branch Deformable Transformer for Image Denoising
  21. Dual-Expert Distillation Network for Few-Shot Segmentation
  22. Image Layer Modeling for Complex Document Layout Generation
  23. LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion
  24. DRFN: A unified framework for complex document layout analysis
  25. Modeling Stroke Mask for End-to-End Text Erasing
  26. Chinese herbal recognition by Spatial-/Channel-wise attention
  27. 3D Clues Guided Convolution for Depth Completion
  28. Document Layout Analysis Via Positional Encoding
  29. A survey of human-in-the-loop for machine learning
  30. Adaptive Multi-Feature Extraction Graph Convolutional Networks for Multimodal Target Sentiment Analysis
  31. Graph Convolution over the Semantic-syntactic Hybrid Graph Enhanced by Affective Knowledge for Aspect-level Sentiment Classification
  32. Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis
  33. Lightweight Network Based Real-time Anomaly Detection Method for Caregiving at Home
  34. Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection
  35. Chinese Herbal Recognition Databases Using Human-In-The-Loop Feedback
  36. Document image layout analysis via explicit edge embedding network
  37. Document Layout Analysis via Dynamic Residual Feature Fusion
  38. MT-YOLOv5: Mobile terminal table detection model based on YOLOv5
  39. Optimizing Speed/Accuracy Table Detection via Knowledge Distillation
  40. TCATD: Text Contour Attention for Scene Text Detection
  41. LSTMVAEF: Vivid Layout via LSTM-Based Variational Autoencoder Framework
  42. CPSPNet: Crowd Counting via Semantic Segmentation Framework
  43. Margin Guidance Network for Arbitrary-shaped Scene Text Detection
  44. Fast video crowd counting with a Temporal Aware Network
  45. Counting Crowds with Perspective Distortion Correction via Adaptive Learning
  46. Counting crowds with varying densities via adaptive scenario discovery framework
  47. Scene Text Recognition with Temporal Convolutional Encoder
  48. Feature Channel Enhancement for Crowd Counting
  49. Adaptive Scenario Discovery for Crowd Counting
  50. Aggregating Rich Deep Semantic Features for Fine-Grained Place Classification