All Stories

  1. Minimizing the pretraining gap: Domain-aligned text-based person retrieval
  2. Harnessing weak pair uncertainty for text-based person search
  3. SynthIR: The First Workshop on Synthetic Content in Information Retrieval Ecosystems
  4. Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
  5. Look, Compare and Draw: Differential Query Transformer for Automatic Oil Painting
  6. Understanding Image Retrieval Re-Ranking: A Graph Neural Network Perspective
  7. Evaluating and enhancing the performance of large language models in thyroid eye disease through customization and Chain-of-Thought strategies
  8. Event-based Lip Reading with Triplane Fusion Network
  9. Progressive Text-to-3D Generation for Automatic 3D Prototyping
  10. CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution
  11. FANet: Fovea Attention Network for Robust Aerial Geo-Localization Across Diverse Weather Conditions
  12. Joint Attribute Graph Reasoning and Aggregation for Composed Image Retrieval
  13. RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning
  14. SAGINGeo: A Space–Aerial–Ground Integrated Framework for VGI Geolocalization in Multidisaster Scenarios
  15. Introduction to the Special Issue on Deep Multimodal Generation and Retrieval
  16. Domain-Agnostic Neural Oil Painting via Normalization Affine Test-Time Adaptation
  17. UniAD: Integrating Geometric and Semantic Cues for Unified Anomaly Detection
  18. Proceedings of the 3rd International Workshop on UAVs in Multimedia: Capturing the World from a New Perspective
  19. The 3rd Workshop on UAVs in Multimedia: Capturing the World from a New Perspective
  20. The 9th AI City Challenge
  21. MORE'25 Multimedia Object Re-ID: Advancements, Challenges, and Opportunities
  22. From Data Deluge to Data Curation: A Filtering-WoRA Paradigm for Efficient Text-based Person Search
  23. CAMeL: Cross-Modality Adaptive Meta-Learning for Text-Based Person Retrieval
  24. EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
  25. Scale Up Composed Image Retrieval Learning via Modification Text Generation
  26. Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
  27. Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
  28. The 2nd Workshop on UAVs in Multimedia: Capturing the World from a New Perspective
  29. The 2nd International Workshop on Deep Multi-modal Generation and Retrieval
  30. Transferring to Real-World Layouts: A Depth-aware Framework for Scene Adaptation
  31. Depth-Aware Blind Image Decomposition for Real-World Adverse Weather Recovery
  32. Self-ensembling depth completion via density-aware consistency
  33. Collaborative group: Composed image retrieval via consensus learning from noisy annotations
  34. Multiple-environment Self-adaptive Network for aerial-view geo-localization
  35. MORE'24 Multimedia Object Re-ID: Advancements, Challenges, and Opportunities
  36. High Fidelity Makeup via 2D and 3D Identity Preservation Net
  37. StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition
  38. Jointly Harnessing Prior Structures and Temporal Consistency for Sign Language Video Generation
  39. Active Discovering New Slots for Task-Oriented Conversation
  40. Learning Cross-View Geo-Localization Embeddings via Dynamic Weighted Decorrelation Regularization
  41. UAVM '23: 2023 Workshop on UAVs in Multimedia: Capturing the World from a New Perspective
  42. Deep Multimodal Learning for Information Retrieval
  43. PiPa: Pixel- and Patch-wise Self-supervised Learning for Domain Adaptative Semantic Segmentation
  44. Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark
  45. Learnable Pillar-based Re-ranking for Image-Text Retrieval
  46. Are Binary Annotations Sufficient? Video Moment Retrieval via Hierarchical Uncertainty-based Active Learning
  47. Context-Aware Pretraining for Efficient Blind Image Decomposition
  48. Multi-view Consistent Generative Adversarial Networks for Compositional 3D-Aware Image Synthesis
  49. Align and Tell: Boosting Text-Video Retrieval With Local Alignment and Fine-Grained Supervision
  50. Progressive Local Filter Pruning for Image Retrieval Acceleration
  51. U-Turn: Crafting Adversarial Queries with Opposite-Direction Features
  52. Soft Person Reidentification Network Pruning via Blockwise Adjacent Filter Decaying
  53. Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis
  54. Each Part Matters: Local Patterns Facilitate Cross-View Geo-Localization
  55. Adaptive Boosting for Domain Adaptation: Toward Robust Predictions in Scene Segmentation
  56. DMRNet++: Learning Discriminative Features with Decoupled Networks and Enriched Pairs for One-Step Person Search
  57. Joint Representation Learning and Keypoint Detection for Cross-View Geo-Localization
  58. Parameter-Efficient Person Re-Identification in the 3D Space
  59. SPG-VTON: Semantic Prediction Guidance for Multi-Pose Virtual Try-on
  60. Self-supervised Point Cloud Representation Learning via Separating Mixed Shapes
  61. Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person Search
  62. Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation
  63. VehicleNet: Learning Robust Visual Representation for Vehicle Re-Identification
  64. University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization
  65. Real-World Automatic Makeup via Identity Preservation Makeup Net
  66. Unsupervised Scene Adaptation with Memory Regularization in vivo
  67. Dual-path Convolutional Image-Text Embeddings with Instance Loss
  68. Going Beyond Real Data: A Robust Visual Representation for Vehicle Re-identification
  69. Thorax disease classification with attention guided convolutional neural network
  70. Bayesian query expansion for multi-camera person re-identification
  71. Unsupervised Eyeglasses Removal in the Wild
  72. Improving person re-identification by attribute and identity learning
  73. Pedestrian Alignment Network for Large-scale Person Re-Identification
  74. Joint Discriminative and Generative Learning for Person Re-Identification
  75. CamStyle: A Novel Data Augmentation Method for Person Re-Identification
  76. Multi-Pseudo Regularized Label for Generated Data in Person Re-Identification
  77. Camera Style Adaptation for Person Re-identification
  78. An improved artificial intelligence scheme for person reidentification
  79. Macro-Micro Adversarial Network for Human Parsing
  80. Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in Vitro