All Stories

  1. From Noise to Order: Learning to Rank via Denoising Diffusion
  2. Failing Forward: Understanding Query Failure in Retrieval, Judgment, and Generation
  3. From Doxa to Logos in Scientific Peer Review
  4. Peerispect : Claim Verification in Scientific Peer Reviews
  5. PeerPrism: Peer Evaluation Expertise vs Review-writing AI
  6. Can QPP Choose the Right Query variant? Evaluating Query Variant Selection for RAG Pipelines
  7. PeeriScope: A Multi-Faceted Framework for Evaluating Peer Review Quality
  8. QueryGym: A Toolkit for Reproducible LLM-Based Query Reformulation
  9. Can LLMs Uphold Research Integrity? Evaluating the Role of LLMs in Peer Review Quality
  10. Query Performance Prediction Using Neural Query Space Proximity
  11. ProActLLM: Proactive Conversational Information Seeking with Large Language Models
  12. Building Trustworthy Peer Review Quality Assessment Systems
  13. RottenReviews: Benchmarking Review Quality with Human and LLM-Based Judgments
  14. A Human-AI Comparative Analysis of Prompt Sensitivity in LLM-Based Relevance Judgment
  15. VAP3: Variation-Aware Prompt Performance Prediction
  16. IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
  17. Benchmarking LLM-based Relevance Judgment Methods
  18. Query Performance Prediction: Theory, Techniques and Applications
  19. Query Performance Prediction: Techniques and Applications in Modern Information Retrieval
  20. Evaluating Relative Retrieval Effectiveness with Normalized Residual Gain
  21. Offline Evaluation of Set-Based Text-to-Image Generation
  22. Reviewerly: Modeling the Reviewer Assignment Task as an Information Retrieval Problem
  23. Enhanced Retrieval Effectiveness through Selective Query Generation
  24. Retrieving Supporting Evidence for Generative Question Answering
  25. Noisy Perturbations for Estimating Query Difficulty in Dense Retrievers
  26. A is for Adele: An Offline Evaluation Metric for Instant Search
  27. Quantifying Ranker Coverage of Different Query Subspaces
  28. A Preference Judgment Tool for Authoritative Assessment
  29. Gender Fairness in Information Retrieval Systems
  30. Addressing Gender-related Performance Disparities in Neural Rankers
  31. Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection
  32. MS MARCO Chameleons: Challenging the MS MARCO Leaderboard with Extremely Obstinate Queries
  33. Matches Made in Heaven: Toolkit and Large-Scale Datasets for Supervised Query Reformulation
  34. BERT-QPP: Contextualized Pre-trained transformers for Query Performance Prediction
  35. On the Orthogonality of Bias and Utility in Ad hoc Retrieval
  36. Geometric Estimation of Specificity within Embedding Spaces