All Stories

  1. Auto-Judge: A Cross-Task Benchmark for Comparing LLM Judges for Citation-Grounded RAG Systems
  2. A Comparative Analysis of Linguistic and Retrieval Diversity in LLM-Generated Search Queries
  3. Principles and Guidelines for the Use of LLM Judges
  4. Enhancing Human Annotation: Leveraging Large Language Models and Efficient Batch Processing
  5. Can Users Predict Relative Query Effectiveness?
  6. sMARE: a new paradigm to evaluate and understand query performance prediction methods
  7. New Perspectives to Query Performance Prediction Evaluation
  8. Is Query Performance Prediction With Multiple Query Variations Harder Than Topic Performance Prediction?
  9. An Enhanced Evaluation Framework for Query Performance Prediction
  10. Utilizing User Email Actions to Improve Ad-Close Prediction
  11. Information Needs, Queries, and Query Performance Prediction