What is it about?
We finally sit down and document how Large Language Model based 'Deep Research' tools perform in practice, under the lens of a real life literature review on a topic not yet covered in existing work. After conducting a systematic literature review of methods that use time as a constraint to increase trust in video object detection, we repeated this review using GPT Deep Research, Gemini Deep Research, Perplexity, Consensus, and Elicit. Not only did the tools completely disagree on what literature was relevant, but even combined together the tools barely scratched the surface of relevant methods. We share our full review of temporal features, including everything missed by the Deep Research Tools.
Featured Image
Photo by National Cancer Institute on Unsplash
Why is it important?
This article highlights the shortcomings of a purely LLM-based literature review. It will hopefully assist researchers to make informed choices about their use of LLMs during their research, upholding the reputation of the scientific community. We also share a novel review of methods to use time as a constraint for improved trust in video object detection, which will provide a solid foundation for future work by researchers interested in the trusted AI space.
Perspectives
This paper began as a curiosity of mine and I am glad to have taken the time to explore it. I was always a bit uncertain of how to balance the risks of AI with the promises of productivity, and this paper provided some solid data to make my own informed decisions. I hope it can assist other researchers in a similar manner and encourage others to undertake similar case studies in their own field.
Jack Napier
Griffith University
Read the Original
This page is a summary of: Critiquing AI Surveys with a Case Study on Temporal Features for Trustworthy Computer Vision, ACM Computing Surveys, October 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3849866.
You can read the full text:
Contributors
The following have contributed to this page







