What is it about?
Short texts are everywhere: tweets, search queries, headlines, and product reviews. But they are hard for computers to organize. They are brief, often messy, and full of words with more than one meaning. For example, “Java” can mean a programming language, an island, or coffee. Traditional methods often fail because they cannot tell these meanings apart. This research introduces SICAR, a new AI method that groups short texts more accurately. SICAR uses context to understand ambiguous words, automatically balances the size of groups, removes noisy or irrelevant items, and keeps improving its labels in repeated rounds. In tests on news, search snippets, social media, and biomedical data, SICAR performed better than existing methods. It improved grouping accuracy by 12% and a key quality measure by 15%. It also filtered out 85% of slang and typographical errors in social media data.
Featured Image
Photo by Max Chen on Unsplash
Read the Original
This page is a summary of: Short-Text Clustering Enhancement Based on Semantic-Aware Iterative Refinement, ACM Transactions on Asian and Low-Resource Language Information Processing, September 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3842665.
You can read the full text:
Contributors
The following have contributed to this page







