What is it about?
Research data centers give researchers access to sensitive microdata in secure environments, but results such as tables, statistics, and models must be checked before they can leave the secure environment. This output-checking process is essential for protecting confidentiality, but it can be time-consuming and difficult to scale. We propose a semi-automated approach that combines three layers: rule-based disclosure checks, a machine-learning model trained on previous release decisions, and a large language model agent that brings the available evidence together and produces structured explanations. Using historical output-checking records from a national statistical office, we evaluate how this architecture can help identify straightforward low-risk cases while directing uncertain or potentially risky outputs to professional reviewers. The system is designed to support human decision-making rather than replace it. [ Ashofteh, A., Carvalho, R., & Campos, P. (2026). Designing LLM Agents for Output Checking in Research Data Centers: Disclosure Risk and Output Validation. In 2026 IEEE 50th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 189-195). (Proceedings of the Annual Computer Software and Applications Conference). IEEE Computer Society. https://doi.org/10.1109/COMPSAC69091.2026.00035 ]
Featured Image
Photo by Laura Rivera on Unsplash
Why is it important?
Research data centers need to balance two competing objectives: enabling researchers to obtain useful results quickly while ensuring that confidential information cannot be disclosed. Manual output checking provides an important safeguard, but increasing demand for access to secure microdata makes the process increasingly difficult to scale. Our work shows how AI-assisted output checking can support this process by automatically identifying obvious disclosure risks, learning from historical institutional decisions, and generating structured explanations for reviewers. Importantly, the proposed approach does not give an LLM autonomous authority to release sensitive outputs. Ambiguous, unsupported, or high-risk cases remain under human control. This provides a practical path toward more scalable output checking while preserving accountability, confidentiality, and professional oversight.
Perspectives
AI agents could substantially reduce the workload involved in checking research outputs, but this is an area where automation must be introduced carefully. A system that incorrectly approves an unsafe result can create confidentiality risks, while an overly restrictive system can unnecessarily delay legitimate research. Our perspective is therefore that LLM agents should act as decision-support tools rather than autonomous gatekeepers. Deterministic rules, machine-learning predictions, LLM-based reasoning, and professional judgment each address different aspects of the output-checking problem. The most promising direction is a governed human-in-the-loop architecture in which AI accelerates routine cases, explains the evidence behind its recommendations, and reliably escalates uncertainty to trained output checkers.
Prof. Afshin Ashofteh
Universidade Nova de Lisboa
Read the Original
This page is a summary of: Designing LLM Agents for Output Checking in Research Data Centers: Disclosure Risk and Output Validation, July 2026, Institute of Electrical & Electronics Engineers (IEEE),
DOI: 10.1109/compsac69091.2026.00035.
You can read the full text:
Contributors
The following have contributed to this page







