What is it about?

Artificial intelligence writing detectors are increasingly being used by schools and universities to identify assignments that may have been generated by AI. However, little evidence exists regarding how accurately these systems evaluate authentic student writing, particularly in multilingual contexts. This study examined two widely used AI writing detectors, ZeroGPT and Copyleaks, by testing them on ten authentic essays written by Filipino undergraduate students. Although every essay was genuinely written by a student, both detectors incorrectly identified half of them as AI-generated. The findings reveal that these tools can produce substantial false positives and may misclassify legitimate student work because of writing characteristics common among non-native English writers. Rather than treating detector results as definitive evidence of misconduct, educators should combine automated outputs with professional judgment and additional evidence before making academic integrity decisions.

Featured Image

Why is it important?

Educational institutions are rapidly adopting AI detection tools, yet relatively little empirical research has examined how these systems perform with authentic student writing from multilingual settings. False accusations based solely on AI detector outputs can negatively affect students' academic records, confidence, and trust in educational assessment. This study provides preliminary evidence that detector outputs should be interpreted cautiously, especially when evaluating students whose writing differs from the English-language datasets used to train many commercial detection systems. The findings support more responsible, transparent, and fair approaches to academic integrity by emphasizing that AI detection should assist not replace human evaluation.

Perspectives

Artificial intelligence will continue to influence educational assessment, but technological convenience should never outweigh fairness. Future research should examine larger multilingual datasets, evaluate additional AI detection systems, and investigate why specific linguistic characteristics trigger false positives. Developing more transparent and equitable detection technologies will help educational institutions balance academic integrity with the protection of students from incorrect accusations.

Mhel Cedric Bendo

Read the Original

This page is a summary of: False Positives in AI Writing Detection: A Small-Scale Empirical Study Using Authentic Filipino Student Essays, ASEAN Journal of Open and Distance Learning, April 2026, Open University Malaysia,
DOI: 10.64233/vyvi9613.
You can read the full text:

Read

Contributors

The following have contributed to this page