
NEW YORK, NEW YORK – FEBRUARY 16: In this illustration, the Claude AI app is seen in the app store on a phone on February 16, 2026 in New York City. According to reports from the Wall Street Journal, the Defense Department used Anthropic’s Claude Ai, via its Palantir contract, to help with the attack on Venezuela and capture former President Nicolás Maduro. (Photo illustration by Michael M. Santiago/Getty Images)
Getty Images
Since the public launch of ChatGPT in 2022, generative AI has increasingly threatened academic integrity, further stoking fears of a literacy crisis and deepening concerns about the authenticity of students’ work. Some schools and university systems have implemented AI detection methods, but these have often proved unreliable, leading many students anxious about the possibility of being falsely accused.
Now, a new detection method originating in the generative AI itself may complicate the educational landscape even further.
Last week, Anthropic announced that Claude would be implementing measures to comply with the EU’s Artificial Intelligence Act by watermarking content generated by the system. The watermark is undetectable to readers and persists across copy-pasted content, even after minor editing. It works by altering the probability of word patterns in a given text, such that a human may not be able to spot the pattern, but a machine with the back-end code could detect its use. Because the watermark is embedded in the word choice itself, detection is more reliable with larger samples of text, whereas shorter passages and grammar and copy editing are unlikely to carry the watermark. While the company has yet to introduce accompanying detection software, it has indicated that a detection API is forthcoming.
For some, these new measures promise to weed out AI use in the classroom, whether by discouraging students from relying on LLMs or equipping professors with a foolproof method of identifying its use. A headline in The Sydney Morning Herald, for instance, declared last week: “AI school essay cheating set to be exposed by watermarks.”
On the one hand, concerns about AI use in higher education are not baseless. Professors have encountered demonstrable issues with widespread cheating, with many forced to devise their own methods for eliminating unethical AI use. The desire for a standard detection method to avoid false accusations while maintaining academic honesty is understandable. But the uncertainties around the functionality of Claude’s watermarking may present new challenges, rather than heralding a new era of certainty around AI detection.
The first issue that watermarking raises lies in the degree and nature of the LLM’s use. Importantly, the watermark will only answer the very dubious question: “What is the likelihood this was partly written by Claude?” It does not indicate to what extent Claude was used (whether to edit or outright generate content), nor can it verify that the text was human-made rather than generated by a different bot. This means that a student who uses Claude to rework a few clunky sentences in their prose is likely to raise the same detection flags as one who used the bot to write their entire essay from scratch. Because the watermark degrades with more thorough editing or in shorter samples of text, its usefulness may also be uneven depending on the department or assignment. In light of this, professors must face the new challenge of determining whether any likelihood detected by the watermark constitutes cheating, or if there is a threshold for acceptable use.
This ambiguity is complicated still further by questions around the availability of detection software. While the release of this software is imminent, Anthropic has not indicated how institutions or users will have access to the product. If the detection software is tied to institutional licensing agreements, it could create a hierarchy in which well-funded research institutions are better equipped to weed out AI generated language than underfunded schools and community colleges. As previous iterations of AI detection became available, students became savvy at outsmarting the software by altering the generated writing—even going so far as to “dumb-down” their writing in order to evade detection. Ironically, it was not just the AI that corroded students’ writing abilities, but the detection measures themselves. Clarity around the institutional use of detection software can help to quell student concerns and discourage intentional writing mistakes as a sign of human-made content.
Finally, as more and more students absorb the vocabulary, syntax, and cadence of AI bots through near-constant exposure to AI writing, it remains to be seen whether watermarking might appear in writing that implicitly follows Claude’s probability patterns, even where the bot was not directly used. Without a detection tool currently available, it is unclear whether false positives will be a feature of this new era of AI use. A new tool marketed as a reliable means of detecting AI generation could carry significant weight in disciplinary hearings over academic dishonesty, making the stakes for false positives incredibly high.
Claude’s watermarking may indeed prove valuable as a deterrent for academic dishonesty and a more effective way of determining the integrity of academic work. But this tool alone is not the answer to the manifold problems that AI has raised within educational institutions. What remains more important for educators, administrators, and institutions is developing clear, reasonable, and informed standards for the acceptable use of AI, instructional methods that are adaptable to this new age of technology, and assessments that set students up for success rather than inviting further confusion and speculation about academic integrity.