Earlier this spring, it seemed all of higher education was talking about Einstein, a tool specifically marketed to students for its ability to autonomously log on to the learning management system Canvas to complete assignments. As agentic AI tools built on systems like Claude Code and Codex began demonstrating the ability to navigate coursework with minimal human input, the conversation in higher education defaulted to prevention. Can we detect these systems? Can we block them? Can we AI-proof assignments?
Those questions are missing the point. If an AI agent can complete an assignment from start to finish, the problem may not be the AI. It may be the assignment (and the classroom culture).
Einstein itself may have been short-lived—its creator took down its website after receiving a cease-and-desist letter from Canvas’s parent company— but with agentic AI here to stay, the bigger, more important questions remain relevant. If a student can outsource classwork entirely to an autonomous agent, what exactly are we measuring in the first place? What kinds of learning activities still make sense when AI can generate polished outputs on demand?
My answer has been to stop grading outputs and start grading thinking.
I want to be direct about something: My goal is not to design assessments that outsmart AI. That is a losing battle, and I have no interest in fighting it. Instead, I design courses where the only thing worth grading is the thinking process itself—and where submitting a polished final answer, however good it looks, is just not enough.
In practice, this means designing courses where the thinking process is visible and where the final answer alone cannot carry the grade. Here are five approaches I use.
Make Participation Count (Literally)
In my biochemistry courses, 25 percent of the final grade comes from in-class debates and participation. That is not a soft credit. It is a quarter of the grade.
This is deliberate. Real-time engagement—being present, responding to a peer’s argument, defending a position under questioning—is something no automated system can do on a student’s behalf. A student cannot send an agent to sit in their seat and argue a position in a live debate. The moment you make live interaction a substantial portion of the grade, the incentive structure shifts. This works for online classes, too.
A Day With Only Markers and Pens
Once a semester/quarter, I give my students a peer-review day with a simple rule: No devices. Highlighters and pens only.
The goal is to teach students to be auditors of each other’s thinking. Reading a peer’s work carefully, marking it up by hand and explaining your feedback out loud requires a level of engagement that clicking through a screen does not. Students learn to catch weak reasoning, spot missing evidence and articulate why something doesn’t land. That skill (being a rigorous reader) is one I want them to carry out of the course.
Reflection Assignment on “Beating the AI”
I ask my students to reflect on AI-generated outputs using a metacognitive scale I built specifically for this purpose. The core question is: Is your own thinking better than what the AI produced, and can you show me why?
This is harder than it sounds. It requires students to compare their own work to the outputs of the AI, read the AI output critically, locate its weaknesses and articulate where their own reasoning (informed by the course, by discussion, by their own background) goes further. The reflection is not about whether the AI was right or wrong. It’s about whether the student can demonstrate something the AI cannot: a perspective that’s actually theirs.
Tie Assignments to Unpredictable Classroom Moments
For years, I have built in classroom moments that an AI agent will definitively not know.
Here is one I use regularly: I ask students whether Dolly the sheep was a true clone. Most people—and most AI systems—will tell you the answer is yes. But an AI agent like Claude might say it is both yes and no. I want my students to say no because I teach about mitochondrial DNA, so for my class the reasoning requires understanding something about mitochondrial inheritance that only makes sense if you were in the room when we discussed it.
You may have your own version of things you say to students who were present, things you emphasize in the classroom. These are not trick questions. They are anchored in the specific intellectual history of the course—things that happened in a live discussion, examples I drew on the board, arguments that emerged from a debate. An AI system that was not present in the room cannot reconstruct that context from a prompt.
Shift Grading Weight Toward Process
In my courses, I regularly ask students to do “chalk tasks”—drawing out a biological process, diagramming a research design or explaining a concept in their own words on the board. The finished drawing is not the point. The point is watching how they think through it.
Could a student theoretically use an AI agent to generate an explanation and then copy it onto the board? Technically yes. But consider what that actually requires: getting the agent to produce something accurate, understanding it well enough to reproduce it and then defending it in real time when I ask follow-up questions. At that point, the “shortcut” has become longer and more demanding than just learning the material. That’s by design.
Assignment Design Alone Isn’t Enough
I want to be honest about something else: You can redesign every assignment in your course and still fail to solve this problem if the surrounding system is working against you.
Students are rational actors. If the primary goal is grades, and the system rewards efficiency, and AI dramatically increases efficiency, then using AI agents is simply the logical strategy. The problem is not that students are lazy or dishonest. The problem is that we have built systems that make shortcuts make sense.
That’s why I think about this in three layers.
Assignment design is the first layer: Build work that makes thinking visible, that requires trade-offs, that creates constraints AI cannot easily navigate.
Course culture is the second: Normalize talking about AI openly. Make intellectual process something the class values together, not just something the instructor enforces alone. My syllabi include explicit language about AI agents because I want my students to understand what these tools actually do, and what it means (ethically and practically) to use them.
Institutional incentives are the third layer, and the hardest: Even when individual faculty members redesign their courses, departments often reward efficiency, accreditation systems demand standardization and professors scale grading in ways that inadvertently undermine changes at the course level. This is a structural problem, and it requires a structural conversation.
The Bigger Picture
The rise and fall of AI detection tools revealed something important: No single intervention—detection, redesign or policy—can hold academic integrity together on its own.
Still, what I keep coming back to is a more fundamental question: What are we actually measuring when we grade student work?
If the answer is “the final product,” AI will keep challenging that model, and we will keep losing. But if we focus on reasoning, live interaction and intellectual development—on the things that are genuinely, stubbornly human—the presence of AI becomes far less threatening.
AI agents can produce answers. They can generate polished paragraphs, solve problem sets and navigate learning management systems. But learning has never been only about answers. It has been about the process of forming them, defending them and revising them in the presence of other people.
That process is still ours to design.
Tina Austin is the author of UnBlooms, a framework that rethinks Bloom’s taxonomy for AI-era learning. She is also a lecturer in computational biology at the University of California, Los Angeles, and works with faculty on designing learning experiences with, without and against AI to improve learning outcomes.