Cheating is not the interesting part of this story. The interesting part is that a language model passing our assignments tells us those assignments were measuring the wrong thing all along, and had been for decades.
I say this as someone who studies agent architectures for a living. When a system can produce the artifact we grade without possessing the capability we claim to be grading, that is not a student integrity problem. That is a specification problem. We wrote a bad reward function and got exactly what we asked for.
What Actually Broke
The shift teachers are reporting is consistent and specific. Written reflections, once attached to nearly every assignment, have moved to in-person conversations. One instructor now requires each student to schedule a fifteen-minute session with a TA after every assignment. Oral exams are back. Projects and live discussion have replaced take-home written products as the primary evidence of learning.
Read that as an engineering response, not a moral one. The old assessment pipeline had a single observable: the submitted document. Everything the instructor believed about a student’s understanding was inferred from that one output. Generative models made the output cheap to produce independently of the understanding, and the entire inference chain collapsed.
The fix teachers landed on is the same fix we reach for in agent evaluation when a benchmark gets gamed. You stop scoring the final artifact and start scoring the trajectory. You add observability. You ask follow-up questions the system cannot have pre-computed.
Homework as Process, Not Product
Teachers describe this transition as treating homework as a learning process rather than a deliverable. That phrasing is doing more work than it appears to.
Consider what a fifteen-minute TA conversation actually measures that a written reflection never could:
- Depth under perturbation. Change one assumption in the problem and watch whether the student’s reasoning bends or shatters. Written submissions are static; conversations are adversarial probes.
- Latency and confidence calibration. A student who understands something answers differently than one reciting it. Hesitation patterns carry signal.
- Repair behavior. When you point out an error, does the student localize it, or rebuild from scratch? This is the single most diagnostic behavior I know of, and no document captures it.
These were always the things we cared about. We proxied them with essays because essays scaled and conversations did not. AI removed the proxy’s validity, and now we are paying the cost we deferred for fifty years.
The Error Analysis Move
One approach I find genuinely sharp: assign error analysis tasks where students identify and correct mistakes in AI-generated solutions. Another is asking students to design their own problem in the style of class material, then solve it.
Both invert the relationship between student and model. Instead of the model producing and the student submitting, the model produces and the student adjudicates. That requires a strictly stronger capability than generation does. You cannot evaluate a chain of reasoning you do not understand. Critics of a domain need more of the domain than producers of average work in it.
Problem design is even more demanding. Generating a well-posed question that is solvable, non-trivial, and stylistically consistent with a body of material requires a model of the material’s structure, not just its surface. This is the difference between pattern completion and having a world model, and it is roughly the same line we argue about in agent research.
Where This Gets Expensive
I want to be honest about the tradeoff, because the optimistic version of this story ignores it. The instructor who made this change described abandoning evidence-based pedagogy practices and redoing most assignments to get there. That is not a small tax. Oral assessment does not scale. Fifteen minutes per student per assignment in a class of two hundred is fifty hours of TA time per cycle.
So the honest framing is that AI did not improve assessment. It made bad assessment visible and forced a switch to better assessment that costs dramatically more. Whether institutions fund that cost is an open question, and the failure mode is obvious: schools that cannot afford conversation-based evaluation will keep grading artifacts and quietly stop believing their own grades.
The Transferable Lesson
Every field that evaluates intelligence through produced artifacts is about to have this same reckoning. Hiring, code review, peer review, professional certification. Each one has a submitted-document choke point, and each one is now measuring something other than what it thinks.
The teachers who moved to conversations got there first, under duress, without a research program behind them. They arrived at process supervision and adversarial probing because the alternative was grading nothing at all. That is a useful result, and it came from practitioners rather than labs.
Worth taking seriously before the rest of us are forced into the same corner.
đź•’ Published: