In August 2026, MIT's own AI committee admitted what many teachers had long suspected: AI can now credibly complete most undergraduate assignments. Essays, problem sets, proofs, code. That is not the end of assessment. It is the end of pretending the old model still works. And if it holds at MIT, it holds even more in a primary school classroom in Brno, Bologna or Bordeaux, where the tasks are simpler.
The real news isn't the report itself but what it reveals. Assessment, which we treat as the final checkpoint of the educational process, is actually its compass. Students, parents, schools, ministries and employers all steer by it. Right now, the needle is spinning in AI’s magnetic field.
For most of history, assessment meant an experienced person judging whether you were ready. China's imperial examinations, from the 7th century until 1905, selected the empire's bureaucrats through essays written under guard. (They also fuelled a thriving trade in miniature cheat books...) Medieval universities used oral disputation: you argued, a master judged.
Mass schooling made that impossible to scale, so the written exam took over. When examiners turned out to disagree wildly about the very same essay, the 20th century answered with multiple-choice tests. Then globalisation brought PISA, and assessment began measuring not only students, but teachers, schools and whole nations.
Each model solved the problem of its era: scale, consistency, comparability. None was a neutral mirror of learning. That is why the current one won't survive an era with tools capable of performing tasks that once required human intelligence.
Most school assessment measures products: essays, solutions, posters. The silent assumption was that if the product exists, someone did the thinking behind it. Generative AI has broken that. A product can now exist without the cognition it was meant to prove. The MIT committee calls it a mismatch between what we want students to learn and how we check it.
The instinctive fixes don't help much. Drag everything back into the exam hall (pen, paper, timer) and you signal that speed and recall matter more than scope, rigour and thought. Reach for AI detectors and you get what MIT advises against: tools that can misidentify the writing of non-native speakers and neurodivergent students as machine text, and a classroom run on mutual suspicion. Meanwhile, outside school, almost nobody does knowledge work without AI anymore. Ban it, and you measure a world your students will never work in. Ignore it, and you mostly measure the AI.
The committee even floated a thought experiment: without grades, most of the incentives to cheat with AI would simply disappear. Nobody is abolishing grades on Monday. But it is worth asking what exactly we are protecting.
Europe has something valuable to offer here. The EU's eight key competences (2018) define competence as a mix of knowledge, skills and attitudes (to which, following OECD’s Global Competence, I add values), not a score. Youthpass, the recognition tool of the European youth programmes, doesn't ask what you scored. It asks: How did you plan your learning objectives? Did you learn things that you did not plan or expect to learn? How did you cope with new and unexpected situations? That is competence-based assessment in a nutshell: evidence of what you did in a real situation, plus reflection on how you got there.
MIT now recommends exploring competency- and mastery-based assessment, oral exams and portfolios. Non-formal Erasmus+ education has practised versions of this for years. It is a better compass, not a magic one, though: harder to standardise, slower to assess, and easy to reduce to a reflection form filled in on the bus home.
Traditional homework rests on a fiction: that it shows what a student can do alone. It always measured the home too: a parent, a tutor, Google, Wikipedia. Now there is a free chatbot writing in your student's voice at 11 pm. The student with a quiet room and educated parents hands in something very different from the one sharing a bedroom with two siblings. We grade the difference and call it achievement.
The research never fully backed the tradition anyway. Harris Cooper's meta-analyses found only a weak link between homework and achievement in primary school, a stronger one in secondary, and diminishing returns as the hours pile up.
So don't abolish homework. Flip its function. Make it input: read, observe, collect data, try something and fail. Then bring the thinking into class, where it becomes visible. MIT lands in the same place: out-of-class work paired with in-class conversation. The work leaves the house; the assessment stays in the room.
AI can produce the product. It cannot produce the student's relationship with their own learning. That is where self- and peer assessment stop being soft extras and become core practice.
Self-assessment is metacognition made visible: comparing work against criteria, predicting results, planning what you would do differently. Assessment researchers Gavin Brown and Lois Harris argue it should be treated as a competence in its own right. It is also hard to outsource. A chatbot can write your essay, but it can't honestly tell you what you found difficult about writing the text.
Peer assessment works for a different reason. To explain why a classmate's work is good, you have to understand the criteria better than you did when you only had to meet them. The assessor often learns as much as the assessed, and research suggests a modest but real boost to achievement. It is also a daily rehearsal of cooperation rather than competition.
Neither works on autopilot. Students overrate themselves when criteria are vague, and peer marking slides into friendship marking without clear rubrics, strong examples and some teacher moderation. AI can help as a draft-checking tool, not a grading oracle. Build the loop: self-assessment, then AI feedback, then peer feedback, then a human conversation with the teacher.
The strongest evidence of learning remains the oldest: can you do it, in the real world, for someone who actually needs it? A campaign for a local NGO, a bridge model that must hold a real weight, a podcast interview. Real projects come with feedback built in. The audience yawns, the bridge collapses, the NGO actually uses your poster.
Biology agrees. Learning isn't only an idea changing in your head; it is pathways forming in the brain through doing. A 2024 EEG study found that handwriting engages far broader brain connectivity than typing. A 2025 MIT Media Lab preprint reported that students writing essays with an LLM showed the weakest connectivity and struggled to quote their own work; the authors called it cognitive debt. These are small studies, warning lights rather than verdicts. But they point in the direction every craftsperson already knows.
So I propose a gradual path to achieve a learning objective for the XXIst century tasks:
Hands first. Draw, build, count, write by hand, measure. This lays the neural groundwork, also physical foundations, and the felt sense of a skill.
Supported. Tools that extend the hand rather than replace it: the calculator after arithmetic, the spellchecker after spelling, AI as a sparring partner that questions your draft instead of writing it.
Delegated. Digital tools take over the mundane, repetitive parts, but only once the learner can critically judge what the machine gives back. Delegation without judgement is dependence.
MIT, for all its AI focus, gives surprisingly analogue advice: handwritten early drafts in class, regular check-ins that show a project growing, feedback at every stage and not only on the final product.
Compare that with what we still often assess: speed tests, or texts, designs and calculations produced only with digital tools. Speed of retrieval is exactly what machines already do better than any of us. Why train children to lose that race?
A compass doesn't tell you where to go; it shows where north is. If our north remains "produce correct outputs, quickly, alone", AI has already arrived there before our students, and the needle will keep spinning. If north becomes "grow into a person who can think, make, judge and cooperate", assessment returns to what it should have been all along: not a verdict at the end of learning, but a way of steering it.
The OECD chose the same metaphor for its Learning Compass 2030: students learning to navigate on their own, rather than following fixed directions. Assessment should point the same way. And you don't need to wait for a ministry reform. Next week, swap one take-home essay for a short oral defence or for an elevator pitch. Add a self-assessment step before you grade. Let a project be judged partly by the people it was made for. Small turns of the dial. That is how you recalibrate a compass.
Łukasz W. Kosowski