Back to Blog
Emerging Skills

What Is Reflective Learning With AI, and Why Institutions Are Asking for It

17 min read

Reflective learning with AI defined from Dewey, Flavell and recent studies, what institutions such as the OECD ask for, and design questions for a programme.

Key Takeaways

  1. Reflective learning is looking back at what you did and why, soon enough to change the next attempt. This page builds that definition from Dewey, Flavell and the book Make It Stick. It is a working definition and not an official one.
  2. AI can raise performance on a task while a person's judgement of that performance drifts. In one study of logical reasoning problems, participants scored about three points better with AI and overestimated their own score by about four.
  3. How AI is set up decides what happens to learning. In a field experiment with nearly a thousand high school maths students, unrestricted GPT-4 access raised practice grades and lowered exam grades once access was taken away. A tutor version that gave hints instead of answers largely removed that drop.
  4. Institutions are asking for judgement and metacognition by name. The OECD's 2026 Digital Education Outlook says a task done well with generative AI does not automatically lead to learning. Indian bodies such as AICTE and the Ministry of Education ask mainly for AI courses and skills.
  5. A reflective design keeps two jobs apart. The AI supplies a likely answer, and the student states a judgement first and compares afterwards.
  6. The evidence is short-term and specific, and none of it shows that reflective designs raise marks. Nothing here promises a result.

In a paper in Computers in Human Behavior (February 2026 issue), participants used AI to solve 20 logical reasoning problems from the Law School Admission Test. In the first study, with 246 participants, task performance improved by three points compared with a norm population. The participants also overestimated their performance by four points. The second study, with 452 participants, replicated these findings. The authors also report that higher AI literacy correlated with lower metacognitive accuracy.

A field experiment in a Turkish high school shows a similar gap in a classroom. Nearly a thousand maths students practised with either no AI, a ChatGPT-style GPT-4 tool, or a tutor version built to give hints. Practice grades rose with both tools. On the exam, taken without AI, the group that had used the ChatGPT-style tool scored 17 percent lower than the group that never had it, according to the PNAS paper. The authors' analysis of student surveys found that students did not perceive any reduction in their learning from copying solutions.

Both studies are narrow. The first used a set of 20 problems and the second a single school and subject. Read together, they say that getting a task done and knowing how well you did it can separate when AI is in the room. A course that adds AI decides, whether or not it means to, what happens to the second of these. Reflective learning with AI is one way of making that decision on purpose.

What reflection and metacognition mean

John Dewey defined reflective thought in How We Think (1910) as "active, persistent, and careful consideration of any belief or supposed form of knowledge in the light of the grounds that support it, and the further conclusions to which it tends". The useful part for a classroom is the phrase about grounds. Reflection asks what supports a belief, and that needs something outside your own feeling of being sure.

John Flavell's widely cited 1979 paper in American Psychologist is a standard starting point for the word metacognition. He describes metacognitive knowledge as what a person believes about themselves and others as thinkers, about tasks and about strategies. He describes metacognitive experiences as the conscious feelings that occur during a task, often about how well it is going. In plain words: what you believe about how you learn, and the sense you have, halfway through, of whether it is working.

The CHI 2024 paper discussed below describes metacognition as the ability to monitor and control one's thoughts and behaviour. Monitoring is noticing how a task is going. Control is changing what you do because of it. Reflection is the slower version of both, done after the attempt.

The book Make It Stick, by Peter Brown, Henry Roediger and Mark McDaniel, describes reflection as a few minutes spent reviewing a recent class or experience and asking questions such as what went well and what could have gone better. It builds the routine from three practices discussed elsewhere in the book. Retrieval is recalling what happened. Elaboration is connecting it to what you already know. Generation is putting ideas in your own words or mentally rehearsing what you would do next time. The book gives the example of a biology professor who sets weekly low-stakes "learning paragraphs" in which students reflect on the previous week. This is a description of a habit, and this page does not treat it as a proven programme.

The same book explains why an outside check matters. It describes calibration as aligning your judgements of what you know with objective feedback, and it says growing familiarity with a text can come to feel like mastery of its content. The research behind that is covered in why you forget what you study.

A working definition for this page follows from these sources. Reflective learning is looking back at what you did and why, against something outside your own sense of knowing, soon enough to change the next attempt. Reflective learning with AI adds one condition: the AI's output is one of the things looked at, and the person's own earlier attempt is another.

What changes when AI is in the loop

More to monitor

Lev Tankelevitch and colleagues argued at CHI 2024, the ACM's human-computer interaction conference, that generative AI (GenAI) systems impose metacognitive demands on users, "requiring a high degree of metacognitive monitoring and control". They list prompting, evaluating outputs, relying on outputs and organising workflows as places where the demand shows up. Microsoft Research lists the paper as a Best Paper at that conference.

This is an argument built from psychology and earlier user studies, and its support strategies are proposals. For a student, each demand is a judgement: what to ask, whether the answer is good, and how far to lean on it. Each can be made badly without any visible sign.

Performance without learning

The PNAS paper by Hamsa Bastani and colleagues randomly assigned classrooms in one Turkish high school to a control group, a "GPT Base" tool that resembled ChatGPT, or a "GPT Tutor" tool with teacher-designed prompts. Practice performance improved by 48 percent with GPT Base and 127 percent with GPT Tutor. When access was removed, the GPT Base group scored 17 percent lower than the control group. The safeguards in GPT Tutor largely removed that drop, although the paper reports no positive effect on the exam either. Students with GPT Base often used it to obtain solutions, which the authors call a "crutch".

The tutor students perceived that they had performed significantly better than the control group, although their exam scores were no higher, according to the full text. PNAS published a correction to the paper in August 2025. It concerns one author's affiliation and does not change the results.

Yizhou Fan and colleagues report a randomised study in the British Journal of Educational Technology with 117 university students revising an essay. The group with ChatGPT, which was set up to give advice and not to write the essay, improved its essay scores more than the other groups. The four groups did not differ significantly in knowledge gain or knowledge transfer. The authors write that AI may promote dependence and potentially trigger "metacognitive laziness". They also note that the study had no dedicated measure of it, so the phrase describes a possibility raised by the pattern and not something measured.

The OECD's Digital Education Outlook 2026 cites the PNAS experiment and uses the phrase metacognitive laziness. It is a summary of studies and not a separate confirmation.

What people report about effort

Hao-Ping Lee and colleagues surveyed 319 knowledge workers for CHI 2025. Higher confidence in GenAI was associated with less critical thinking, and higher self-confidence was associated with more. These are self-reports from people at work, recruited online, who used GenAI at least weekly. They are not students, and an association in a survey cannot show that AI causes less thinking.

Judging the output

Bearman and colleagues argue in Assessment & Evaluation in Higher Education that students need evaluative judgement, "the capability to judge the quality of work of self and others". They propose three targets: judging AI outputs, judging AI processes, and having AI assess a student's own judgements. They argue that existing formative assessment can interrupt uncritical use. This is a conceptual paper, so it offers an argument and no experiment.

Why institutions are asking for it

The OECD's Digital Education Outlook 2026, released on 19 January 2026, says emerging evidence suggests that general-purpose GenAI tools can enhance students' performance on tasks without necessarily leading to learning gains. It says tools designed or used with an intentional pedagogical purpose tend to show sustained improvements in learning. Its companion Insights document advises teachers to scaffold student use of GenAI with drafts, prompts and reflections. It advises students to use GenAI as a learning partner and to balance it with independent practice without GenAI.

The EC-OECD AI literacy framework, published in 2026 for primary and secondary education, warns that without guidance learners may rely on AI in ways that diminish reflection, persistence or independent reasoning. UNESCO's AI competency framework for students sets out 12 competencies across four dimensions, one of them a human-centred mindset, and emphasises critical judgement of AI solutions. It is written for school curricula, so applying it to higher education is an extension.

India shows a different emphasis. The All India Council for Technical Education (AICTE) declared 2025 the Year of AI for approved institutions, and the Ministry of Education set up a Centre of Excellence in AI for Education at IIT Madras. In May 2026 a meeting on revamping the AI curriculum listed a separate workstream for non-STEM disciplines covering AI awareness, foundational AI literacy and applied use in non-technical roles. A survey study in Humanities and Social Sciences Communications, run between May 2023 and April 2024 with urban Gen Z students and some older participants, found that students were motivated mainly by expectations of AI's performance, its ease of use and the social context. It measured adoption, not learning.

These Indian items concern teaching AI, building courses and adoption. The request for reflection, judgement and metacognition is stated more directly in the international guidance. The two are related, because a course in using AI raises the question of what the student does with the output.

Who decides and what the AI predicts

Economists Ajay Agrawal, Joshua Gans and Avi Goldfarb interpret recent AI as an improvement in prediction. Their model treats judgement as a separate input, and their abstract says judgement is costly and that prediction and judgement are complements as long as judgement is not too difficult. In a later paper they describe judgement as what is exercised when the objective cannot be coded. Their setting is decisions in organisations and not learning.

Borrowed for a classroom, the lens looks like this. A chatbot's answer is a prediction of what a good answer looks like. Whether it is good for this case, this marking scheme or this client is a judgement, and the student is the one who must make it. A tool that answers first and sounds sure joins the two, and the student may never exercise the second. That reading is this page's own, and none of the studies above tested it. It fits the "crutch" behaviour in the PNAS experiment, and it fits Bearman and colleagues' concern that humans should stay the arbiters of quality.

QuestionChatbot that answersReflective design
Who states the first judgementThe AI, in its replyThe student, before opening the AI
What the AI providesThe likely answerA comparison point, such as a critique or a counter-argument
What is keptUsually the final textThe attempt, the AI output, what changed and why
What a teacher can seeThe finished workThe reasoning as well

Design questions for a programme

The list below is a practical framework assembled from the sources above and Make It Stick. It has not been tested as a package.

  1. Where does the student commit first? A written answer, a prediction or a confidence level, recorded before the AI is opened.
  2. What does the AI provide? A critique, a counter-argument or a reading of the marking scheme gives the student something to compare against. A finished answer gives something to copy.
  3. How is the output checked? Against a source, a worked solution, a peer or a teacher. Bearman and colleagues call this evaluative judgement.
  4. When is the AI off? Some practice done without it shows the student what they can do alone. The PNAS exam and the OECD's advice on independent practice both point here.
  5. Do students meet AI errors on purpose? A planned flawed output that they must catch gives practice in the judgement Tankelevitch and colleagues describe.
  6. What do students record about their own thinking? The attempt, the AI output, what they changed and why, and how confident they were before and after.
  7. Where does that record live? If it can be retrieved weeks later, the student can look back. If it is lost, the reflection was a one-time exercise.
  8. What is assessed? If only the final artefact is marked, a student has every reason to outsource it. Marking the reasons as well changes the incentive, though graded reflection can also teach students to write what a marker likes.
  9. Who decides, and is it written down? The decision and the reason, in the student's own words.

A session built on these could run as follows. The student attempts the problem or drafts a first paragraph and writes down how confident they are. They then ask an AI for a critique or a counter-argument and not for a solution. They compare the two and write what they would change, what the AI got wrong and why. A week later they try to reproduce the reasoning without either. This follows the retrieval, elaboration and generation practices in Make It Stick, and it is a sketch to adapt.

How to tell whether reflection is happening

Observable signs are more useful here than scores. None of the following has been validated as a measure, and each is a prompt for a conversation with the student.

  • A first attempt exists before the AI's output, and the student can say how the two differ.
  • The student gives reasons for keeping or changing something, and the reasons refer to evidence.
  • The student names an AI error they caught and how they caught it.
  • Stated confidence and results can be set side by side across several weeks, so the student can see whether they tend to run high or low.
  • The reasoning can be reproduced without the tool, in a closed-book setting.
  • Questions change from "what is the answer" towards "what would make this wrong".
  • Notes refer back to earlier notes, which shows that someone is reading their own record.

The announcement

On 25 September 2026, Gradeless AI announced a partnership with Tata Indian Institute of Skills, under which students and faculty at the institute receive access to Rehearsal and coaching on using generative AI in learning (read the announcement). Tata IIS describes itself on its own site as an initiative of the Tata Group that, in partnership with the Ministry of Skill Development and Entrepreneurship, is setting up IIS Mumbai and IIS Ahmedabad. The research above stands on its own and does not draw on the announcement.

Where Rehearsal fits

This article is published by Rehearsal, which makes the app described in this section.

Everything above can be done with a notebook and a phone. The one thing reflection needs is something to look back at. Rehearsal is a place where a person's own recordings, notes and saved material are kept so they can look back at their own thinking. It records lectures and meetings and keeps the recordings with playback and transcripts. It holds PDFs, screenshots and links, and selected files, including from WhatsApp and Telegram, can be forwarded into it.

A student who follows the session above could keep each step there: the first attempt as a voice note, the AI's output as a saved file and the reasons as a note. The MCP (Model Context Protocol) connector lets ChatGPT and Claude read the material the student chooses to save, so a question put to an AI can be answered from the student's own records. Asking across saved material and the connector are Unlimited features, and there is a free plan.

Rehearsal is a capture-and-recall habit. It does not coach you, score you or predict a result, and nothing on this page says it improves marks or any other result. An app that keeps records does not make anyone reflect. The person still has to look. The phrase "second brain" is explained at second brain app, and the connector at Rehearsal MCP.

Common misconceptions

"Reflection is a diary entry at the end of term." Dewey's definition asks for consideration of grounds and consequences. Make It Stick describes a few minutes soon after the class or experience, with questions that involve retrieval.

"Reflective learning means using less AI." The OECD advises balancing AI-assisted work with independent practice and using AI as a learning partner. In the PNAS experiment, the same model with hint-giving safeguards largely avoided the exam drop.

"AI literacy is the same as calibration." The EC-OECD framework for primary and secondary education describes AI literacy as knowledge, skills and attitudes that let learners engage, create with, manage and shape AI while critically evaluating its benefits and risks. The Computers in Human Behavior paper reports that higher AI literacy correlated with lower metacognitive accuracy in its samples. The abstract does not say how AI literacy was measured, so read the full paper before generalising.

"A chatbot can do the reflecting for you." It can write a reflection. In the retrieval, elaboration and generation account in Make It Stick, the value comes from the person doing the recalling and the rephrasing. That is reasoning from the book, and no study cited here tested AI-written reflections.

"If it feels clear, I know it." In the PNAS experiment the tutor group felt they had done better than their exam scores showed, and in the Computers in Human Behavior paper participants overestimated their scores.

Questions people ask

What is reflective learning?

Reflective learning is looking back at what you did and why, against something outside your own feeling of knowing, soon enough to change the next attempt. The definition on this page draws on Dewey, Flavell and Make It Stick.

What is reflective learning with AI?

It is reflective learning in which the AI's output is one of the things looked at, and the person's own earlier attempt is another. The person states a judgement before the AI's answer, then compares.

What is metacognition?

Flavell describes it as knowledge about how you and others think, about tasks and about strategies, together with the conscious experiences you have during a task about how well it is going. In practice it means monitoring how a task is going and changing what you do.

Does AI make students think less?

The studies cited here give a mixed and specific picture. In one Turkish high school, unrestricted GPT-4 access was followed by lower exam scores. In one essay study, ChatGPT users improved their essays more, with no significant difference in knowledge gain or transfer. In a self-reported survey of workers, higher confidence in GenAI went with less critical thinking. None of them shows that AI use causes less thinking in general, and the hint-giving tutor result suggests design matters.

Is reflective learning the same as AI literacy?

They overlap without being the same. AI literacy in the EC-OECD framework includes engaging with, creating with, managing and shaping AI. Reflection concerns what the learner did with the output and how well they can judge their own work.

Why are institutions asking for it?

The OECD reports that task performance gains with general-purpose GenAI do not necessarily lead to learning gains, and it advises scaffolding student use with drafts, prompts and reflections. UNESCO's student framework emphasises critical judgement. The Indian documents cited above ask mainly for AI courses and skills.

How can a teacher tell whether it is happening?

Look for a first attempt before the AI's output, reasons for changes, AI errors the student caught and reasoning that can be reproduced without the tool. These are prompts for conversation and have not been validated as measures.

Does it need special software?

No. A notebook and a rule about writing the attempt first will do. Software helps when the records need to be searched later.

Does it raise marks?

No study cited here shows that a reflective design raises marks. The evidence on hints versus answers and on independent practice is suggestive, is short-term and comes from specific settings.

Where to go next

Tags

reflective learningmetacognitiongenerative AI in educationAI literacyhigher education India

Reading ≠ Speaking

Put what you learned into action with AI-powered mock interviews — free to start.

Start A Rehearsal — Free