Skip to content

September 5, 2026 · Edition #97

The questions disappeared.

AI made help faster. It also moved confusion somewhere managers cannot see.


A few years ago, I built an AI tutor for the university where I teach.

The first result that made me happy was simple: fewer questions reached professors.

That sounds terrible when I write it like that.

But the tutor was solving a real problem. Students could ask questions about their courses, get an explanation at midnight or work through something they did not understand.

A working student should not have to wait two days because a link is broken or a definition in a lecture is unclear. Professors should not spend Monday morning answering the same basic question twenty times either.

That is a worthwhile result. Looking at it now, though, the quieter inbox can mean two very different things. Students may understand the course better, or they may still need the same help and ask for it somewhere the teacher cannot see.

Those situations look identical on a support dashboard.

If you manage junior team members or run AI assistants inside your organization, you're already running the same experiment. Their questions are moving from meetings and Slack into private AI conversations, and the silence can look a lot like independence.

The quiet classroom

We treated student questions like support tickets. A question arrived, someone answered it and the queue got shorter.

That works as an operational metric. It tells us very little about learning.

When a student asks me a question, I learn something too. I see the step they skipped or the assumption underneath the wrong answer. Sometimes five students ask what appear to be different questions, and I realize the problem is my lesson.

The question asks for help, but it also shows the teacher where the learning broke.

An AI tutor handles the first part very well. The student gets an immediate answer, can ask without feeling embarrassed and does not have to arrange their confusion around a professor's calendar.

The second part is easier to lose.

The conversation happens in private. My inbox improves, the student moves on, and I may teach the same weak explanation again next term because I never saw the pattern.

Those questions are useful feedback for anyone responsible for helping someone improve. When AI absorbs them, that feedback loop gets weaker.

I had built something that made help more available to every student while making each student's struggle less visible to the teacher.

That sentence has been bothering me.

What happened after the mistake

Two large studies published in August made the problem much clearer.

The first followed Khanmigo, Khan Academy's AI tutor, across 18 middle schools for two years.

Almost every student tried it. Yet the median student used it on only about a third of the days they practised. When they made a mistake, they opened the tutor in just 17 percent of those sessions.

Students assigned access improved modestly, but their gains looked much like ordinary Khan Academy practice without the AI tutor.

The tool was available. Students simply did not use it very often at the moment when help might have mattered most.

A second experiment put another AI tutor directly inside the practice loop for more than 6,000 maths students.

This group moved more slowly. They attempted fewer questions and spent longer on the ones where the AI helped them.

That would look worrying on a dashboard built around progression and volume.

Then you look at what happened after an error. Their next attempt was more likely to be correct, and they needed fewer tries to recover.

Did that improvement last? Maybe. The delayed test found a small gain in one part of the material when AI sat inside the mastery workflow. It was only marginally significant, so I would not build a theory around it.

Still, it showed me where to look.

We keep judging an AI tutor by the quality and speed of its answer. The more interesting moment comes next.

What did the student do after the model replied?

The answer is not the end

Most AI products try to shorten the distance between a request and an output. That makes sense when you are searching for a document or formatting a report.

Learning is awkward because the shortest path to the output can skip the thing the student came to build.

If a tutor gives someone the right formula in three seconds, the answer looks excellent. If the student first chooses the wrong formula, understands why it failed and then solves a slightly different problem, the session looks slower.

The second session tells us much more about what changed in the learner.

I am not arguing that every difficulty is educational. Some friction is simply bad design. Students learn nothing while waiting forty-eight hours for a basic explanation, searching for a missing file or recopying work they already understand.

The useful distinction is where the AI saves time.

It can remove the waiting around an attempt. The danger begins when it removes the attempt itself.

Education researchers have studied this for years through scaffolding, fading and the gradual transfer of responsibility from teacher to learner. None of that is a new AI discovery.

What is new is the scale. We can now put a patient tutor within reach of every student, while the product team watches a dashboard built largely around speed and use.

If I rebuilt our tutor today, I would make the student's next move part of the product.

Routine questions would still get a direct answer. There is no educational prize for making someone struggle to find the right chapter.

For a question tied to the learning itself, the tutor would first ask for a small attempt. The student could share a prediction, the formula they think applies, three lines of code or the exact point where they became unsure.

The tutor would then respond to something real. It might explain the missing idea, offer a hint, show an example or challenge the student's reasoning.

The session would continue until the student tried again.

That changes what we can measure. We can record what the learner attempted, how much help the tutor supplied and what the learner could do afterward. A few days later, an unseen problem using the same skill at a similar level of difficulty would tell us what carried over.

This would not produce one magical learning score. I distrust those even more than I did before writing this letter.

It would give us a record of how responsibility moved during the work.

That matters because better students do not necessarily use less AI. They often ask harder questions. Their usage can rise while the tutor does less of the thinking for them.

The same problem at work

Most readers of this letter do not run universities. But almost every company is running a version of this experiment now.

People ask AI questions they used to ask a colleague. A new analyst can query an internal assistant about the forecast. A salesperson can ask it what a policy means. A junior engineer can paste in an error rather than walking over to someone senior.

This can be genuinely useful. The employee gets help immediately, while the experienced person gets a little more time to do their own work.

It can also create a false impression of independence.

The new hire asks fewer questions in meetings, so the manager assumes the onboarding is working. Meanwhile, the same confusion may be repeating every day inside a private chat window.

Questions used to tell a manager where the documentation was weak, which part of the process nobody understood and where a new hire kept making the same mistake. When those questions disappear, the employee may move faster while the company loses a useful early-warning system.

Your internal AI assistant is becoming a private tutor whether you designed it that way or not.

It should help the person in front of it, but it should also return the pattern to the organization. Which policy keeps confusing the sales team? Which error returns after the explanation? Where are people repeatedly asking the AI to take over rather than help them make the next move?

That does not require a manager reading everyone's private conversations. The system can surface aggregated patterns, with a small and consented sample used to check whether its classifications are accurate.

The same principle applies in a university or a company. AI should help the individual without making the organization blind to where people are struggling.

What I would measure now

I still care about fewer questions reaching professors.

Student waiting time matters, and so does faculty time. I do not regret solving either problem.

The mistake was asking an operations number to answer an education question.

The quiet inbox told me how much work had left the professor's desk. It could not tell me what had stayed with the student.

An AI answer shows us what the model can do.

I would now stay long enough to see the student's next attempt.

Have a great weekend.

Stay sharp.

— Charafeddine (CM)


↑ All editions Older →
Charafeddine Mouzouni — AI Scientist and Founder

Start with one email.