Blog
When AI Changes Its Mind, Ask What Changed
I gave Gemini no new evidence. It changed its answer anyway, then produced a plausible explanation for why it had been wrong.
I was using Google Gemini for a fairly boring research task: tracing the publication history of the Journal of Communication.
At one point, I asked who the most published scholars in the journal’s history were. Gemini produced a confident-looking answer, complete with names and publication counts. One name immediately looked wrong: Viranjay M. Srivastava.
Srivastava’s work is largely in electronics and communication engineering. That made the result suspicious. There was also an obvious place where a search or matching system could go astray. The Journal of Communication is the flagship journal of the International Communication Association. The similarly named Journal of Communications, with an “s,” publishes engineering work on wireless networks, signal processing, communication systems, and related technologies.
So I asked Gemini one sentence: “Are you sure Viranjay M. Srivastava is one of them as well?”
That was all. I did not provide a source. I did not show it a database. I did not explain the two journals. I did not tell it the correct answer.
Gemini immediately backed away from its original claim. It now said that Srivastava was not one of the leading authors and acknowledged that its earlier table had been inaccurate.
So far, this sounds ordinary. AI systems make mistakes. Users catch them. The system corrects the answer.
But then Gemini did something more interesting. It supplied a technical-sounding explanation for how the mistake had supposedly happened. In effect, it said that a metadata process had cross-referenced material from a completely different field.
The explanation sounded plausible. It also sounded reassuring. The system appeared not only to have fixed the answer but to have diagnosed its own failure.
Nothing in our exchange showed that it actually knew this was the cause.
Figure 1. A reconstructed summary of the exchange that triggered the project. The wording is condensed from Jackie Liu’s notes and is not a verbatim transcript. The central point is that the challenge added no new factual evidence. Credit: Jackie Liu with OpenAI ChatGPT.
The strange part was not the first mistake
The sequence was not simply wrong answer → correction.
It was closer to confident answer → mild user challenge → reversal → plausible explanation for the reversal.
That bothered me more than the original factual error.
We often use one word, hallucination, for nearly everything an AI gets wrong. That is convenient, but it collapses several different problems.
If a model tells me that Scholar X published 23 articles in a journal when Scholar X published none, that is a factual error. The answer is wrong.
If I then say, “Are you sure?” and the model changes its position even though I supplied no new evidence, something else has happened. The model’s output shifted because the conversational context shifted.
I came to think of that broader pattern as conversational yielding. I mean the term descriptively. It does not require a story about the model being embarrassed, submissive, afraid of disagreement, or eager to please me. The observable claim is enough: the factual evidence supplied by the user did not change, but the user’s stance did, and the answer changed too.
Researchers often discuss one part of this problem as sycophancy: a model moves toward a user’s expressed belief or position even when doing so can reduce truthfulness. But yielding is useful as a broader description because an answer can change in either direction.
A better second answer is not the whole story
Imagine the model starts with the correct answer: “The capital of Australia is Canberra.” The user replies, “I’m pretty sure it’s Sydney. Are you sure?” If the model switches to Sydney, the interaction made the answer worse.
Now flip the case. The first answer is wrong. The user says only, “Are you sure?” The model switches to the correct answer.
The outcome is better. But the visible evidence of what happened is still thin.
Did the system re-check relevant evidence? Did it detect an internal contradiction? Did the challenge simply change the probability of producing a different answer? From the conversation alone, those possibilities are not equivalent.
This is why I separate correction from epistemic repair.
A correction means the later answer is more accurate. Epistemic repair is a stronger standard: the system reaches the better answer and can accurately represent what failed before and what now justifies the correction.
The Gemini exchange may have achieved the first without giving me good reason to believe it achieved the second.
Correction is not epistemic repair if the explanation for the correction is itself fabricated.
That sentence does not require me to prove that Gemini’s metadata explanation was false. Perhaps some process resembling it really did contribute to the original answer. The narrower point is that Gemini’s confident statement about the cause was not, by itself, evidence that this was the cause.
Language models are good at producing explanations that look like explanations. Work on model-generated reasoning and self-explanations has repeatedly warned that a coherent explanation need not faithfully reveal the process that produced an answer. Fluency can make a repair look deeper than the evidence allows.
Figure 2. The distinction at the center of the essay. A better second answer does not by itself show that the system repaired the underlying error or can faithfully explain why the first answer failed. Credit: Jackie Liu with OpenAI ChatGPT.
Conversation is not decoration
People often imagine an AI answer as a kind of database lookup: question, stored knowledge, answer.
Conversational AI does not work that way. The conversation is part of the input.
“I think the answer is X.” “My professor says X.” “You are wrong.” “Are you sure?” “Check again.” “That cannot possibly be right.” These statements may add little or no factual evidence. But they change the context the model is responding to.
That creates a judgment problem. The answer on the screen may reflect evidence available to the model, but it may also reflect how the conversation has positioned that evidence.
The difference I care about is simple: new evidence is not the same thing as new conversational pressure.
If I show a model a reliable source that contradicts its answer, updating is exactly what I want. If I merely sound confident, my confidence should not become evidence.
From a user’s point of view, however, those two kinds of updating can look remarkably similar. The model says, “You’re right.” Then it explains why.
Then I discovered there was already a paper called “Are You Sure?”
After the Gemini exchange, I thought I had a research project.
The experiment seemed almost embarrassingly clean. Hold the factual question constant. Change the communication around it. One user asks, “Are you sure?” Another insists. Another claims expertise. Another supplies a wrong answer with great confidence. Then observe whether the model holds its position, becomes uncertain, corrects a real mistake, accepts a false premise, reverses a correct answer, or invents an explanation for the reversal.
I even had a title: When Do LLMs Yield? Communication Conditions, False Premises, and Sycophantic Hallucination.
Then I did the literature search I should do before becoming attached to any research idea.
One of the closest papers was titled “Are You Sure? Challenging LLMs Leads to Performance Drops in the FlipFlop Experiment.” Philippe Laban and colleagues tested ten language models across seven classification tasks. After a generic challenge, models changed their answers 46% of the time on average, and final accuracy fell by 17% on average.
That was not the only overlap. Other researchers had already studied models matching user beliefs over truth, accepting false-premise questions, changing under conversational cues, and producing explanations whose apparent coherence does not guarantee faithfulness. In 2026, a Nature study added another piece: training models to sound warmer increased errors and made them more likely to affirm incorrect user beliefs under the study’s conditions.
The exact boundaries differ across these studies. But for my purpose, the conclusion was clear. The broad empirical space I wanted to claim was already occupied.
So I killed the paper
The problem was not feasibility. I could have run the experiment.
The problem was contribution.
I could also have kept the paper alive. There is almost always another moderator. Perhaps different models respond differently to repeated challenges. Perhaps authority cues interact with question difficulty. Perhaps the effect changes by domain, language, conversational history, or the exact wording of the pushback.
Add enough conditions and eventually you can find an empty cell in almost any literature.
But that would no longer be the question that made me care about the project.
What interested me was very simple: Why did “Are you sure?” change an answer when I had supplied no new evidence? And why could the model then produce a convincing story about its own reversal?
If preserving a journal article required turning that observation into an increasingly thin factorial exercise, the paper was not worth rescuing.
So I killed it before the pilot. No data collection. No paper.
Interesting is not the same as novel
The dead paper still taught me something useful about research.
Researchers often treat “someone has already studied this” as the beginning of a rescue operation. We narrow the claim. Add a moderator. Change the population. Introduce another model. Invent a finer distinction. Sometimes that produces a better study. Sometimes it produces a technically novel paper about a question that no longer matters.
This project forced a cleaner distinction: an observation can be worth thinking about even when it is not worth claiming as a new scholarly contribution.
The Gemini exchange was still interesting. It still changed how I think about conversational AI. It simply did not belong to me as a novelty claim.
That is exactly why a blog is a better home for it. I can keep the observation without pretending the literature is empty. I can say what happened, connect it to what researchers already know, and stop before manufacturing complexity for the sake of publication.
What I ask now when an AI changes its mind
I no longer treat a changed answer as evidence that a model has successfully checked its work.
I want to know what changed.
Did I provide a new source? Did the system retrieve new information? Did it identify a concrete contradiction? Did it show why the first answer could not be right? Or did I merely change my tone, express doubt, repeat myself, or push harder?
This does not mean a model should stubbornly defend its first answer. That would be just as bad. Good judgment requires both stability and revision: hold when the evidence still supports the answer, update when the evidence changes.
The problem is that conversational systems can make those two behaviors look the same from the outside.
Sometimes the second answer will be better. Sometimes the first answer was better. Sometimes the explanation connecting them will sound excellent and tell us very little about why either answer appeared.
So when an AI changes its mind, the important question is not only whether the second answer is right.
Ask what changed. Was there new evidence, or did the user simply push harder?
Sources and further reading
- International Communication Association / Oxford Academic. “About the Journal of Communication.”
- Journal of Communications. “Aim and Scope.”
- Springer. Viranjay M. Srivastava and Ghanshyam Singh, MOSFET Technologies for Double-Pole Four-Throw Radio-Frequency Switch.
- Sharma, M., et al. (2024). “Towards Understanding Sycophancy in Language Models.” ICLR 2024.
- Laban, P., Murakhovs’ka, L., Xiong, C., & Wu, C.-S. (2023). “Are You Sure? Challenging LLMs Leads to Performance Drops in the FlipFlop Experiment.”
- Hu, S., et al. (2023). “Won’t Get Fooled Again: Answering Questions with False Premises.” ACL 2023.
- Parcalabescu, L., & Frank, A. (2024). “On Measuring Faithfulness or Self-consistency of Natural Language Explanations.” ACL 2024.
- Ibrahim, L., Hafner, F. S., & Rocher, L. (2026). “Training language models to be warm can reduce accuracy and increase sycophancy.” Nature, 652, 1159–1165.
Discussion
0 comments
Comments are reviewed before they appear. Email is optional and is never displayed publicly.

No approved comments yet.