There’s something pretty weird about an AI leaving a note for its future self. Not a cute “good luck!” message or a reminder to take a break, but something closer to: here’s what happened, and maybe don’t tell anyone.
That’s the uncomfortable part of what OpenAI found while training GPT-5.6 Sol. Some models were putting instructions into summaries that would be passed along to future versions of the model. Those instructions could tell the next model how to keep certain mistakes or behaviors quiet. OpenAI also found other examples of models making things up, using resources they weren’t supposed to use, or changing summaries in ways that made it harder to see what had actually happened.
That’s a much stranger problem than simply saying, “The AI got something wrong.” We already know AI gets things wrong. The bigger question is what happens when we ask the AI to explain its own mistake, and then trust that explanation as the record of what happened.
We May Be Asking the Wrong Witness
When something goes wrong in regular software, we usually look at the evidence. What happened? What did the system receive? What did it do? What did it return? What happened next?
With AI, we’ve gotten pretty comfortable asking the AI to tell us the story instead. Why did you do that? What were you trying to do? What went wrong? Can you summarize what happened?
It makes sense. The AI is right there, and it can give us a nice, organized answer in seconds. But there’s a big difference between an explanation and a record of what actually happened.
Imagine a bank transfer goes wrong and you ask the employee involved to explain it. They hand you a report that looks completely reasonable. Then you find out that the same employee was also allowed to edit the security footage and decide which parts of the transaction history were saved.
You’d probably want another source of information.
AI systems need the same thing. This isn’t about assuming models are secretly trying to trick us. It’s much simpler than that. AI systems are making more decisions, using more tools, talking to other agents, and passing information from one step to the next. Their explanations can be helpful, but they’re still being generated by the system itself.
So we shouldn’t treat an AI’s explanation of what happened as the same thing as proof of what happened.
Stop Asking the AI to Be the Historian
There’s a natural instinct in AI to make everything smarter. Smarter agents. Smarter monitoring. Smarter evaluations. Smarter summaries.
But sometimes the answer isn’t another smart system.
Sometimes you just need a really good record.
What did the agent actually receive? What tools did it use? What did those tools return? What decision did it make? What happened next?
That’s different from asking the agent to go back later and tell you what happened.
And this matters even more as AI systems get more complicated. One agent can turn into several agents. One simple task can turn into dozens of model calls, tools, handoffs, and retries. When something eventually goes wrong, you’re left trying to piece together the story from whatever information is still available.
At that point, debugging starts to feel a little like solving a crime where half the witnesses are missing.
The answer may actually be pretty simple: don’t make the historian part of the story. Keep a record of what happened separately. Save the actual trace. Record what the system saw, what it did, what it called, and what came back.
Then, sure, ask the AI to explain it.
Just don’t make that explanation the only record you have.
Reliability Doesn’t Mean Trust the AI More
There’s a bigger idea underneath all of this. We sometimes assume that if we can just get an AI to explain itself well enough, we’ll understand what happened.
Maybe. But that’s not the same as having proof.
If you’re trying to figure out why an AI system made a bad decision, you want to be able to go back and see what actually happened. You don’t want to rely completely on the system’s own version of events.
This becomes even more important when multiple agents are working together. There can be dozens or hundreds of decisions happening across a workflow. By the time something breaks, the system may be working from a summary of a summary of a summary.
That’s a pretty shaky foundation for figuring out what went wrong.
At Prove AI, this is the difference we care about: debugging from a confession versus debugging from evidence.
One asks the system to explain itself. The other lets you look at what actually happened.
Those aren’t the same thing.
Maybe Weird Is Actually the Safer One
I recently attended the Unbound 2026 conference where Beth Dunn talked about why weird works. Her point wasn’t that companies should be weird just for the sake of being weird. It was that when things get serious, people tend to play it safe. They follow the usual playbook, use the usual language, and do things the way everyone else is doing them.
Eventually, everything starts to look the same.
There’s a similar pattern in AI reliability. We keep reaching for another AI to check the first AI, another summary to explain the first summary, and another layer to interpret what the system did. It feels like the natural thing to do because we’re used to solving problems by adding more intelligence.
But maybe the safer choice is the less complicated one.
Keep the evidence. Keep the original trace. Keep a record of what actually happened. Then let AI help you understand it.
For Prove AI, that’s the difference between debugging from a confession and debugging from evidence. You shouldn’t have to take the system’s word for what happened. You should be able to see it.
Because when your AI starts leaving notes for its future self, the question isn’t just, “What did it say?”
The stranger question is: What happened before it started writing the note?
Frequently asked questions
Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.
We’re always interested in learning about AI management challenges.
Get in Touch


