Assertion 1
Switching on thinking, the mode where a model reasons step by step before answering, recovered roughly 40 to 65 percent of the stored-but-unretrieved facts in models tuned for it, but 11 to 12 percent still refused to surface even then, and the 2 to 5 percent of facts that were never encoded stay out of reach no matter how hard the model thinks Google Research Blog.
Assertion status: No spot-check verdict is published for this assertion.
The authors state that Gemini-3-Pro and GPT-5 encode 95–98% of facts but fail to directly recall 26–34% of them.
No stored spot-check names this claim in this edition.
The authors claim that even with 'thinking' enabled, Gemini-3-Pro and GPT-5 fail to recall 11–12% of facts.
No stored spot-check names this claim in this edition.
The authors state that 'thinking' recovers roughly 40–65% of encoded-but-not-directly-known facts in thinking-optimized models.
No stored spot-check names this claim in this edition.