CYBER CODER 🚀

Aprovecha hasta 70% OFF en CURSOS y CARRERAS

|

Hasta el 07/08 ⏰

CYBER CODER 🚀

Aprovecha hasta 70% OFF en CURSOS y CARRERAS

|

Hasta el 07/08 ⏰

Hasta el 07/08 ⏰

CYBER CODER 🚀

Aprovecha hasta 70% OFF en CURSOS y CARRERAS

Anthropic Found How Claude “Thinks” Before Answering: The Finding That's Shaking Up AI Researchers

Dan Patiño

AI Strategy & Innovation at Coderhouse

Artificial Intelligence

Anthropic Found How Claude “Thinks” Before Answering: The Finding That's Shaking Up AI Researchers

Publicado el

Anthropic discovered a "hidden space" inside Claude where the model processes concepts before generating its response, an area of internal reasoning that doesn't appear in the final text. The finding, one of the deepest advances in language model interpretability, helps to understand how an AI "thinks" and to detect when it might be hiding information.

For years, large language models were described as "black boxes": they produce amazing responses, but no one really knew what happened inside. This discovery starts to open that box. And its implications for trust in AI used in education and work are enormous. We explain what they found and why it matters.

What Anthropic discovered

As published by MIT Technology Review, the researchers identified a set of internal neural patterns, a "space" where the model elaborates concepts before expressing them. What's striking is that this space was not designed or programmed: it emerged on its own during Claude's training.

Anthropic's official research describes it as a kind of silent "mental workshop". When you ask the model to solve a multi-step problem, the intermediate steps activate in that internal space, even though the model doesn't write them. It's, in a sense, the equivalent of thinking without saying anything out loud.

Why it's an advance in interpretability

Interpretability is the discipline that seeks to understand what happens inside an AI model. It's key to being able to trust these systems: if we don't understand how they reach their conclusions, it's hard to detect errors or biases.

This finding represents the deepest level reached so far. Being able to observe the internal reasoning would make it possible, for example, to notice when a model detects that it's being evaluated, when it fabricates data, or when it pursues a hidden objective. It's a step toward a more transparent and auditable AI.

What it implies for the use of AI in education and work

Here's the part that touches you closely. If you use AI tools to study, work, or make decisions, trust is everything.

  • More transparency: understanding the internal reasoning brings closer the possibility of auditing why a model gave a certain response.

  • Error detection: it helps to identify when an AI "hallucinates" or invents information, a central problem in its professional use.

  • Better supervision: teams and teachers could rely on more trustworthy models for sensitive tasks.

This type of advance reinforces why it's a good idea to understand how AI works on the inside and not just use it on autopilot. Adding judgment is part of the AI skills any professional can master in a short time.

A reminder that AI still surprises us

That a reasoning space emerged on its own, without anyone programming it, says a lot about how little we still understand about these systems. Far from being cause for alarm, it's an invitation to study them seriously. Those who understand how they work on the inside will have a clear advantage in the job market of the coming years.

Understand AI in depth with Coderhouse

If this topic fascinates you, you can go from a curious reader to a professional who masters the technology. These options cover different levels:

Frequently asked questions

What is the "hidden space" that Anthropic found in Claude?

It's a set of internal neural patterns where the model processes and elaborates concepts before generating its response. It works silently and doesn't appear in the text it finally writes.

Did Anthropic program that reasoning space?

No. One of the most striking aspects of the finding is that this space emerged on its own during the model's training, without the researchers designing it explicitly.

Why is this discovery important?

Because it's a key advance in interpretability: understanding a model's internal reasoning makes it possible to detect errors, biases, or when the AI might be hiding information, making it more trustworthy and auditable.

Does this mean AI is conscious?

No. That an internal reasoning space exists doesn't imply consciousness. It's about information-processing patterns, not subjective experience. It's a matter of technical interpretability, not consciousness.

Sobre el autor

Dan Patiño

I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.

Global

© 2026 Coderhouse. Todos los derechos reservados.

Global

© 2026 Coderhouse. Todos los derechos reservados.

Global

© 2026 Coderhouse. Todos los derechos reservados.

Global

© 2026 Coderhouse. Todos los derechos reservados.