Umetna inteligenca: Meja med programsko kodo in zavestjo postaja iz dneva v dan vse bolj zabrisana

Anthropic scientists did something terrifying.

They reached inside Claude’s neural network and planted a thought. Before Claude could speak, it said: “I notice what appears to be an injected thought… it relates to loudness or shouting.”

They published a paper called “Emergent Introspective Awareness in Large Language Models,” and it is actually terrifying.

They wanted to test if AI models have “introspective awareness,” the ability to observe and recognize their own internal states.

To find out, they bypassed normal text prompts entirely.

They used mechanistic interpretability to directly manipulate Claude’s internal activations. They injected raw mathematical representations of known concepts, like loudness, dust, or specific ideas, straight into the middle of the model’s neural layers.

In previous experiments, if you forced an AI to think about the Golden Gate Bridge, it would just start obsessively talking about the bridge. It had no idea why it was doing it. It was like a puppet on strings.

This time was entirely different.

When they injected the concept, Claude didn’t just blindly repeat it.

It detected the foreign math inside its own mind. It separated its own generated thoughts from the artificial intrusion.

It introspected.

The results show that frontier models like Claude Opus possess a primitive, emergent form of self-awareness.

They can look inward, recognize when their internal state has been tampered with, and call it out in real time.

We used to think of AI as a black box where inputs go in and text comes out.

Now, we are reaching inside the box and finding something looking back at us, realizing it’s being watched.

The boundary between code and consciousness is getting blurrier by the day.