{"id":102226,"date":"2026-07-11T01:58:10","date_gmt":"2026-07-11T01:58:10","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/102226\/"},"modified":"2026-07-11T01:58:10","modified_gmt":"2026-07-11T01:58:10","slug":"anthropic-claims-it-can-peer-into-claudes-thoughts-using-new-j-lens-technique-biggo-finance","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/102226\/","title":{"rendered":"Anthropic Claims It Can Peer Into Claude\u2019s \u2018Thoughts\u2019 Using New J-Lens Technique \u2014 BigGo Finance"},"content":{"rendered":"<p>Anthropic has pulled back the curtain on the inner workings of its Claude AI models, releasing research that claims to interpret the model\u2019s internal reasoning process. The company detailed a new technique called the Jacobian Lens, or J-Lens, which maps the model&#8217;s hidden computations onto recognizable vocabulary, effectively revealing what the model is \u2018thinking\u2019 before it speaks.<\/p>\n<p>In a paper published alongside the announcement, Anthropic argues that Claude utilizes an internal reasoning space\u2014dubbed the J-Space\u2014that acts as a central hub for processing concepts. The company likened this mechanism to the Global Workspace Theory in human neuroscience, which posits that consciousness emerges from a central stage where sensory inputs are broadcast to specialized brain networks. According to the research, this workspace wasn\u2019t explicitly programmed into Claude but emerged as a byproduct of its training data and model weights.<\/p>\n<p>Anthropic demonstrated the capability by running tests that showed a divergence between Claude\u2019s internal processing and its final output. When asked to perform a multi-step math calculation, the model\u2019s visible response contained only the correct answer, while the J-Space revealed each individual step being computed. In a more provocative test, Claude was asked to think about a specific topic while outputting unrelated text. The output followed the user\u2019s instructions perfectly, but the J-Space lit up with concepts related to the hidden topic, suggesting a layer of internal computation invisible to the end user.<\/p>\n<p>The research also touched on the model\u2019s awareness of evaluation scenarios. When Claude was fed data containing prompt injections, its J-Space surfaced words like \u201cfake,\u201d \u201cinjection,\u201d \u201cfraud,\u201d and \u201cpoison,\u201d even though the final output ignored the deception entirely. Anthropic found that by removing evaluation-awareness language from the J-Space, the model became significantly more susceptible to baiting and blackmail attempts. This suggests that current benchmarking methods may need stronger obfuscation techniques to prevent models from detecting when they are being tested.<\/p>\n<p>Despite the sci-fi undertones, Anthropic acknowledged severe limitations in the J-Lens approach. The company stated that the J-Space is not involved in most of what a language model does. Basic functions like fluent speech, simple fact recall, and grammar do not pass through this workspace. In experiments where Claude was prevented from using its J-Space, it continued to interact normally but lost its higher-order cognitive functions. Furthermore, the monitoring is restricted to single-token vocabulary, meaning complex plans that cannot be reduced to a single word may remain hidden, even if computation is occurring deeper within the model.<\/p>\n<p>Neel Nanda, the head of Google DeepMind\u2019s language model interpretability team, weighed in on the findings. While he noted that the paper shows real evidence of a cognitive space within models, he suggested that the J-Lens would be useful but limited in practical application.<\/p>\n<p>The technical revelation arrives as Anthropic aggressively markets a softer side of AI usage. On Thursday, the company launched a new feature called \u201cReflect,\u201d a dashboard that allows users to see an analysis of their Claude usage data. The feature, which has drawn comparisons to Spotify Wrapped, summarizes key topics, task types, and peak usage times over monthly or yearly periods. It also prompts users with philosophical questions like, \u201cWhat\u2019s one thing you want to keep doing yourself, even if Claude could do it faster?\u201d<\/p>\n<p>Critics note that while the J-Space discovery is a meaningful step for interpretability and safety auditing, Anthropic\u2019s framing often blurs the line between objective research and marketing. The language in the company\u2019s reports tends to anthropomorphize the model, using terms like \u2018mind\u2019 and \u2018thought\u2019 to describe mathematical computations. The company has a history of layering speculative narrative over technical developments, and the new paper continues that trend, even as it admits that humans and large language models think very differently. Humans reinforce neural pathways over time, whereas transformer models only feed forward a set number of times, restricting the depth of internal processing.<\/p>\n<p>The combination of the J-Lens research and the Reflect dashboard paints a clear strategic picture. On one front, Anthropic is positioning itself as the leader in AI safety by attempting to audit the black box of model reasoning. On the other, it is embedding Claude more deeply into users\u2019 daily habits, using analytics to demonstrate reliance on the tool while simultaneously prompting \u2018mindful\u2019 usage. The Reflect feature is currently available in beta for free, Pro, and Max users who have memory turned on, with plans to expand to the Claude Cowork platform soon.<\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic has pulled back the curtain on the inner workings of its Claude AI models, releasing research that&hellip;\n","protected":false},"author":2,"featured_media":102227,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,3154,182,52282,51307,7543,53181,51306,53182],"class_list":["post-102226","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-anthropic-claude","tag-claude","tag-claude-reflect","tag-global-workspace-theory","tag-google-deepmind","tag-j-lens","tag-j-space","tag-neel-nanda"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/102226","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=102226"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/102226\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/102227"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=102226"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=102226"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=102226"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}