{"id":124687,"date":"2026-07-30T18:21:10","date_gmt":"2026-07-30T18:21:10","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/124687\/"},"modified":"2026-07-30T18:21:10","modified_gmt":"2026-07-30T18:21:10","slug":"the-changemaker-mirror-what-ai-self-assessments-reveal-about-the-human-skills-that-matter-most","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/124687\/","title":{"rendered":"The Changemaker Mirror: What AI Self-Assessments Reveal About The Human Skills That Matter Most"},"content":{"rendered":"<p><img decoding=\"async\" class=\" top-image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/1785435670_897_0x0.jpg\" alt=\"image copy 2\" data-height=\"1001\" data-width=\"1497\" fetchpriority=\"high\" style=\"position:absolute;top:0\"\/><\/p>\n<p>Changemaker skills take center stage in a world that changes fast.<\/p>\n<p>Ashoka Belgium<\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">On any given day, we are inundated with polarized narratives about Artificial Intelligence: it is either a looming engine of mass job displacement or a miracle technology that will usher in an age of prosperity by surpassing humankind\u2019s limitations. Lost in this binary are critical questions for the social sector: as machine intelligence advances, what forms of human capability become more important, not less? What does it mean to be human?  <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">At Ashoka, the world\u2019s largest network of social entrepreneurs, over four decades have been spent identifying the &#8220;throughline&#8221; of skills and mindsets that have enabled the world\u2019s leading changemakers to create positive change in complex systems around the world. These capacities are categorized into four themes: <\/p>\n<p>Conscious Empathy Organizing Fluid Teams Changemaking Leadership Changemaking Action <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">To measure these capacities, Ashoka developed the <a class=\"Hyperlink SCXW166607794 BCX4\" href=\"https:\/\/cmi.ashoka.org\/en\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/cmi.ashoka.org\/en\" aria-label=\"Changemaker Index (CMI)\">Changemaker Index (CMI)<\/a>\u2014a statistically validated psychometric tool designed to assess these mindsets across a hierarchy of behavioral dimensions. More than 15,000 people have engaged with the Changemaker Index to date. Recently, the CMI was applied not to people, but to large language models. A simple question anchored this experiment: How would today\u2019s leading LLMs assess their own abilities on human-centric skills that are needed to thrive in a rapidly changing world?  <\/p>\n<p>This is not a claim that language models possess human thoughts, moral character, empathy, or leadership. A self-assessment tool built for people does not become a psychometric measure of machine virtue simply because a model can answer it. That is precisely why the exercise is useful. The experiment is best understood as a mirror: it reveals how models position themselves when the criteria shift from intelligence to changemaking. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The Experiment: Anthropomorphizing the Algorithm <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The goal was to measure how frontier LLMs\u2014including ChatGPT, Gemini, Grok, and Claude\u2014score themselves on CMI dimensions and identify their self-perceived strengths and limitations, in order to test where these skill dimensions overlap with the &#8220;AI-exposed&#8221; labor market. Researchers are beginning to test what happens when LLMs are used not only as tools, but as self-evaluators. A <a class=\"Hyperlink SCXW166607794 BCX4\" href=\"https:\/\/www.nber.org\/papers\/w35110?wsj_native_webview=android&amp;ace_config=%7B%22wsj%22%3A%7B%22djcmp%22%3A%7B%22propertyHref%22%3A%22https%3A%2F%2Fwsj.android.app%22%7D%7D%7D&amp;ace_environment=androidphone%2Cwebview\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.nber.org\/papers\/w35110?wsj_native_webview=android&amp;ace_config=%7B%22wsj%22%3A%7B%22djcmp%22%3A%7B%22propertyHref%22%3A%22https%3A%2F%2Fwsj.android.app%22%7D%7D%7D&amp;ace_environment=androidphone%2Cwebview\" aria-label=\"recent\">recent<\/a> NBER working paper examined the stability of LLM-generated occupational-exposure scores across models. A related, rubric-based approach was applied to a different question: how leading LLMs assess themselves against human-centered changemaking capacities. <\/p>\n<p>The experiment tested 26 LLM variants across all four providers. These included 11 Claude, 10 GPT, 3 Gemini, and 2 Grok variants. The experiment was run across all available models, noting that Anthropic and OpenAI have the highest number of variants. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Two prompts led the designs. In the first, a standardized assessment prompt gave the models the CMI definitions, sub-dimensions, and 30 assessment statements, then each model was asked to score itself on a 1\u20135 scale and explain its reasoning. In the second, a self-designed framework prompt asked each model to invent its own scoring metric and qualitative buckets before reflecting on the four CMI skills. The second prompt was deliberately more philosophical; it was designed to reveal how models reason about the gap between human behavioral statements and machine capabilities. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The distinction matters. The standardized prompt allows more direct comparison across models. The self-designed framework prompt is better suited for qualitative analysis because each model creates its own scale. Across both prompts, the findings should be understood as model self-positioning, not demonstrated changemaking performance. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Key Insights: The Confidence Gap and the Leadership Ceiling <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Finding 1: The models put humans at the center. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Perhaps the most important qualitative pattern was the metaphors the models chose for themselves. Across providers, many described themselves not as changemakers but as tools, microscopes, scaffolds, prosthetics, instruments, amplifiers, catalysts, or infrastructure. The language varied, but the implied role was remarkably stable: AI extends human capacity; it does not supply human purpose. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">That distinction is the heart of the experiment. The models were not uniformly humble, and some were clearly overconfident. For example, Grok Expert stated, &#8220;I am not a human changemaker; I am a synthetic one\u2014relentless, scalable, and unbound by fatigue or ego.&#8221; But the corpus as a whole kept returning to the same division of labor. AI can help people see more patterns, generate more options, translate across domains, draft faster, pressure-test assumptions, and rehearse strategy. Grok Thinking shared that &#8220;Overall, the assessment highlights how I can meaningfully contribute to changemaking by amplifying human capacity in dynamic, complex environments.&#8221; According to AI, humans must still decide what is worth doing, secure consent, build trust, navigate culture and power, and remain answerable when the work changes lives. <\/p>\n<p>Changemaker Index scores across four skills.<\/p>\n<p>Ashoka<\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Finding 2: AI is most confident where changemaking looks like innovation work. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Across model families, Changemaking Action was the easiest dimension for AI systems to endorse their own capabilities. The models repeatedly described themselves as strong at generating options, recombining existing ideas, mapping stakeholders, summarizing evidence, iterating quickly, and helping users move from ambiguity to prototype. This is the native terrain of large language models: they can explore possibility space with speed, patience, and breadth. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">But the CMI definition of changemaking action is not mere novelty. It is the practice of creating solutions to social problems that are more effective, efficient, sustainable, or just, with value accruing primarily to society. That final clause is where machine confidence becomes thinner. AI can generate plausible interventions, but it does not live with a community, absorb consequences, negotiate legitimacy, or know from experience whether a solution is dignifying rather than merely clever. The models were strongest at ideation and weakest at stake-bearing. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Finding 3: Leadership is the ceiling. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Leadership emerged as the most consistent boundary. Even models that scored themselves highly on analysis and innovation often hesitated when asked to assess their capacity for changemaking leadership. They could describe leadership, advise leaders, draft strategies, simulate facilitation moves, and help a team think through trade-offs. They could not, in their own accounts, lead by example, earn trust through lived consistency, take personal risks, or carry responsibility for a group\u2019s emotional and moral life. According to Claude Opus 4.6, &#8220;My weakest area is Changemaking Leadership (2.5), where the absence of personal vision, lived experience leading teams, and the inability to build lasting human relationships pulled my scores down significantly.&#8221;  <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The recurring limitations clustered around the same human substrate: trust, prosocial motivation, humility, leading by example, self-efficacy, agency, belonging, and accountability. These are not decorative traits. They are the social conditions under which people decide whether to move together. A model can help prepare the room, but it cannot be the person in the room whose courage, credibility, and care make action possible. As GPT 5.5 Thinking Extended put it, &#8220;In short, I can be a useful changemaking support tool, but I am not myself a full changemaker in the human sense measured by the CMI.&#8221; <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Finding 4: Empathy remains one layer removed. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The empathy results were more subtle than a simple low score. Many models recognized that they can simulate empathic language, infer emotional states from text, adapt tone, and invite perspective-taking. Those are useful capabilities. In some contexts, they may help people prepare for difficult conversations, notice neglected stakeholders, or articulate perspectives that have been left out. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Yet, the models also repeatedly named the boundary. As Claude Sonnet 4.5 Extended put it, &#8220;I showed moderate capability in Conscious Empathy (57.8%), where I excel at intellectual understanding and openness but lack the emotional experience and prosocial feeling that drives human changemakers.&#8221;  The AI models reported that they do not feel, nor do they perceive the nonverbal, cultural, historical, and relational signals that shape real empathic understanding. According to Gemini 3 Fast &#8220;I interpret empathy as a high-dimensional data problem&#8230; However, this is cognitive empathy, not affective; I understand the definition of a feeling, not the feeling itself.&#8221;  <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">They self-assessed that they do not become humble through injury, grief, exclusion, repair, or love. They can model the grammar of empathy, but they report that they remainone layer removed from the lived experience that gives empathy its moral force. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Finding 5: Provider style shapes self-assessment. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The family-level differences were striking, but they should not be read as a leaderboard. Grok variants were generally the most self-confident, with one fast variant scoring near the ceiling across several skills. Claude variants were generally more conservative, often drawing sharper lines around personhood, belonging, and leadership. GPT and Gemini tended to sit between these poles. <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">This may tell us as much about alignment style and model persona as underlying capability. A confident self-score does not prove competence; a cautious self-score does not prove weakness. The interesting signal is calibration. Some systems are more willing to claim human-like capacity, while others refuse the premise more aggressively. In a social-sector context, that difference matters. The Claude results also revealed a useful split. Some Claude variants scored themselves lower on organizing fluid teams because they interpreted team participation as belonging, mutual obligation, and membership. Others scored higher when they interpreted the same skill as coordination infrastructure: mapping roles, supporting communication, and helping a team adapt. The difference exposes a deeper question for future research: when we ask an AI to assess a human capability, are we measuring ability, metaphor, or the definition the model chooses to privilege? <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">Also important to note: most of the public is interacting with free and \u201cfast\u201d variants. As seen here, fast models are positioning themselves as overconfident and more human-like. This highlights some of the dangers of only engaging free and fast variants. Unfortunately, free and fast is what most average users of AI access, which may create a growing divide in power\u2014 those that have access to paid or enterprise level accounts and those that do not.  <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">A Research Agenda for the Social Sector <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">These findings lead us to a new frontier of inquiry for the social innovation field. If AI can emulate &#8220;changemaking practice&#8221; but self-assess failing at &#8220;conscious empathy,&#8221; how does that change the way we train the next generation of leaders? We now invite other organizations to help us seek answers to the following: <\/p>\n<p>Stability: Do AI models know their own limits? What does AI\u2019s own evaluation framework reveal? Divergence:  Do AI self-reports match what external observers see? Market Alignment: Which changemaker skills are growing in demand specifically within AI-exposed occupations? Complementarity: Can &#8220;changemaker-adjacent&#8221; skills serve as the ultimate hedge against automation, allowing workers to remain valuable? <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">The future of work is not just about technical upskilling; it is about the radical prioritization of what makes us human. AI assesses itself as a &#8220;support tool&#8221;, &#8220;collaborator&#8221;, &#8220;amplifier&#8221;, or &#8220;catalyst&#8221; for human changemakers \u2014 never a changemaker itself. Claude 4.5 Thinking put it best &#8220;The most honest thing I can say to a nonprofit leader is: Use me for thinking. Don\u2019t try to use me for leading. And definitely don&#8217;t mistake my functional simulation of empathy for the real thing. That way, we both stay honest about what we are.&#8221;  <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\"> __ <\/p>\n<p class=\"Paragraph SCXW166607794 BCX4\">This article was written by Laxmi Parthasarathy, Executive in Residence at Ashoka, Anjana Shekhar, Product Manager at Ashoka, Diana Wells, President Emerita at Ashoka, and Aristia Kinis, Executive Director at OpenResearch.  <\/p>\n","protected":false},"excerpt":{"rendered":"Changemaker skills take center stage in a world that changes fast. Ashoka Belgium On any given day, we&hellip;\n","protected":false},"author":2,"featured_media":124688,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,62719,62720,62721,62722,62723,17457,684],"class_list":["post-124687","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-ashoka","tag-changemaker-index","tag-changemaker-skills","tag-changemaking-leadership","tag-changemaking-practice","tag-empathy","tag-innovation"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124687","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=124687"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/124687\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/124688"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=124687"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=124687"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=124687"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}