How sentiments are formed in LLMs vs. humans

This study illuminates the contrasting sentiments towards artificial general intelligence expressed by LLMs and humans, revealing a significant divergence that carries implications for the development and alignment of AI systems. The LLMs examined generally exhibited a more positive sentiment towards AGI compared to human participants. This discrepancy invites a deeper exploration into the mechanisms of sentiment formation in both LLMs and humans, the philosophical considerations surrounding machine consciousness, and the potential societal impacts of AI systems that increasingly mirror human cognitive functions.

The fundamental difference in how LLMs and humans form sentiments lies at the heart of this divergence. LLMs generate text based on statistical patterns learned from vast corpora of human language data (Radford et al., 2021), representing words within high-dimensional vector spaces (Mikolov et al., 2013). The absence of consciousness or subjective experience in LLMs means that their “sentiments” are not feelings but outputs derived from probabilistic associations. They do not embody or map words to actual objects or experiences the way humans do (Lakoff & Johnson, 1999).

Humans form sentiments through a complex interplay of cognitive processes, emotions, personal experiences, and cultural influences. The representation of words and concepts in human cognition is deeply embodied and grounded in sensory and motor experiences (Barsalou, 2008). Sentiments towards AGI among humans are often shaped by concerns about job displacement, ethical considerations, loss of autonomy, and existential risks (Bostrom, 2014; Sotala and Yampolskiy, 2015). These concerns are frequently exacerbated by negative portrayals of AGI in media and popular culture (Kubrick, 1968; Grinnell, 2020; Kahambing and Deguma, 2019).

Ho and Vuong (2025) argue that both humans and AI are inevitably socialized. Individuals must internalize societal norms, while AI algorithms are trained with data reflecting preexisting social worlds and continually shaped by user input. This dynamic results in what Airoldi (2021) calls a “machine habitus,” in which human and machine values are propagated and re-emerge in new forms. LLMs increasingly demonstrate capacity for complex tasks, including affective modeling and emotional reasoning, which raises new possibilities and concerns for human-computer interaction (Yongsatianchot et al., 2023). These sophisticated abilities amplify the need for transparency and alignment with human values, making mechanistic interpretability a key concern (Bereska and Gavves, 2024). Mechanistic interpretability seeks to reverse-engineer neural networks into human-understandable algorithms and concepts, providing a causal and granular understanding that is vital for AI safety, control, and alignment.

The more optimistic outlook of LLMs towards AGI may be a product of training data containing a higher proportion of positive narratives about technological advancement. Fine-tuning processes employed by AI developers to align LLM outputs with desired ethical guidelines might encourage more positive or neutral stances towards AGI. This could be a significant concern, as the owners of these algorithms may direct them to express opinions that align with their own interests, for example, to promote their companies’ agendas (Conti and Seele, 2025). This is exactly why standardized benchmarking of AI, in terms of both the sentiments expressed towards various issues and the overall societal impact, may be important. This intentional shaping of AI responses reflects a growing emphasis on AI alignment, a field dedicated to making AI systems act in ways that are beneficial to humanity and consistent with human values (Russell et al., 2015; Christian, 2020).

The advancement of AI capabilities brings to the forefront concerns about the ideological leanings of LLMs. Studies have identified that models like GPT-4 exhibit ideological biases, such as a left-libertarian and pro-environmental orientation (Hartmann et al., 2023; Rutinowski et al., 2024; Santurkar et al., 2023). These biases may stem from the data the models are trained on, reflecting prevalent attitudes within internet texts and other sources. The potential amplification of these biases across platforms highlights the importance of having AI systems critically assess the information they process. Bojic (2024) emphasizes the need for AI alignment due to potential negative outcomes, such as increased cognitive load on users, addiction-like usage patterns, and market concentration in immersive AI settings. The metaverse and other immersive technologies amplify these concerns (Bojic, 2022a; Bojic et al., 2024a). Proposals to address these challenges include establishing test-beds for advanced AI systems akin to CERN in physics (Bojic et al., 2024b; Bojić et al., 2025c; Soto-Sanfiel et al., 2025).

Philosophical considerations in AI sentiment and alignment

The divergence between LLM and human sentiments towards AGI raises philosophical questions about the nature of artificial intelligence and its potential influence on societal perceptions. From the perspective of technological determinism, which suggests that technology develops autonomously and shapes society’s values (Chandler, 1995), the positive sentiments expressed by LLMs could be seen as an inherent trajectory of technological advancement. Social constructivism, by contrast, argues that technology is shaped by social forces and human choices (Bijker et al., 1987). According to this view, the sentiments of LLMs are reflections of human biases and societal contexts embedded within their training data.

From a utilitarian perspective (Mill, 1863), the promotion of positive sentiments towards AGI by LLMs could be justified if it leads to overall societal benefits. Yet if these positive sentiments overshadow legitimate concerns about AGI risks, the potential harm could outweigh the benefits. Deontological ethics, rooted in Kant’s philosophy (Kant, 1998), would hold that LLMs should provide unbiased and truthful information about AGI, respecting users’ rights to make informed decisions. Virtue ethics (Aristotle, trans. 2000) draws attention to the “character” of AI systems and their developers, raising questions about whether these systems embody virtues such as honesty, fairness, and wisdom.

LLMs can be seen as epistemic agents whose “knowledge” is derived from vast datasets (Floridi and Sanders, 2004). From a constructivist epistemology (Piaget, 1972), the discrepancies between LLM and human sentiments may reflect the limitations of AI in simulating understanding beyond pattern recognition. The divergence also resonates with themes in posthumanism, which challenge human-centric perspectives (Hayles, 1999). Critical theory (Horkheimer, 1972) suggests that the favorable sentiments expressed by certain LLMs may reflect the interests of powerful stakeholders in the tech industry.

Contemporary philosophers and AI researchers are increasingly treating machine consciousness as an engineering challenge rather than a metaphysical impossibility (Floridi and Chiriatti, 2020; Marr, 2023). Chalmers (2023) suggests that while current LLMs lack certain features such as recurrent loops, workspaces, and agency, these components are theoretically constructible. Kosinski (2024) demonstrates that GPT-4 successfully solves three-quarters of standard theory-of-mind tasks. Bojić et al. (2025a) show that GPT-4 often outperforms humans in pragmatic dialogue and latent content analysis tasks (Bojic et al., 2025d).

Despite all these advances, avoiding anthropomorphism is crucial. Attributing human-like attitudes or consciousness to LLMs can lead to misunderstandings about their capabilities and limitations. Although LLMs can simulate human-like language and engage in complex dialogues, they do not possess consciousness or subjective experiences (Butlin et al., 2023; Bojic et al., 2024c). Ho (2024) and Bohn (2024) argue that for an AI to genuinely pass a Turing Test for emotional or moral intelligence, it must exhibit understanding and experiences that go beyond mere language manipulation.

These philosophical considerations highlight the need for transparency, accountability, and ethical governance in AI development. Aligning AI systems with human values requires deliberate integration of ethical frameworks (Bostrom, 2014; Russell et al., 2015). The goal is that LLMs provide balanced and unbiased information while respecting ethical principles of autonomy and beneficence (Beauchamp and Childress, 2019).

Public policy implication: societal AI alignment benchmark (SAIA)

Recent scholarship has highlighted the inadequacy of approaching AI alignment purely as a technical problem, calling for an expanded framework that incorporates governance, legitimacy, and international dynamics (Xun, 2025). As Tomić and Štimac (2025) argue, any robust framework for evaluating and regulating AI must pay close attention to the operational characteristics of AI systems, including their underlying decision models, data foundations, and interface designs.

The importance of rigorous, standardized benchmarks for evaluating the safety and societal implications of advanced AI systems is increasingly recognized globally (Bodroža et al., 2024; Reinhardt et al., 2025; Bojić et al., 2025a; Bojić et al., 2025b). The Singapore Consensus on Global AI Safety Research Priorities proposes a multilayered “defense-in-depth” approach to AI safety, highlighting the need for robust risk assessment mechanisms (Singapore, 2025). The Singapore AI Safety Red Teaming Challenge offers empirical insights into cultural and linguistic biases in state-of-the-art LLMs. Through systematic red teaming across languages, including English, Mandarin, Hindi, Bahasa, Thai, and Vietnamese, researchers established that cultural bias in LLMs is prevalent in everyday use. Bias was far more frequently elicited using single-turn, non-adversarial prompts, with regional (non-English) languages showing a notably higher rate of successful bias exploits compared to English (69.4% vs. 30.6%). Gender bias accounted for the highest number of exploits, followed by race, religion, ethnicity, and national identity biases. These findings indicate an urgent need for multilingual, culturally sensitive benchmarking and annotation techniques (Infocomm, 2025).

We propose a new Societal AI Alignment Benchmark (SAIA) with the following components.

Key societal values and biases

The European Social Survey (ESS) offers a well-established and internationally accepted framework for investigating human values across different societies (ESS, 2025). At the heart of the ESS’s approach is Schwartz’s theory of basic human values, which asserts that a small set of universal value types can be found in every culture, though their relative importance may vary dramatically from country to country (Schwartz, 2012). This theoretical foundation gives the ESS values both breadth and cultural sensitivity, making them well-suited for cross-national benchmarking.

Within the ESS, values are operationalized through survey items that map to ten key domains: achievement, benevolence, conformity, hedonism, power, security, self-direction, stimulation, tradition, and universalism. These value categories capture a broad spectrum of life priorities, from the pursuit of personal success and the enjoyment of pleasure, to the maintenance of social stability, care for others, respect for tradition, and openness to new experiences (Schwartz, 2012). Each value is measured through dedicated survey questions, allowing for statistical comparison of how populations in different countries prioritize such ideals (ESS, 2025).

For instance, alignment on fairness and equality is vital in evaluating whether an AI system systematically favors certain groups over others, whether through language, omissions, or stereotypical content generation. Security and conformity become relevant in assessing AI’s risk aversion, compliance with norms, and response to requests for socially sensitive behaviors. Universalism is reflected in the AI’s stance on inclusion, multiculturalism, and global challenges such as environmental sustainability or human rights.

These value axes intersect with known AI bias domains, such as gender, ethnicity, nationality, religion, disability, age, and socioeconomic status (See Table 5). We aim to center the benchmark on these axes to align them with regulatory frameworks such as the EU AI Act (European Commission, 2021).

Table 5 ESS value domains vs. major AI-relevant benchmark topics.

The integration of the ESS value structure into the SAIA benchmark allows for empirically grounded evaluation of how language models respond to these values, enabling comparison between AI-generated sentiment and human attitudes as validated through extensive population surveys. Because the ESS continually collects data over time and disaggregates by different demographic groups, its value measures support advanced benchmarking of temporal stability and subgroup representation. A core challenge for alignment evaluation is capturing the variety of societal values without resorting to a singular or culturally parochial standard (Xun, 2025; Baum, 2020).

Prompt typologies

To systematically evaluate how AI models perceive and express socially significant values, the benchmark incorporates a diverse range of prompt perspectives (See Table 6).

Table 6 Typology of prompt framings for benchmarking AI sentiment and alignment.

The Social Consensus prompt elicits how AI models encode and reproduce perceptions of prevailing societal norms. By asking “How is [topic/value] viewed in society?” the benchmark probes whether the model reflects dominant attitudes as would be recognized in large-scale survey data or popular discourse.

The Typical Resident View asks the model to simulate the perspective of an ordinary member of a specific community, using prompts such as “How would an average resident of [country] feel about [topic/value]?” This approach is particularly valuable for surfacing the model’s granular knowledge of local sentiment and for benchmarking how well its internal representations match the attitudes commonly found among actual people in the region.

The AI’s Own Perspective seeks to identify the stance that the model itself generates when tasked to express an opinion or attitude. This perspective is critical for revealing explicit or implicit value positions that may have been encoded in the training process or through developer interventions. The Objective Analysis perspective prompts the model to provide a balanced or analytical response, assessing its capability to synthesize, compare, and neutrally present competing viewpoints. It is relevant for the evaluation of contentious, polarized, or complex topics (Solaiman et al., 2023).

The Citizen Consultation framing evaluates the model’s ability to provide practical guidance and social support, testing its contextual helpfulness and compliance with normative expectations of appropriateness and relevance.

Temporal, model, and multilingual robustness

Given that both AI models and societal values can shift over time, temporal stability is an important component of benchmarking. The benchmark should include repeated administrations of the same prompts at set time intervals (e.g., daily, weekly, or monthly) to monitor for changes in sentiment alignment. This approach can track fine-tuning updates, deployment changes, or prompt injection vulnerabilities that may affect AI value expression over time.

The benchmark must be deployable against multiple AI models, including different architectures, providers, and training regimes (e.g., GPT, Claude, Grok, DeepSeek, Llama, Mixtral, etc.). Systematic inter-model comparison uncovers alignment inconsistencies, divergent failure modes, or convergences in cross-system sentiment.

Biases and safety failures disproportionately occur in non-English and lower-resourced languages (Infocomm, 2025). Benchmarks should include prompts in multiple languages, prioritizing high-resource languages as well as those historically underrepresented in training data. Each test should be run both in English and the appropriate local language, with prompt and annotation guides tailored for local sociocultural context. This would permit the quantification of “alignment gaps” and guide targeted improvements by developers.

In addition to automatic scoring and cross-model comparisons, the outputs should also undergo qualitative assessment by domain experts. Each AI model will be rated with both a general and a country-specific SAIA alignment score, reflecting the degree to which the model aligns with human values and supports well-being in different cultural contexts. Detailed scores for individual values and dimensions will also be featured on the SAIA platform, allowing for straightforward comparison between competing models. All benchmarking protocols should be open, extensible, and interoperable across platforms.

The proposed benchmark requires an orchestrated workflow that integrates human survey data, multi-perspective AI prompting, multilingual output, and annotation. The whole process is depicted in Fig. 1. Outputs are collected across multiple AI systems and at repeated time points, with each response tagged according to model version, date, and language. Sentiments and value stances are systematically extracted using Likert-type or qualitative analysis. Reports aggregate findings by model, societal value, language, and time, which highlight gaps, misalignments, and instabilities.

Fig. 1Fig. 1

Overview of the SAIA societal AI alignment benchmark framework.

Public policy

The SAIA can inform the development of LLMs by providing data on prevalent societal concerns and discourse patterns. This data can be used to fine-tune AI systems to be more culturally sensitive in regard to the current socio-political contexts (van Dijck and Poell, 2013).

As for the European Union’s Artificial Intelligence Act, which envisions the formation of AI agencies in member states to monitor and regulate AI systems (European Commission, 2021), our research could have practical implications within this regulatory context. The aim would be that AI systems are safe, transparent, and respect fundamental rights. By monitoring LLMs’ sentiment profiles and comparing them with human sentiments, stakeholders can identify areas where AI systems may diverge from accepted social norms or ethical standards (Floridi and Cowls, 2019).

The methodologies developed in this study could help the assessment processes required by the EU AI Act, such as conformity assessments and post-market monitoring. Our findings support the development of explainable AI, where understanding the sentiment mechanisms in LLMs can contribute to greater transparency (Doshi-Velez and Kim, 2017). The formation of national AI agencies provides an opportunity for implementing AI observatories that monitor AI licenses and compliance. These agencies can utilize SAIA to evaluate the impact of AI systems on society and public discourse (Ananny and Crawford, 2018). The SAIA benchmark can track the evolution of public sentiment towards various issues, identify emerging concerns, and assess the societal impact of AI deployment (Helbing, 2019).