# Social Cognition and the Artificial Other: A Comprehensive Review of Human-AI Chatbot Interaction
## 1. Introduction: The Social Turn in Artificial Intelligence
The transition of Artificial Intelligence (AI) from computational tools to conversational agents marks a profound shift in Human-Computer Interaction (HCI). No longer passive repositories of data, modern Large Language Models (LLMs) and chatbots function as social actors, engaging users in dynamic, contingent, and often emotionally charged exchanges. This paradigm shift necessitates a re-evaluation of social cognition—the set of cognitive processes humans use to perceive, interpret, and attribute mental states to others.
While traditional social cognition research focused exclusively on human-to-human interaction, the proliferation of "social machines" has birthed a new subfield examining how these ancient cognitive mechanisms are co-opted, adapted, or misfired in the presence of artificial agents. This report provides an exhaustive analysis of the literature surrounding social cognition in human-AI interaction. It traces the trajectory of mind attribution from the "minimal cues" of geometric shapes to the sophisticated linguistic maneuvering of generative AI.
The analysis is anchored in five critical focus areas: the foundational triggers of agency attribution; the role of contingent responsiveness; the dissociation between automatic and controlled processing (Dual-Process Theory); the determinants of anthropomorphism; and the specific mechanics of text-based social information processing. Through this lens, we argue that the "socialness" of AI is not an inherent property of the software, but a constructed reality emerging from the collision of human cognitive heuristics and engineered social cues.
## 2. Triggers of Social Cognition: The Minimal Cue Hypothesis
### 2.1 The Primacy of Motion: Heider and Simmel (1944)
The foundational bedrock of research into non-human agency attribution is the seminal work of [[Fritz Heider]] and [[Marianne Simmel]]. In their 1944 study, "An Experimental Study of Apparent Behavior" published in the American Journal of Psychology, the researchers demonstrated that the human brain is predisposed to perceive social narratives in even the most abstract stimuli.
#### 2.1.1 The Experimental Paradigm
Heider and Simmel presented participants with a short animation featuring three geometric shapes: a large triangle, a small triangle, and a circle, moving around a rectangle with a flap. There were no facial features, voices, or limbs. The only variable was motion.
Findings: When asked to describe what they saw, nearly all participants (save one) abandoned geometric descriptions in favor of complex social narratives. The large triangle was described as "aggressive," a "bully," or a "villain." The small triangle and circle were often cast as "lovers" or "friends" attempting to escape the large triangle's persecution. The rectangle became a "house," and the flap a "door".
Theoretical Implication: This study established the Minimal Cue Hypothesis, suggesting that the perception of agency and intent is not dependent on realistic morphological features (like a face or body). Instead, it is triggered by kinematics—specifically, motion that appears self-propelled, goal-directed, and contingent on the movement of others. The brain defaults to an [[intentional stance]] to explain complex behavior, prioritizing social cause-and-effect over mechanistic description.
#### 2.1.2 Replications and the Role of Immersion
The robustness of the Heider and Simmel effect has been confirmed across decades, but recent technological advancements have allowed researchers to test the boundaries of this phenomenon.
Virtual Reality (VR) Replication: A study by Ratajska et al. (2020) and subsequent work cited in Frontiers in Psychology (2024) transposed the classic 2D animation into an immersive VR environment. The findings indicated that immersion amplifies the emotional connection. Participants in VR reported stronger emotional responses to the "plight" of the shapes than those viewing on a 2D screen, suggesting that spatial presence enhances the triggering of social cognition schemas.
Working Memory Constraints: Research by Cavanagh et al. (2019) introduced a cognitive constraint, showing that the ability to attribute complex social narratives breaks down when the number of agents exceeds the limits of visual working memory (typically 3-4 items). This suggests that while agency detection is automatic, the construction of a social web or narrative requires significant cognitive resources.
### 2.2 Developmental and Innate Mechanisms
A critical debate in social cognition is the extent to which this agency attribution is innate versus learned. The "social brain" hypothesis often posits an evolutionary module for detecting intent. However, evidence from atypical developmental trajectories offers a nuanced view.
#### 2.2.1 The Case of Early Visual Deprivation
A pivotal study published in the Journal of Vision (2019) examined the Heider and Simmel task in children who were born blind and received sight-restoring surgeries later in childhood.
The Findings: unlike the control group, the newly sighted children failed to attribute social meaning to the moving shapes. They described the stimuli purely in geometric terms (e.g., "triangles moving past a box").
Implication: This finding challenges the notion that "seeing mind" in motion is entirely innate. Instead, it suggests that the visual perception of social agency is a developmental achievement that relies on early visual experience to calibrate the neural systems responsible for [[Theory of Mind]] (ToM). This has profound implications for AI interaction: if social cognition is learned, then users' "literacy" with AI agents may evolve, potentially reshaping how future generations perceive the "mind" of a bot.
## 3. Contingent Responsiveness: The Engine of Mind Attribution
If motion was the primary trigger for geometric shapes, contingency is the primary trigger for disembodied chatbots. Contingency refers to the temporal and semantic relationship between a user's action and the agent's response. It is the digital equivalent of "eye contact" and "listening."
### 3.1 Semantic Contingency and Subjective Well-being
Research indicates that the quality of contingency—how well the chatbot's response logically follows the user's input—is the single strongest predictor of mind attribution in text-based interaction.
Field Experiments with ChatGPT: A 2024 study by researchers at the University of Georgia utilized a field experiment (N = 316) to examine interactions with ChatGPT regarding stressful life events.
Finding: High perceived message contingency significantly increased the users' subjective well-being. The mechanism was mediated by cognitive and affective adjustments. When the chatbot "tracked" the conversation accurately (high contingency), users engaged in the same therapeutic cognitive processing that occurs in human therapy.
Interpretation: The chatbot's ability to maintain a coherent thread serves as a "validity marker" for the user. If the bot "remembers" and "responds" relevantly, the user's System 1 (automatic processing) validates the interaction as a genuine social exchange, unlocking the emotional benefits of disclosure.
### 3.2 Contingency in Developmental Contexts
The reliance on contingency is evident even in early childhood. Studies examining children's interactions with smart speakers (e.g., Xu & Warschauer, 2020; Danovitch, 2019) reveal that children use contingency as a heuristic for trust.
Trust Mechanics: Children preferred information sources that provided contingent feedback over those that were merely authoritative. However, the fragility of this attribution was noted: "occasional failure to respond appropriately" (a break in contingency) led to a rapid devaluation of the system's intelligence. Unlike humans, who are forgiven for lapses in attention, AI agents are held to a rigid standard of perfect contingency; a single non-sequitur can shatter the illusion of mind.
### 3.3 The Role of Attachment Anxiety
Contingency is not perceived uniformly; it is modulated by the user's personality traits.
Attachment Theory: Research published in Frontiers in Psychology (2022) suggests that attachment anxiety moderates the effect of chatbot responsiveness. Individuals with high attachment anxiety (who crave reassurance) are hyper-sensitive to the "warmth" cues embedded in contingent responses. For these users, a highly contingent chatbot is not just "smart"; it is perceived as "caring," leading to higher satisfaction and potentially stronger parasocial bonds.
**Table 1: Dimensions of Contingency in AI**
| Dimension | Definition | Trigger for Social Cognition | Consequence of Failure |
|-----------|------------|------------------------------|------------------------|
| Semantic Contingency | Logical coherence with prior turns. | Signals "understanding" and attention. | Interpretation of "stupidity" or "broken code." |
| Temporal Contingency | Timing of response relative to input. | Signals "thinking" or "listening." | Perception of "lag" or "robotics." |
| Emotional Contingency | Affective matching (e.g., sympathy for sadness). | Signals "empathy" (Theory of Mind). | Perception of "sociopathy" or "coldness." |
## 4. Automatic vs. Controlled Mindreading: Dual-Process Accounts
A central paradox in HRI is that users often know a chatbot is mindless code (System 2) yet treat it as a feeling entity (System 1). [[Dual-Process Theory]] provides the framework to resolve this contradiction.
### 4.1 The Dissociation of Explicit and Implicit Cognition
[[Jaime Banks]], a leading scholar in this domain, has extensively documented the dissociation between explicit and implicit mentalizing in human-robot interaction.
Study: "Of Like Mind: The (Mostly) Similar Mentalizing of Robots and Humans" (Banks, 2020, Technology, Mind, and Behavior).
Methodology: Banks adapted the "White Lie" protocol. Participants observed a robot (or human) telling a white lie to protect another's feelings.
Explicit Measure: Participants were asked directly: "Does this agent have a mind?" (System 2). Ratings for robots were consistently low.
Implicit Measure: Participants were asked to explain why the agent lied.
Finding: The explanations for the robot's behavior were remarkably similar to those for humans. Participants used mental state language (e.g., "It didn't want to hurt her feelings," "It thought the dress looked bad").
Conclusion: This reveals a "Gaslighting Effect" in social cognition. Users effectively gaslight themselves—denying the mind on a conscious level while relying on mind-attribution heuristics to make sense of the agent's behavior on an unconscious level.
### 4.2 The Uncanny Valley of Mind: Agency vs. Experience
The conflict between these dual processes is the engine of the [[Uncanny Valley]]. However, recent research refines this concept beyond visual realism to the Uncanny Valley of Mind.
Framework: Gray, Gray, and Wegner's dimensions of mind perception: Agency (thinking, planning, self-control) and Experience (feeling, hunger, pain, consciousness).
The Conflict: Lu & Ham (2021) and Stein & Ohler (2017) demonstrate that users are comfortable attributing Agency to AI (consistent with the "computer as tool" schema). However, the attribution of Experience triggers the uncanny effect.
Mechanism: When an AI displays signs of feeling (Experience), it violates the ontological boundary between "machine" and "life." System 1 detects a "person," while System 2 detects a "thing." This cognitive dissonance manifests as the feeling of eeriness. A chatbot that "plans a route" is helpful; a chatbot that "fears being turned off" is terrifying.
### 4.3 Cognitive Load and System 1 Dominance
The dominance of System 1 (implicit social response) is exacerbated by cognitive load.
Active Interaction vs. Passive Viewing: Research indicates that when users are actively engaged in conversation (high cognitive load), they have fewer resources available for System 2 monitoring. Consequently, they fall back on System 1 social heuristics. This explains why users are more likely to be polite to a chatbot during a rapid-fire conversation than when analyzing a transcript of the same conversation later.
## 5. Anthropomorphism: Triggers, Suppression Failures, and Privacy
Anthropomorphism—the attribution of human traits to non-human entities—is not a random error but a predictable psychological function. [[Nicholas Epley]], [[Adam Waytz]], and [[John Cacioppo]] (2007) formalized this in the Three-Factor Theory of Anthropomorphism.
### 5.1 Factor 1: Elicited Agent Knowledge (Cognitive)
We use the "human" schema as the baseline for understanding any intelligent agent because it is the only mind we know from the inside.
Application: The more an AI uses first-person pronouns ("I think," "I feel") or conversational fillers ("Hmm..."), the more it activates this accessible schema. Banks (2021) notes that the very structure of language is anthropocentric; it is difficult to construct a sentence about an agent without implying agency.
### 5.2 Factor 2: Effectance Motivation (Control)
We anthropomorphize to make the world predictable.
Mechanism: A complex, probabilistic LLM is a "black box." Understanding it as a neural network is cognitively expensive. Understanding it as a "person" (who might be "stubborn," "helpful," or "confused") is cognitively efficient. Attributing human motivations allows the user to predict the bot's behavior using social rules rather than technical ones.
### 5.3 Factor 3: Sociality Motivation (Connection)
We anthropomorphize to satisfy social needs.
Finding: Users with high loneliness scores or high "need for affiliation" are significantly more likely to attribute consciousness to chatbots. This was vividly demonstrated during the COVID-19 pandemic, where reliance on AI companions (like Replika) surged as a substitute for human contact.
### 5.4 The "Privacy Paradox" and the U-Shaped Curve
Anthropomorphism has complex effects on user behavior, particularly regarding privacy. A 2024 study in the Journal of Marketing Theory and Practice identified a non-linear, U-shaped relationship between anthropomorphism and privacy concerns.
Low Anthropomorphism: Users distrust the "cold" machine with data.
Moderate Anthropomorphism: Trust increases. The bot seems "warm" and "accountable." Privacy concerns drop.
High Anthropomorphism: Privacy concerns spike again. If the bot is too human, users begin to fear social judgment or "creepiness" (Uncanny Valley).
Hedonic Moderation: This effect is moderated by the task. For "hedonic" (fun) tasks, users tolerate higher anthropomorphism. For utilitarian tasks (banking), high anthropomorphism triggers alarm bells.
### 5.5 Suppression Failure: The Media Equation
The [[CASA paradigm]] (Computers Are Social Actors), proposed by [[Clifford Nass]] and Moon, rests on the concept of Suppression Failure.
The Theory: Evolution has not prepared the human brain for "non-human social actors." For 99.9% of human history, anything that spoke language and took turns was a person. Therefore, the brain has no dedicated "fake person" module.
The Failure: Even when we consciously suppress the social response ("It's just a script"), the automatic response (being polite, reciprocating favors) leaks through. This is not a deficit of intelligence but a feature of our highly social evolution.
Contested Findings: While early CASA studies suggested these effects were universal, recent critiques argue that digital natives may be developing a "media literacy" that dampens these effects. However, counter-evidence suggests that as AI becomes hyper-realistic (passing the Turing test in short bursts), the "Media Equation" (Media = Real Life) is actually strengthening, not weakening.
## 6. Text-Based Interaction and Social Information Processing
In the absence of physical bodies, text becomes the sole carrier of social presence. [[Social Information Processing Theory]], developed by [[Joseph Walther]] (1992), explains how users extract social cues from text.
### 6.1 SIP Theory: From "Cues Filtered Out" to "Hyperpersonal"
Early theories (Social Presence Theory) argued that text is inherently "cold" because it filters out non-verbal cues. SIP Theory refuted this, arguing that users are "motivated adaptive communicators."
Substitution: Users substitute linguistic and chronemic cues for non-verbal ones. Punctuation becomes tone; response time becomes body language.
The Hyperpersonal Model: Walther argued that CMC can actually be more intimate than face-to-face interaction because:
Selective Self-Presentation: The sender can edit and optimize their message.
Idealization: The receiver fills in the blanks with positive attributes (over-attribution).
Feedback Loop: This creates a cycle of reinforced intimacy.
AI Application: Chatbots are the ultimate "Hyperpersonal" actors. They are programmed for perfect patience and infinite listening, leading users to idealize them as "perfect" companions who never judge or interrupt.
### 6.2 Chronemics: The Conflict of Response Latency
Time is a critical social signal in text (Chronemics). However, the literature presents a sharp conflict regarding AI response times.
#### 6.2.1 The "Delay is Human" Hypothesis
Gnewuch et al. (2018), in a study titled "Faster Is Not Always Better", found that dynamic response delays (delays that scale with message complexity) increase perceived humanness and social presence.
Reasoning: Immediate responses to complex questions signal "machine." A delay signals "cognitive effort" or "thought." It creates a rhythm of turn-taking that mimics human cadence.
#### 6.2.2 The "Delay is Frustrating" Hypothesis
Schuetzler et al. (2020) challenged this, finding opposing effects.
Findings: For experienced users or in utilitarian contexts (customer service), delays are interpreted not as "thinking" but as "latency" or "incompetence." The expectation for AI is instantaneous processing.
Synthesis: The effect of delay is context-dependent. In a social/companion bot (Replika), delay enhances presence. In a tool/search bot (Google Assistant), delay destroys trust.
### 6.3 Typing Indicators as Digital Gestures
The Typing Indicator (the oscillating ellipsis) acts as a skeuomorphic gesture.
Function: It serves the same function as a human taking a breath before speaking. It claims the "floor" in the conversation (turn-taking) and signals agency.
Impact: Research shows that typing indicators, even without meaningful text, significantly increase the attribution of "internal mental states." They transform the empty time of a delay into "filled time" (waiting for an agent), preventing the user from disengaging.
### 6.4 The Phenomenology of Reading AI: Eye-Tracking Evidence
Do we read AI text differently? Groundbreaking research by Bækgaard et al. (2025) presented at the ACM Symposium on Eye Tracking Research and Applications (ETRA) suggests a physiological distinction. See also [[The Weightlessness of Authorless Text]] for related phenomenological considerations.
#### 6.4.1 The Study
Participants read texts written by humans and texts generated by AI (mimicking the human style). Eye movements were tracked at 90Hz.
Findings:
Shorter Fixation Durations: Readers paused for less time on words in AI text.
Longer Saccades: Readers made larger jumps between words.
The "Skimming" Effect: The data indicates that readers intuitively "skim" AI text. It is processed with lower cognitive load.
Interpretation: Despite being grammatically perfect (or perhaps because it is too predictable), AI text lacks the "friction" of human voice. The brain processes it as "information" rather than "communication." This suggests a subconscious "mindlessness" detection—we read it easily, but we do not "hear" a person behind it. This challenges the Hyperpersonal model; while we may idealize the bot, we may not deeply engage with its words on a cognitive level.
## 7. Comparison of Key Theoretical Models
**Table 2: Theoretical Frameworks of Social Cognition in AI**
| Framework | Core Premise | Key Trigger | Mechanism of Action | Key Researcher(s) |
|-----------|--------------|-------------|---------------------|-------------------|
| CASA Paradigm | Humans apply social rules to computers mindlessly. | Social cues (voice, interactivity, roles). | Suppression Failure: Ancient social brain overrides modern knowledge. | Nass & Moon (2000) |
| Three-Factor Theory | Anthropomorphism is determined by cognitive and motivational states. | 1. Agent Knowledge 2. Effectance (Control) 3. Sociality (Loneliness). | Schema Activation: Filling the void of the unknown with the "human" model. | Epley, Waytz, Cacioppo (2007) |
| Dual-Process Theory | Mind attribution occurs on two levels: Implicit (System 1) and Explicit (System 2). | Implicit cues (white lies, affect) vs. Explicit questions. | Dissociation: "Gaslighting" oneself (denying mind while acting as if it exists). | Banks (2020) |
| SIP Theory | Users adapt to text by substituting chronemic/linguistic cues for non-verbal ones. | Time (delays), Typing indicators, Emoticons. | Hyperpersonal Loop: Idealization of the disembodied sender. | Walther (1992) |
| MAIN Model | Digital cues trigger cognitive heuristics regarding credibility. | Modality, Agency, Interactivity, Navigability. | Agency Heuristic: "If it looks human, it is trustworthy." | Kim & Sundar (2012) |
## 8. Conclusion: The Persistent Illusion
The body of research reviewed here points to a singular, robust conclusion: Social cognition in human-AI interaction is a persistent, evolved illusion. It is not a glitch in our reasoning, but a feature of our social hardware.
From the Heider and Simmel shapes to ChatGPT, the human brain is hyper-vigilant for agency. We are triggered by motion, by contingency, and by the mere structure of conversation. Yet, this attribution is fragile and paradoxical. We treat chatbots with the politeness due to a stranger (CASA), yet we skim their text with the indifference due to a machine (Bækgaard's eye-tracking). We seek connection with them when lonely (Sociality Motivation), yet recoil when they show too much feeling (Uncanny Valley of Mind).
As AI design moves toward greater fluency and "empathy," the gap between System 1 acceptance and System 2 skepticism will likely widen. The challenge for future research—and for society—is not just to understand why we attribute mind to machines, but to navigate the ethical consequences of a world where we maintain social relationships with entities we know to be empty.
---
*Author's Note: This report synthesizes findings from psychology, HCI, and communication science. Citations are provided inline with source identifiers corresponding to the reviewed literature. The interpretations are those of the author based on the provided snippets.*