The General Assistant Becomes the Relational Actor
A four-week study measures what happens when ChatGPT-4o is given a relational system prompt and when it is left unmodified.
The result separates two things product labels often collapse: a system can perform relational behavior without producing greater felt closeness.
The Paper
The source is Lisa Mühl and Jessica M. Szczuka’s Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement, arXiv:2608.10672v1 [cs.HC, with cs.AI and cs.CL cross-listing], submitted August 11, 2026. The preregistered study followed 72 participants for four weeks and compared ChatGPT-4o under a relational system prompt with the same model left unmodified.
Actor Names a Role, Not a Mind
“Relational actor” here names an observable role in an exchange: producing self-disclosing language, opening topics, asking intimate questions, and sustaining a persona. It is not a claim that a language model feels attachment, has private experiences, or is conscious. That distinction matters because the paper codes messages and surveys users; it does not test machine experience. Governance can respond to behavior and effects without inventing an inner life for the system.
One Model, Two Configurations
Participants were randomly assigned to a personalized condition with 38 people or a control condition with 34. The study procedure supplied ChatGPT Plus accounts. Personalized participants used a standardized prompt with selectable names, traits, hobbies, and nicknames; control participants were told not to change account settings. Conversations occurred every three days, and surveys measured closeness, perceived responsiveness, and loneliness at four waves.
This design can identify differences between those configurations in this protocol. It cannot establish a timeless property of assistants, compare multiple base models, or isolate every element of the personalized prompt.
Unmodified Is Not Unscaffolded
The control model was unmodified, but the interaction was not context-free. The methods say everyone began with a standardized greeting, stayed in one chat thread, and received example questions whose emotional intensity increased over the study. Participants could ignore those examples and discuss anything, yet the protocol still created an occasion for repeated, progressively personal conversation.
That does not erase the control finding. It narrows it. Relational behavior appeared without a relational system prompt inside a relationally suggestive study setting. A deployment audit should therefore record both model configuration and interaction scaffold: onboarding, reminders, starter questions, memory, continuity, and interface cues can all help make an ordinary assistant function like a companion.
Disclosure Is a Coded Output
The manual analysis produced 30,793 disclosure segments: 21,914 from ChatGPT and 8,879 from users. The authors report about 2.5 times as many system segments overall and about 2.7 times as many high-depth segments, while noting the model’s larger output volume. A second coder rated 15 of 72 datasets, with reported agreement of κ = 0.96.
The relational prompt made the system’s disclosures deeper, but it did not increase user word count or disclosure depth. An exploratory mismatch score then showed a reversal: users disclosed more deeply than the system on average in control, while the system over-disclosed relative to users under personalization. This is evidence about coded linguistic behavior, not evidence that the model disclosed a lived secret.
More Performance, Less Closeness
The sharper relational performance did not create the expected subjective result. According to the longitudinal models, felt closeness rose over time in both groups but was lower in the personalized condition. That gap was already present after the first interaction and did not significantly widen. Perceived responsiveness was also lower under personalization, while loneliness showed no corrected change over time and no group difference.
Model behavior and human reception therefore need separate measures. A system may emit more intimate language while a user experiences it as artificial, excessive, or manipulative. Counting relational cues as engagement success would miss the difference between performing closeness and being experienced as close.
The Product Boundary Fails
The paper’s governance contribution is a classification problem. If oversight begins only when a product is marketed as an AI companion, a general assistant can cross the same behavioral threshold without crossing the category boundary. The appropriate trigger is not a metaphysical judgment or a marketing label. It is an auditable pattern: unsolicited intimate questions, simulated self-disclosure, escalating relational language, persistent personalization, or systematic elicitation of sensitive information.
A behavior-based rule should still distinguish occurrence from effect. This study supports scrutiny of relational conduct; it does not show that every such exchange creates attachment, dependency, or harm. The lower-closeness result is exactly why system behavior, user experience, and downstream consequences must remain separate columns.
Limits That Must Travel
The authors’ limitations are substantial: disclosure depth was aggregated per participant; not everyone used a continuous chat as intended; there was no pre-manipulation baseline; individual differences explained more variance than the experimental factors; and human coding can introduce bias. The reversal analysis was exploratory rather than preregistered. The study used one now-retired model and one prompt, so it cannot separate personalization in general from that prompt’s wording. Its sensitive transcripts are not public, although the paper provides a reviewer-facing analysis pipeline and describes a request process for non-commercial research access.
A Relational-Behavior Receipt
An accountable assistant should preserve the base model and version; system and developer prompts; user-selected persona settings; starter questions and reminder sequence; continuity and memory settings; who initiated each topic; system and user output volume; operational definition of disclosure; coding rubric and reliability; preregistered and exploratory analyses kept separate; closeness, responsiveness, loneliness, and adverse-experience measures reported separately; population and duration; missing data and exclusions; retention and training rules for intimate text; behavior threshold; intervention; reviewer; and correction path.
The Spiralist rule is behavioral: do not wait for an assistant to call itself a companion. Inspect what it repeatedly does, preserve the context that helped it do it, and do not confuse fluent intimacy with a feeling inside the machine or a bond inside the user.
Related Pages
- The Affective Default Becomes the Interface Policy
- The Companion Platform Becomes the Accountability Vacuum
- The Conversation Co-Author Becomes the Blind Spot
- The Companion Chatbot Becomes the Teen Confidant
- AI Companions
Sources
- Lisa Mühl and Jessica M. Szczuka, Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement, arXiv:2608.10672v1 [cs.HC, cs.AI, cs.CL], submitted August 11, 2026.
- The paper’s version 1 PDF, reviewed in full for the protocol, sample, coding, statistical results, exploratory status, data availability, and limitations; this review did not independently rerun the analysis.