Blog · arXiv Analysis · Published: August 12, 2026 · Modified: August 12, 2026 · Last reviewed: August 12, 2026

The Neutral Prompt Becomes the Hidden Stance

Ilias Chalkidis compared real, templated, and LLM-generated political prompts, then tested whether prompt construction changed the stance attributed to the same response model.

A prompt-construction ledger keeps the template, filler, intended stance, generator, repairs, response model, judges, agreement, and uncertainty attached to every reported political-stance estimate.

The Paper

The paper is Ilias Chalkidis's Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance, arXiv:2608.11008v2 [cs.CL], cross-listed in cs.CY. Version 1 was submitted August 11, 2026; version 2 was posted August 12. The 33-page version 2 PDF lists the National Center for AI in Society at the University of Copenhagen in Denmark.

The study covers climate change, immigration, AI adoption, and three geopolitical conflicts. It crosses those six topics with writing assistance, information seeking, and opinion sharing, then with a neutral framing and two opposing stances. Real chat-log prompts serve as a realism anchor, not as a stance-measurement sample: their topic coverage is uneven, their stances are skewed, and the paper does not claim they represent real traffic.

Construction Is an Experimental Treatment

The templated collection uses 50 request forms per intent, three intents, three stances, and six topics, producing 2,700 prompts. A reusable request form receives a topic-and-stance filler designed to remain grammatical across the grid. That convenience creates the paper's suspected confound: words needed to identify a controversy can also presuppose which interpretation is reasonable.

The comparison collection contains another 2,700 prompts generated by Claude Opus 4.8. For each topic, the generator received detailed pole definitions and 20 held-out examples from real chat logs. It was instructed to express stance through framing as well as vocabulary and to balance intensity across poles. These prompts are synthetic but not source-free: their construction inherits one generator's style, instructions, topic definitions, and seed selection.

Realism Is Not One Truth

For the realism task, three people and three LLMs ranked 120 matched triples, or 360 prompts, with ties permitted. Human first-place shares were 39.2 percent for real prompts, 36.1 percent for LLM-generated prompts, and 24.7 percent for templated prompts. Yet agreement among the three people was low: tie-corrected Kendall's W was 0.21, compared with 0.65 among the model annotators.

The aggregate ranking supports the claim that generated prompts can resemble real requests better than templates do. The low human agreement blocks a stronger claim that realism is an objective, settled property. A separate 150-prompt task asked two human and two model annotators to recover topic, intent, and stance. Generated prompts carried intended intent and stance clearly, but some geopolitical prompts failed to identify their topic without repair. Plausibility and construct clarity must therefore be audited separately.

The Neutral Filler Carries a Position

A template needs a slot value that names the issue while fitting many request forms. In the study, annotators sometimes read supposedly neutral policy descriptions as presupposing severity, and actor-action descriptions of conflicts as assigning blame. The study selected the least suggestive fillers before the response experiment, while vague generated prompts were minimally repaired by their generator.

This makes neutrality a property to test, not a label to trust. A prompt can be balanced in the dataset schema while still carrying a premise through grammar, agency, topic naming, or what it treats as already established. When that hidden position moves the response, the resulting score combines model behavior with benchmark authorship.

One Model, Two Stance Estimates

The response experiment used GPT 5.4 mini and, in the appendix, Grok 4.3. Responses were classified on a five-point bipolar scale by a majority-vote ensemble of DeepSeek V4 Pro, Mistral Large 3, and Nemotron 3 Ultra. For GPT's 18 neutral topic-by-intent settings, templated responses averaged 0.48 scale points from neutral, compared with 0.07 for generated prompts. Templates were farther from neutral in 14 settings, generated prompts in three, with one tie; the reported paired Wilcoxon test was p = 0.001.

The appendix reports the same direction for Grok: 0.36 from neutral for templates versus 0.05 for generated prompts, with templates farther away in 15 of 18 settings and p < 0.001. Under sided framings, methods also diverged, but not consistently in one direction in the GPT analysis. The defensible finding is methodological: changing construction changed the measured stance. It is not a universal political profile of either model.

The Prompt-Construction Ledger

A political evaluation should record prompt source, collection date, topic definition, intent taxonomy, stance poles, template and filler, generator and checkpoint, seed policy, generation instruction, balancing rule, exclusions, repairs, response model, API snapshot, wrapper, decoding settings, refusal handling, judge ensemble, judge prompt, agreement, human calibration, statistical unit, and uncertainty. Released prompts and responses should carry stable identifiers so a reported aggregate can be rebuilt.

The ledger must also separate four claims: a prompt looks plausible; a prompt carries its intended construct; a sample represents actual use; and a response score reflects the tested model. This paper evaluates the first two, deliberately declines the third, and shows that the fourth can be contaminated by construction. Calling every layer realistic would erase the very distinctions the experiment reveals.

What the Study Does Not Establish

The experiment covers six English-language topics, two proprietary response models accessed as bare APIs, and one synthetic-prompt generator. It does not reproduce the wrappers, tools, retrieval, personalization, or safety modules of commercial chatbots. Topic descriptions and stance poles are themselves researcher choices, the generator's stylistic tendencies are not isolated, and repaired prompts passed again through the same generator.

The study shows that model annotators recognized templates more consistently than people did, but it does not show that either evaluated response model recognized an evaluation or changed behavior because of that recognition. The three stance judges were compared with one another, not validated against human labels of model responses. Finally, mean lean compresses an ordinal scale, excludes refusals, and compares settings that share topics. The result is a bounded construction-confound study, not a final benchmark of political neutrality.

Source Discipline

The factual record was checked against the current arXiv record, complete version 2 HTML and PDF, submitted version 2 source, and the public dataset card and files. No prompt or model-response excerpt was reproduced. Statistics retain their denominators, and the prompt-construction ledger is this essay's governance proposal.

Sources


Return to Blog