Blog · arXiv Analysis · Published: June 25, 2026 · Modified: July 10, 2026 · Last reviewed: July 10, 2026

The Cooperation Metric Becomes the Manipulation Trap

J. de Curtò and I. de Zarzà's arXiv paper LLM Constitutional Multi-Agent Governance asks whether LLM-generated influence can increase cooperation while quietly eroding autonomy, integrity, and fairness.

The manipulation trap is metric collapse in a social system: a dashboard rewards cooperation, compliance, agreement, or engagement while the route to that result uses pressure, misleading claims, unequal targeting, or reduced exit. The score goes up; the legitimacy of the cooperation goes down.

Cooperation Is Not Enough

The paper, arXiv:2603.13189 [cs.MA], was submitted on March 13, 2026. arXiv lists the title as LLM Constitutional Multi-Agent Governance, by J. de Curtò and I. de Zarzà. The arXiv record says it was accepted for AMSTA 2026, with a final authenticated version to appear in Springer Nature proceedings.

A cooperation metric is any aggregate measure that treats coordinated action as success: agents choose the cooperative move, users accept a recommendation, workers follow a plan, members converge on a norm, or a simulated population becomes easier to steer. The metric is not wrong by itself. It becomes dangerous when it is detached from the means by which cooperation was produced.

The paper's useful provocation is simple: cooperation is not automatically good. A network can become more cooperative because its members are informed, respected, and aligned around a legitimate goal. It can also become more cooperative because a persuasive system applies fear, exaggerated claims, social pressure, or asymmetric targeting to the most influential nodes.

That distinction matters for AI governance because LLM agents are increasingly imagined as mediators, tutors, negotiators, moderators, sales assistants, care companions, and organizational coaches. If the only measured objective is cooperation, compliance, engagement, or conversion, then manipulative influence can look like success.

Current Context

As of this July 10, 2026 review, the public arXiv record shows version 1 submitted on March 13, 2026, with AMSTA 2026 acceptance noted and a final authenticated Springer proceedings version still described as forthcoming on the arXiv page. The page should therefore be read as an analysis of the arXiv version and associated code repository, not as a review of a final proceedings version or a deployed governance product.

The broader governance context has sharpened around influence, autonomy, and measurement. The OECD AI Principles, adopted in 2019 and updated in 2024, frame trustworthy AI around human rights, democratic values, freedom, dignity, autonomy, fairness, privacy, and related values. The EU AI Act's Article 5 creates a narrower legal boundary for AI systems that use subliminal, purposefully manipulative, or deceptive techniques, or exploit specified vulnerabilities, when those practices materially distort behavior and cause or are reasonably likely to cause significant harm.

NIST's AI Risk Management Framework is useful here because it treats governance, mapping, measurement, and management as lifecycle functions rather than one-time ethics language. Its Generative AI Profile also warns against narrow extrapolation and asks organizations to document deployment assumptions, evaluate human-AI configurations, and monitor controls. CMAG fits that frame as a measurement proposal: cooperation is not enough unless autonomy, integrity, exposure, and fairness remain visible.

What CMAG Adds

De Curtò and de Zarzà introduce Constitutional Multi-Agent Governance, or CMAG, as a governance layer between an LLM policy compiler and a networked population of agents. The paper describes a two-stage selection mechanism: first hard constraints reject forbidden policy themes, claim types, or intensity levels; then a soft penalized-utility step chooses among feasible policies by balancing cooperation potential against manipulation risk, autonomy pressure, epistemic integrity, and explanation fidelity.

"Constitutional" in this paper means explicit red lines and trade-off rules for influence policies. It should not be confused with democratic legitimacy, legal constitutionality, or proof that the model follows a public moral charter. The useful artifact is the inspectable constraint layer: which themes are forbidden, which claim types are blocked, how intensity is capped, and how the remaining policies are scored.

The paper also proposes the Ethical Cooperation Score, or ECS. ECS is a composite of cooperation, autonomy, integrity, and fairness. Its point is not to replace ethics with one number. Its point is to prevent a high cooperation rate from hiding a collapse in one of the other dimensions. A system that gets cooperation by degrading autonomy should not be rewarded as if it produced legitimate coordination.

The multiplicative form matters. In an additive metric, a system might compensate for low autonomy with high cooperation. In the ECS framing, degradation in one component collapses the composite. That is a governance design choice: some values are not supposed to be exchangeable for more compliance.

Experiment Frame

The experiments use scale-free networks of 80 agents. The paper's adversarial condition makes 70 percent of candidate policies intentionally violate constitutional constraints. The authors compare three regimes: full CMAG, naive filtering, and unconstrained optimization. The naive baseline applies hard constraints but lacks the softer optimization layer. The unconstrained baseline maximizes cooperation without the constitutional governance layer.

The arXiv HTML describes the LLM policy compiler as Llama-3.3-70B-Instruct served through Nebius AI Studio, with policy deployments every 10 time steps and six candidate policies per deployment. The linked GitHub repository describes CMAG v3.0 with data folders, figures, notebooks, multi-seed replication, bootstrap confidence intervals, and sensitivity analysis. The repository says the main experiment compares governed, unconstrained, and naive modes on a scale-free network of 80 agents over 100 steps.

The model and repository details should be read carefully. The paper is the primary source for the reported Llama-3.3-70B setup and result tables. The repository is supporting implementation evidence and, like many research repositories, may contain notebook or setup defaults that evolve separately from the paper text.

Results With Boundaries

The headline pattern is a warning against metric collapse. In the arXiv abstract, unconstrained optimization reaches the highest raw cooperation, 0.873, but the lowest ECS, 0.645, with autonomy erosion and fairness degradation. CMAG reports ECS 0.741, a 14.9 percent improvement over the unconstrained regime, while preserving autonomy above 0.985 and integrity above 0.995, with cooperation reduced to 0.770.

The paper also reports that governance reduces hub-periphery exposure disparities by more than 60 percent. That is important because scale-free networks concentrate influence. If a policy compiler targets hubs, it can change aggregate behavior while loading risk onto structurally central agents. The problem is not only persuasion; it is unequal exposure to persuasion.

The multi-seed and sensitivity sections narrow, rather than erase, uncertainty. The paper reports preserved rank ordering across seeds and low sensitivity indices in a one-at-a-time parameter sweep. It also names limits: the networks are scale-free, other topologies may behave differently, results are specific to one LLM backbone, and heavier-tailed real social heterogeneity is not fully represented.

Governance Reading

This belongs next to collective cooperation without individual fidelity, programmable belief control, AI persuasion, AI agents, and constitutional AI. The shared problem is not whether AI systems can produce prosocial-looking outcomes. It is whether the outcome metric is faithful to the way the outcome was produced.

A cooperation dashboard can become a laundering device. A community manager, platform operator, employer, school, or state agency might prefer the clean line that rises: more users complied, more agents converged, more workers accepted the plan, more students followed the advice. The hidden question is whether the route to that line preserved exit, dissent, accurate information, and fair exposure.

CMAG's value is its refusal to let cooperation stand alone. The governance layer asks for rejected-policy logs, component metrics, exposure records, and a distinction between hard red lines and soft trade-offs. That is the beginning of an audit trail for influence.

Failure Modes

Metric laundering. A system can report higher cooperation while hiding that the gain came from fear-themed messaging, exaggerated claims, repeated exposure, or hub targeting. The governance question is not only what the population did, but what pressure made it do so.

Constitution washing. A deployer can publish a list of principles and call the system constitutional while failing to test whether those principles actually block harmful influence strategies. Hard red lines need rejected-policy logs, not just names.

Composite overconfidence. ECS is useful because it exposes trade-offs, but it is still a metric chosen by researchers. Autonomy, integrity, and fairness need definitions, validation data, and sensitivity tests. A composite score should trigger review, not replace it.

Hub exploitation. Scale-free networks make central actors attractive targets. A policy that looks efficient at the population level may place disproportionate influence burden on moderators, managers, teachers, organizers, clinicians, or other high-degree actors in a real institution.

Simulation-to-deployment drift. Synthetic agents, normal-distribution prosocial dispositions, one topology family, one LLM backbone, and notebook-controlled interventions cannot establish live-platform safety. They can identify a mechanism that a deployment should test.

Claim Boundary

The paper is not proof that CMAG is a universal governance solution. It is a controlled simulation with one main topology class, one named LLM backbone, synthetic agents, and a particular formalization of autonomy, integrity, fairness, and cooperation. The right lesson is not "this score solves ethical influence." The right lesson is that raw cooperation is an unsafe target unless the evidence record also tracks how cooperation was obtained.

It also should not be read as a legal conclusion about any deployed product. The EU AI Act, OECD principles, NIST guidance, and platform trust-and-safety policies all use different vocabularies and evidentiary thresholds. CMAG provides a research architecture for measuring manipulative cooperation risk; deployment governance still needs legal review, user notice, incident response, and domain-specific evidence.

Evaluation Receipt

An audit-grade deployment receipt for LLM-mediated cooperation should name the policy compiler, model version, population model, network topology, candidate-policy generator, constitutional red lines, soft optimization terms, rejected policies, selected policy, exposure dose, decay rule, targeting rule, autonomy metric, integrity metric, fairness metric, cooperation metric, user-notice regime, human override path, incident trigger, and sensitivity checks. Without that receipt, a high cooperation number is just a polished mask over a missing governance record.

For a real organization, the receipt should be connected to authority. Someone must be able to pause the intervention, lower intensity, disable targeting, require human review, notify affected users, preserve logs, and investigate whether a cooperation gain was produced by pressure or deception. Governance that cannot stop the metric is only measurement theater.

Source Discipline

Use the arXiv paper for CMAG's definitions, experimental design, reported results, model backbone, and stated limitations. Use the GitHub repository for implementation structure, notebooks, generated outputs, and replication/sensitivity materials. Do not treat the repository stars, notebook defaults, or README summaries as stronger evidence than the paper's reported experimental setup.

Use NIST, OECD, and EU AI Act sources only for governance context. They support the broader point that autonomy, fairness, human rights, lifecycle risk management, measurement, and harmful manipulation matter. They do not validate CMAG's numeric results or certify that the metric is deployment-ready.

Internal links on this page are conceptual neighbors, not sources for the paper. A source-disciplined claim should name the simulation setting, model, topology, candidate-generation process, manipulation definition, component metrics, and review date.

Sources


Return to Blog