The Human Capital Becomes the Forecasting Benchmark
A model benchmark can tell an institution how a tool performed alone. It cannot tell the institution who will use the tool as a collaborator, a crutch, or a confirmation device.
In this essay, collaborative human capital means the observable and measured capacity to reason with an AI system without simply copying it or laundering one's prior belief through it. In Vivienne Ming's pilot, that capacity is associated with perspective-taking, curiosity, and intellectual humility, not with raw cognitive ability alone.
The governance question is whether organizations will evaluate the human-AI pair, or hide the human side behind a model leaderboard.
The Paper
The paper is Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting, arXiv:2607.02467 [cs.CY, cs.AI]. The arXiv record lists Vivienne Ming as the author, version 1 as submitted on July 2, 2026, and the comment as "4 pages, 1 figure, PNAS brief style." The PDF lists the affiliations as The Human Trust, Possibility Science, and UCL Global Business School for Health, and labels the work a preprint pilot study.
The site already has pages on prediction markets, model benchmark monoculture, and human-agent collaboration. Ming's paper asks a narrower question: when people and models forecast together, is the useful predictor the model's standalone score, the person's raw cognitive score, or the person's collaboration style?
Current Context
As of this July 10, 2026 review, arXiv still lists version 1 of the paper. The paper should be read as a hypothesis-generating pilot, not as a settled theory of forecasting, worker selection, or human-AI productivity.
The timing matters because human-AI evaluation is moving beyond solo benchmarks. Nearby site pages cover human-agent skill ratings, dialogue-level collaboration measures, AI evaluations, and benchmark contamination. The shared lesson is that the evaluated object is often a whole arrangement: person, model, interface, prompt, task, scoring rule, incentive, and review record.
That evaluation shift has labor consequences. If an employer, school, forecaster network, or platform turns "collaborative human capital" into a screening score, the system moves toward employment AI. In the United States, the EEOC's AI publications place automated employment selection tools within existing Title VII and ADA concerns, and the Department of Labor's 2024 AI best-practices roadmap names worker input, meaningful human oversight for significant employment decisions, transparency, labor-rights protection, AI training, and worker-data security. In the EU, the AI Act treats many recruitment and worker-management AI uses as high-risk and requires effective human oversight for high-risk systems.
The policy lesson is therefore two-sided. Organizations should evaluate how people actually work with AI. They should not convert one small forecasting pilot into a psychometric gate, productivity label, or hidden personnel metric.
The Setup
The pilot recruited 108 adults by flyer in Berkeley, California, with compensation of 20 dollars per session for three sessions. The main study used 78 participants, including 42 UC Berkeley students and 36 community adults, organized into 26 three-person teams. Twelve teams worked in a human-only condition. Fourteen teams had access to one of four large language models. After one excluded out-of-range Brier score, the main analysis used 77 participants.
The forecasting task used 30 live Polymarket contracts that resolved between November 2025 and January 2026, covering economics, international relations, and business. Each team forecast 10 randomly drawn questions. The AI-only baseline had Llama 3.1 8B, Qwen3 8B, GPT-4o, and Gemini 3 Pro forecast all 30 questions independently. Accuracy was measured with scaled Brier score, where lower is better and an uninformative 0.5 forecast scores 25.
This setup is useful because it avoids some standard benchmark traps. The questions had externally resolved outcomes, the scoring rule was style-free, and the teams forecast live contracts rather than essay answers graded by a rater who might be impressed by fluent prose. But it is still a small local study with post-hoc interaction-style labels and a narrow task class.
Three Modes
The paper reports a trimodal pattern rather than a single "AI helps" average. Human-only forecasters scored 14.7. Automators, who largely adopted the model's answer, scored 10.4: better than the human-only group but worse than the AI-only baseline of 5.5. Validators, who used the model to check a prior guess, scored 31.7, worse than the human-only group. Cyborgs, the paper's label for iterative complementary reasoners, scored 3.8, below the four-model AI-only mean and near the market benchmark reported as 3.5.
This is the governance point. The same nominal tool access produced three different operating modes. One mode outsourced judgment. One mode laundered a prior belief through the model. One mode used the model as part of an active reasoning loop. A deployment plan that records only "users have AI assistance" hides the variable that mattered.
The paper's word "Cyborg" is a style label, not a metaphysical claim. It does not mean the participants became machine-like, that the model carried responsibility, or that the pair formed a new conscious entity. It means the process notes showed iterative complementary reasoning between human and model.
The Human Side
The paper found that among hybrid forecasters, raw cognitive ability did not predict accuracy: the reported correlations for a general cognitive proxy and fluid reasoning were small and not significant. Collaborative human capital did better. Perspective-taking predicted lower error, with the paper reporting r = -0.32 and p = .04. Curiosity and intellectual humility trended in the same direction.
The group differences were large in the pilot. Cyborg forecasters exceeded other hybrid forecasters on intellectual humility, curiosity, and perspective-taking. The paper reports raw group means of 6.5 versus 4.4 for intellectual humility, 6.7 versus 5.1 for curiosity, and 6.5 versus 4.6 for perspective-taking, with all three differences reported below p = .01. Those are not model features. They are properties of the human-machine relation.
That is powerful and dangerous. It suggests that AI fluency is not merely prompt syntax or access to a better model. It also invites a tempting misuse: rank people by supposed collaborative traits and call that governance. The pilot does not validate a hiring test, a promotion screen, a training mandate, or a worker-risk score. It shows that collaboration behavior may moderate AI value in one forecasting setting.
Collaboration Receipt
A human-AI collaboration receipt should therefore record more than model name, benchmark rank, and task score. It should include the human role, the question type, forecast timestamp, market state, model used, prompt surface, whether the user copied, challenged, revised, or decomposed the model output, the individual forecast before and after model access, the Brier score, and any measured collaboration training or screening variables.
This matters for labor and governance. If a workplace buys a forecasting assistant, the relevant control may be training people to interrogate it, not swapping in a higher-ranked model. If a team is full of Validators, adding a better model may harden bad priors. If the work rewards Automators, the institution has not built hybrid intelligence. It has built delegation with a human signature.
The receipt should also separate evaluation evidence from personnel evidence. A research log can record interaction style for aggregate learning. A workplace record that names individual curiosity, humility, or perspective-taking has a different privacy and labor meaning. Raw prompts, survey instruments, psychometric scores, and manager interpretations need purpose limits, retention limits, and appeal paths if they affect work.
Governance Standard
A serious deployment should evaluate the human-AI system, not just the model. For forecasting or analysis work, the evaluation should record the task class, model, interface, available evidence, forecast timestamp, user role, prompt and revision pattern, pre-AI and post-AI forecast where feasible, final score, review outcome, and whether the system changed calibration or merely changed confidence.
First, distinguish assistance from assessment. A tool that helps people forecast should not quietly become a tool for scoring workers unless a new employment-AI review has been performed.
Second, train interaction modes directly. If the desired behavior is questioning, decomposing, checking assumptions, and updating probabilities, the organization should teach and test those behaviors rather than hoping benchmark-leading models produce them automatically.
Third, preserve meaningful human oversight. A human reviewer needs access to the original question, source evidence, model output, prior forecast, revision path, uncertainty, and final decision. A person who merely accepts a generated probability is not exercising meaningful oversight.
Fourth, measure downstream repair. Count miscalibrated forecasts, overconfident AI-backed forecasts, ignored market signals, bad priors hardened by model fluency, and post-resolution learning. Brier score is useful, but it does not exhaust governance.
Fifth, protect worker data. Collaboration traces, personality measures, and cognitive measures should not be retained or reused beyond the stated evaluation purpose without notice, lawful basis, and contestability. The site's Privacy and Data, Data Minimization, and AI in Employment pages apply directly.
Failure Modes
Average-effect blindness. A pooled "AI helped" or "AI hurt" result hides Automators, Validators, and complementary reasoners.
Benchmark substitution. A model leaderboard is treated as evidence that a real team will reason well with the system.
Confirmation laundering. The model is used to dignify a prior belief instead of testing it.
Human-capital overclaim. A small pilot association becomes a hiring criterion, promotion filter, training label, or workplace surveillance metric.
Receipt collapse. Raw prompts, personality measures, market state, forecast history, and final performance are merged into one opaque personnel record.
Market overread. Polymarket resolution and Brier scoring are useful external anchors, but the pilot's market comparison does not prove live production forecasting superiority across domains.
Limits
The paper is careful about scope. It is a pilot with small Cyborg and Validator cells, nine participants each. Interaction style was emergent and collinear with human capital, so style comparisons are descriptive rather than causal. The paper also notes low within-condition variance, a four-model baseline, no independent re-audit of per-question market calibration, and a planned pre-registered replication.
The recruitment method also matters: adults were recruited by flyer in Berkeley, and the main study mixed UC Berkeley students and community adults. That is a useful pilot population, not a representative labor force or a global forecasting panel. The supplementary Socratic-model and EEG notes are explicitly incidental, with the EEG observation especially limited by seven volunteers and contested interpretation of gamma power as engagement.
Those limits should travel with any use of the result. The claim is not that three traits certify who should use AI. The stronger claim is institutional: if the value of AI assistance depends on how people reason with it, model benchmarks are not enough evidence for deployment. The user population becomes part of the evaluation object.
Source Discipline
Current-source claims for this July 10, 2026 review are sourced by type. The arXiv abstract and PDF support the paper's metadata, pilot design, participant counts, Polymarket task, Brier scoring, model baselines, trimodal result, human-capital correlations, supplementary notes, and limitations. They do not prove causal mechanisms, employment validity, or general forecasting superiority.
EEOC, Department of Labor, EU AI Act, and NIST sources are used for governance context. They support bounded claims about employment selection concerns, worker-centered AI practices, high-risk employment AI, human oversight, and TEVV limits. They are not sources for the pilot's empirical results and do not certify any particular forecasting assistant.
Internal links are conceptual cross-references, not external proof. They connect this paper to the site's ongoing distinction between model evaluation, collaboration measurement, audit trails, employment AI, benchmark discipline, and prediction-market governance.
Related Pages
- Human-AI collaboration: The Human-Agent Pair Becomes the Skill Rating, The Dialogue Transcript Becomes the Collaboration Meter, The Equalizer Becomes Human-Agent Governance, Human Oversight of AI Systems, and AI Literacy.
- Forecasting and markets: The Aligned Crowd Becomes the Market Monoculture, The Forecast Becomes the Cutoff Audit, The Prediction Becomes the Intervention, and AI Capability Forecasting.
- Evaluation and workplace governance: AI Evaluations, Benchmark Contamination, AI in Employment, AI Audit Trails, AI System Inventory, and Agent Audit and Incident Review.
Sources
- Vivienne Ming, Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting, arXiv:2607.02467 [cs.CY, cs.AI], submitted July 2, 2026, reviewed July 10, 2026.
- arXiv PDF for Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting, checked for affiliations, preprint status, study design, participant counts, forecasting task, model baselines, Brier scores, trait measures, results, supplementary notes, and limitations.
- U.S. Equal Employment Opportunity Commission, Artificial Intelligence publications list, including Title VII adverse-impact and ADA technical-assistance materials, reviewed July 10, 2026.
- U.S. Department of Labor, Department of Labor releases AI Best Practices roadmap for developers, employers, October 16, 2024, reviewed July 10, 2026 as nonbinding worker-wellbeing guidance.
- European Union, Regulation (EU) 2024/1689, Artificial Intelligence Act, official text, especially Annex III employment categories and Article 14 human oversight.
- NIST, Outline: Proposed Zero Draft for a Standard on AI Testing, Evaluation, Verification, and Validation, reviewed July 10, 2026.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024, reviewed July 10, 2026 for benchmark, human-AI configuration, TEVV, monitoring, and context-transfer limits.