The Model Choice Becomes the Dependency Map
Zackary Okun Dunivin traced which large-language-model families researchers report using, rather than merely mentioning, across a large corpus of scientific articles.
A model-dependency map treats each choice as a scientific-infrastructure record: model, release, access path, research role, institutional setting, and the conditions needed to inspect or rerun the work.
The Paper
The paper is Zackary Okun Dunivin's Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science, arXiv:2608.11090v1 [cs.CY], submitted August 11, 2026. The 21-page version 1 PDF lists the Institute for Social Sciences at the University of Stuttgart.
The study began with 21,338,177 unique full-text records in a July 21, 2026 snapshot of S2ORC and supplemented them with OpenAlex metadata. Its dated analytic corpus contains 157,446 papers published from January 2023 through June 2026 and 2,276,136 occurrences classified as model use. Those numbers describe the observable corpus, not every scientific paper published during the period.
A Mention Is Not a Use
The measurement pipeline first recognized candidate model names, then asked whether a matched term really denoted the intended model and whether the paper's authors used it as an instrument or object of study. A third classifier filtered passages resembling AI-disclosure statements, which usually described writing or workflow assistance rather than the scientific use under study.
The identity and author-use classifiers were each trained on a 1,500-passage silver-label set created with DeepSeek-V4-Flash after comparison against separate 150-passage author-coded sets. The disclosure classifier instead used section-derived proxy labels. Five-fold validation reported F1 scores of 0.94 for model identity, 0.93 for author use, and 0.96 for disclosure-like text. These are meaningful checks, but they do not make automated extraction error-free. The important move is conceptual: visibility, citation, workflow assistance, and operational use are different relations.
Family Count Changes the Meaning
The paper separates 63,172 single-family papers from 94,274 papers using two or more families. It treats the split as a crude proxy: a single-family paper is more likely to reflect a narrow applied choice, while a multi-family paper is more likely to benchmark or compare models. The authors explicitly warn that this is not a validated classification of research purpose.
In the 2026 partial-year cohort, 44.0 percent of 16,125 single-family papers used a coded open-weight family. Qwen alone appeared in 22.0 percent, compared with 8.6 percent for Llama, 3.0 percent for DeepSeek, and 0.8 percent for Mistral. Among 29,000 multi-family papers, 87.2 percent included at least one open-weight family. That larger number means something different: inclusion in a comparison set does not show that the family was the primary research instrument.
Open Weight Is Not One Constituency
The study defines open weight narrowly: trained parameters are available for download or local execution. It does not infer that training data, training code, filtering procedures, documentation, or license freedoms are open. Release-specific coding takes priority, and generic references to mixed-access families are not automatically counted as open weight.
This matters because the aggregate category can hide ecosystem substitution. The reported growth did not spread evenly across downloadable models; it was concentrated in particular Chinese-developed families, especially Qwen. A rising open-weight share can therefore reflect capability, cost, language coverage, platform availability, or institutional embedding without proving that researchers have adopted open science as a norm. The category records access to parameters; it does not reveal the reason for selection.
Geography Is Association, Not Motive
For 48,129 single-family papers with the required model, date, author, field, and affiliation data, a date-adjusted logistic model with OpenAlex subfield random effects estimated that papers linked to Chinese institutions had 2.23 times the odds of selecting an open-weight family, with a reported 95-percent interval of 2.12 to 2.34. A complementary multinomial model estimated 2026 use of Chinese-developed open-weight families at 37.1 percent for papers with a Chinese institutional link and 9.2 percent for papers with no observed China link.
These are associations after specified adjustments, not a causal explanation. The institutional measure concerns the first or last author, while a separate name-origin proxy is explicitly not a measure of nationality, citizenship, ethnicity, identity, or location. The paper cannot separate price, performance, language, access, professional networks, or institutional support. Geography is part of the observed dependency pattern; it is not permission to assign motives to people.
The Model-Dependency Map
A scientific paper should preserve more than a family label. A model-dependency map would record the exact release and checkpoint, developer, weight-access status, license, inference provider or local serving stack, access date, quantization, fine-tuning or adapter, retrieval sources, prompt and decoding settings, research role, data-handling path, compute requirements, known failure checks, and replacement conditions. It should distinguish a primary instrument from an experimental object, comparison baseline, and manuscript-assistance tool.
The record needs three linked layers: the claim the model helped produce, the runnable configuration needed to examine that claim, and the industrial ecosystem on which the configuration depends. Open weights may improve preservation or local execution while leaving other dependencies intact. A proprietary endpoint may remain repeatable for a time while giving its provider unilateral power to alter access. Neither label alone settles reproducibility, sovereignty, safety, or scientific adequacy.
The paper's map is aggregate and retrospective. The practical next step is local and prospective: make model choice visible at the moment a study commits to it. Then future reviewers can ask whether a result survived a version change, whether a model family was selected or merely benchmarked, whether an access regime constrained participation, and which dependencies moved when the scientific instrument changed.
What the Study Does Not Establish
The paper says corpus coverage and authorship proxies limit the precision of its analysis. The 2026 data stop in June, published papers lag model use, family aggregation conceals release-level variation, and model-use classification can be wrong. The institutional regression further excludes papers without the required linkage and country information; the paper reports residual missingness concentrated in recent records.
Denominators must therefore stay attached to every percentage. The 44.0-percent figure describes all coded single-family papers in the 2026 partial-year corpus; a 35.8-percent figure elsewhere in the paper describes the narrower institutional-analysis subset. The study supports a careful account of recorded model use and geographic association. It does not measure all science, establish researchers' motives, prove a causal national effect, or show that downloadable weights are equivalent to open science.
Source Discipline
The factual record was checked against the arXiv abstract, complete version 1 HTML and PDF, submitted TeX source, tables, and technical appendix. No study text was quoted. Percentages retain their populations and the model-dependency map is this essay's governance proposal.
Related Pages
- The Open-Weight Model Becomes the Release Boundary
- The Chokepoint Becomes the Open Model
- The Open-Weight Model Becomes the Governance Horizon
- Open-Weight AI Models
- The Governance Question Becomes the Geographic Bias Test
Sources
- Zackary Okun Dunivin, Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science, arXiv:2608.11090v1 [cs.CY], submitted August 11, 2026.
- Dunivin, version 1 HTML, reviewed for corpus construction, classification, descriptive results, statistical models, metadata limitations, and discussion.
- Dunivin, version 1 PDF, reviewed in full for the 21-page paper's affiliation, tables, figures, methods, results, and technical appendix.
- Dunivin, version 1 submitted source archive, checked for the manuscript text, exact table values, model specifications, and sample-flow records.