Enterprise Telemetry Becomes the Work Census
A study of more than 17 million ChatGPT Enterprise messages turns product telemetry into a map of organizational AI use: who is active, how often, and for which classified tasks.
The map is census-like, but it is not a census of the workforce, work quality, productivity, or value. Governing it begins by keeping those boundaries visible.
The Paper
The source is Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, and Gawesha Weeratunga's How Organizations Use AI: Evidence from ChatGPT, arXiv:2608.12236v1 [econ.GN, with cs.AI and cs.HC], submitted August 12, 2026. The record calls it a working paper whose results may change. All five authors list an OpenAI affiliation; Holtz and Tambe also list academic affiliations and disclose that they contributed as paid OpenAI contractors.
That provenance does not erase the evidence, but it belongs beside it. The product owner can assemble administrative records outsiders cannot independently reproduce, while also controlling the study's measurements, classifications, and disclosure boundaries.
What Was Measured
The paper's data section separates four samples rather than treating one dataset as universal. Its broad organization-week panel covers centrally administered, paid ChatGPT Enterprise workspaces adopted from January 1, 2024 through March 31, 2026. It excludes personal accounts, API use, other subscription plans, other vendors, and internally built tools.
At 26 weeks after adoption, the worker-characteristics sample contains 1,764 organizations and 17,446,551 messages, with an industry assignment and at least some usable job-title data. A narrower task sample contains 973 organizations and 8,696,657 classified messages because its classifier was available only from October 30, 2025. These are large observational samples inside a bounded product population.
The Census Metaphor
The growth analysis reports that aggregate output tokens grew roughly sevenfold from June 2025 to March 2026, while output in a fixed cohort of earlier adopters grew about fourfold. Its role analysis finds active use across functions and seniority levels; among active users, inferred early-career workers and trainees sent roughly eight to nine more messages per week than the within-firm average. Its task analysis distributes messages across writing, technical, communication, research, planning, legal, financial, and other categories.
Placed together, organization, role, seniority, frequency, and task labels resemble a work census. Yet the rows do not enumerate jobs performed, outputs delivered, errors introduced, authority exercised, or value created. They enumerate interactions within one service. A measurement system decides what those traces mean.
The Missing Denominator
The paper's role limitations are explicit: job-title coverage is incomplete and recorded at one point in time, and the researchers do not observe each role's full workforce denominator. A category's share of weekly active users is therefore not the percentage of that category using ChatGPT, and it cannot show whether a role is overrepresented among users. Intensity comparisons also concern people who were active, not everyone's probability of becoming active.
This blocks seductive headlines. Seven percent of the average firm's observed active users being classified as early-career or trainees is not a seven-percent workforce adoption rate. Eight or nine extra messages is not eight or nine units of productivity. The paper itself says message volume measures usage intensity, not complete economic importance. A missing denominator is not a minor statistical caveat; it defines which labor claims are unavailable.
The Classifier Is Part of the Claim
The job-title appendix says gpt-5-mini classified a supplied title into department, seniority, manager status, and job-title class. Display categories then collapse several outputs: for example, student, trainee, intern, entry, and associate labels become “early-career / trainee.” The validation section publishes frequent example titles subject to a suppression rule, but it does not report a held-out accuracy, agreement rate, or confusion matrix.
The task appendix says each turn receives exactly one value in a two-level, 60-category taxonomy and that the classifier was evaluated on an internal human- and model-labelled benchmark. It publishes no benchmark score there. A multi-purpose or ambiguous turn must nevertheless land in one bin. Classification uncertainty is therefore inside every role-by-task result, even when a chart presents clean categories.
Privacy Is Not a Labor Charter
The methods say the researchers used de-identified data, reported only aggregates, classified message content automatically, securely linked aggregate organizational usage to financial data, and did not manually review individual customer messages. Those are meaningful study safeguards, and they should be stated rather than replaced with a vague accusation that researchers read everyone's chats.
They do not answer every workplace-governance question. The paper does not describe worker notice, consent or opt-out, retention, access rights, or whether employers may use similar labels in performance decisions. Aggregate research can reduce disclosure risk while still teaching institutions how to compare worker groups. A privacy-preserving study design is not, by itself, authorization for employee scoring or managerial surveillance.
What Employers Must Not Infer
The data do not establish that frequent users are more productive, less capable, more replaceable, or more deserving of promotion. They do not show that an “early-career” classification describes actual tenure or age. They do not measure downstream work products, productivity effects, changes in routines, wages, job loss, or whether model output was accepted. The paper's conclusion states these outcome limits directly.
The public-company comparison is also associational. Observed adopters are larger and more valuable, but the authors warn against causal interpretation; firms without a ChatGPT Enterprise match may use competing products, personal accounts, or other OpenAI services. “Uses more” cannot be silently upgraded to “performs better,” and “not observed here” cannot become “does not use AI.”
The Work-Measurement Receipt
A work-measurement receipt should record the product boundary, observation period, inclusion rule, adoption horizon, unit of analysis, active-user definition, missing-title coverage, available and missing denominators, classifier model, prompt, taxonomy, validation set and scores, aggregation and suppression rules, uncertainty, data-access purpose, retention, prohibited inferences, reviewer, and correction or contest path. Every dashboard derived from workplace AI telemetry should carry the receipt forward.
The receipt also separates research from employment action. A descriptive aggregate does not authorize individual performance review. Worker-level evaluation needs a necessity test, labor and privacy review, access controls, appeal, and evidence that the metric measures what the decision claims.
What the Study Establishes
The paper's reported scope establishes that ChatGPT Enterprise activity, among observed adopting organizations, is growing and distributed across many classified roles and tasks. It also demonstrates how vendor telemetry can make organizational AI use legible at a scale surveys rarely reach. It does not establish a representative census of work, a causal productivity effect, or a forecast of labor displacement.
The durable lesson is methodological and political: once workplace conversations become countable, the categories used to count them become part of workplace power. Keep the trace distinct from the task, the active user distinct from the workforce, and the message distinct from human worth.
Related Pages
- The Enterprise Role Matrix Becomes the AI-Native Work Map
- The Receptivity Index Becomes the Adoption Margin
- The Enterprise Connector Becomes the Permission Map
- Algorithmic Management
- Privacy and Data Stewardship
- Research and Editorial Integrity
Sources
- Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, and Gawesha Weeratunga, How Organizations Use AI: Evidence from ChatGPT, arXiv:2608.12236v1 [econ.GN, cs.AI, cs.HC], submitted August 12, 2026.
- Authors' version 1 HTML paper, reviewed in full, including sample construction, role and task results, conclusion, job-title classifier prompt and validation section, and task-classifier appendix.
- Authors' 69-page version 1 PDF, checked against the HTML record for pagination, figures, tables, footnotes, appendices, disclosures, and limitations.