Blog · Review Essay · Modified July 10, 2026 · Last reviewed July 10, 2026

Data Grab and the Extraction Layer of AI

Ulises A. Mejias and Nick Couldry's Data Grab argues that Big Tech's power begins before prediction, personalization, or automation. It begins with the routine capture of human life as data. Read after the spread of generative AI, the book is less a general privacy complaint than a theory of the extraction layer beneath machine intelligence.

For this review, the extraction layer is the chain that turns life into model-ready leverage: collection, permission, retention, brokerage, labeling, embedding, scoring, memory, and reuse. The governance question is not only whether the final AI system behaves well. It is whether the upstream taking was legitimate, contestable, proportionate, and reversible.

The Book

Data Grab: The New Colonialism of Big Tech and How to Fight Back was published by the University of Chicago Press in 2024. The publisher lists Ulises A. Mejias and Nick Couldry as authors, gives the print ISBN as 9780226832302, the ebook ISBN as 9780226832319, and lists the book at 224 pages. Amazon lists the first edition with ISBN-10 0226832309, ISBN-13 978-0226832302, University of Chicago Press as publisher, and March 14, 2024 as the publication date.

The book extends the authors' earlier argument in The Costs of Connection. That earlier work named the data relation: the social arrangement in which ordinary life is converted into a continuous source of extractable value. Data Grab is more directly polemical. It asks what resistance would look like once data extraction is understood not as a bad bargain between user and service, but as a structural claim made by firms over social life.

Current Context

As of July 10, 2026, the book reads against a more concrete regulatory record than it had at publication. The FTC's September 2024 staff report on social media and video streaming services described "vast surveillance" with weak privacy controls and inadequate safeguards for children and teens. The FTC's January 2025 final order against Gravy Analytics and Venntel restricted sensitive location-data uses. California's DROP system is live for residents, with data brokers required to begin processing deletion requests on August 1, 2026. The EU Data Act has applied since September 12, 2025, and treats access, use, sharing, switching, and interoperability as governance questions rather than mere product features.

AI-specific governance now points in the same direction. NIST's AI Risk Management Framework and Generative AI Profile frame trustworthy AI as lifecycle work across design, development, use, and evaluation. The European Commission's AI Act overview says the Act entered into force on August 1, 2024, with prohibited-practice rules and AI-literacy obligations applying from February 2, 2025, GPAI governance obligations from August 2, 2025, and further high-risk obligations phasing in under the implementation timeline. Those instruments do not abolish extraction. They make extraction inspectable: provenance, minimization, retention, deletion, logging, contestability, and vendor exit become safety evidence.

Extraction Before Intelligence

The AI relevance is immediate. Public debate often begins at the model layer: benchmark scores, hallucination, alignment, safety testing, synthetic media, agentic workflows, and copyright disputes. Mejias and Couldry push the analysis downward. Before a model can predict, summarize, classify, rank, personalize, recommend, or automate, something has to become data. Text, clicks, location traces, images, purchases, messages, biometrics, worker activity, classroom interactions, health records, and platform behavior become inputs for systems that are later sold back as intelligence.

The path from life to leverage has several stages. First, instrumentation makes behavior recordable through apps, sensors, accounts, cookies, workplace suites, learning platforms, devices, and public-service portals. Second, permission systems make the capture appear normal through terms, contracts, procurement clauses, consent banners, API access, and background data sharing. Third, transformation turns records into labels, embeddings, segments, scores, training examples, saved memories, and model features. Fourth, action returns those derivatives as ranking, pricing, management, fraud detection, targeting, recommendation, eligibility, or automated assistance. A data grab is not only the first collection event. It is the whole route by which a trace becomes institutional power.

This is why the book belongs beside Atlas of AI, The Age of Surveillance Capitalism, and Ghost Work. It refuses the clean diagram in which AI arrives as a cloud service floating above society. The service depends on histories of capture, labeling, cleaning, brokerage, moderation, profiling, infrastructure, and asymmetrical consent. A chatbot interface may look like conversation, but its institutional precondition is a much older question: who had the power to collect and reuse the traces of human activity?

The Weight of the Analogy

The book's central analogy is deliberately heavy. "Colonialism" is not a decoration to make privacy sound dramatic. The authors use it to mark appropriation, enclosure, unequal exchange, dependency, and the normalization of extraction by powerful institutions. That framing is useful because it breaks the consumer myth. A person does not face Big Tech as a sovereign shopper comparing neutral offers. They often face platforms as conditions of work, speech, learning, mobility, entertainment, government access, and social recognition.

The analogy also needs discipline. Historical colonialism involved land seizure, racial hierarchy, slavery, military violence, imposed law, and dispossession in forms that should not be flattened into a metaphor for data collection. Data Grab is strongest when it treats colonialism as a claim about continuity and mutation, not sameness. It asks readers to see digital extraction as part of a longer political economy of taking, categorizing, governing, and profiting from lives that did not meaningfully consent.

The Governance Reading

The current regulatory record makes the book harder to dismiss, but also more demanding. It is no longer enough to say that platforms collect too much data or that AI systems need ethical principles. A serious review has to produce evidence: source category, collection context, legal basis or consent claim, fields collected, sensitive attributes, transformations, retention period, derivative artifacts, training or retrieval permission, recipients, deletion path, appeal path, and audit owner.

For AI procurement, that evidence should become an extraction ledger. The ledger should follow data from original capture into brokered files, feature stores, vector indexes, training corpora, evaluation sets, prompt logs, saved memories, agent traces, and downstream reports. If a deletion request, opt-out, contract termination, or source-quality failure cannot propagate to derivatives, the system has not solved the extraction problem; it has hidden it behind technical form.

For public institutions, the threshold should be higher. A school, benefits agency, health system, library, court, workplace, or city office should not make access to essential services depend on unnecessary capture. Nor should it buy privately extracted data to avoid the legal friction that would apply if the same information were collected directly. Procurement should require data minimization, purpose limits, subprocessor lists, retention schedules, deletion tests, incident reporting, audit logs, and an exit plan that preserves necessary public records without preserving avoidable surveillance.

Those controls do not prove that the problem is solved. They prove where the political fight has moved: from abstract concern to operational rights. Governance that starts only when a model is deployed arrives late. Data collection, retention, reuse, brokerage, training, and feedback loops are already governance decisions.

Where the Book Needs Care

The book's risk is over-consolidation. "Big Tech" is a necessary shorthand, but the extraction stack includes advertisers, brokers, cloud providers, app developers, device makers, public agencies, schools, employers, consultants, and contractors. Some are dominant platforms; others are small systems plugged into larger infrastructures. A good politics of data has to know which actor can be constrained by which lever: procurement, labor law, privacy law, antitrust, data protection, sectoral regulation, union bargaining, public-interest technology, or refusal.

The book's resistance program is morally clear, but it sometimes needs more institutional engineering. Collective resistance cannot depend only on awareness. People need rights to inspect, challenge, delete, port, withhold, negotiate, and audit data practices. Workers need protection when refusing surveillance. Communities need funding and technical capacity to build alternative data institutions. Public agencies need rules that stop them from laundering private extraction into public administration.

The opposite risk is romantic refusal. Not every record is a theft. Science, medicine, accessibility, public administration, safety, labor enforcement, journalism, and historical memory can require data. The line is not data versus no data. The line is whether collection is necessary, proportionate, purpose-limited, inspectable, time-bounded, contestable, and governed by people other than the party profiting from capture.

What This Changes

Data Grab sharpens a recurring problem across this library: the machine's apparent intelligence can hide the social arrangements that made it possible. Once extraction is normalized, the later system looks less like taking and more like service. The recommendation feels helpful. The score feels objective. The generated summary feels efficient. The agent completing a task feels inevitable.

The book's useful lesson is not that all data should disappear. The lesson is that records need politics. Ask who collected the data, who had a realistic choice, who can reuse it, who profits, who is exposed, who can refuse, what derivatives survive deletion, and what collective power exists when the answer is abusive. AI governance that skips those questions is not governance of intelligence. It is permission for extraction to keep calling itself innovation.

Source Discipline

This review separates book metadata, theory, regulator findings, legal timelines, and operational recommendations. University of Chicago Press and Penguin support bibliographic and publisher-framing claims. Couldry and Mejias's Internet Policy Review work supports the concepts of datafication, data colonialism, and data extraction as a social order. FTC, CPPA, European Commission, EUR-Lex, and NIST sources support current governance context. None of those sources proves that every AI system has the same data relation or the same remedy.

Claims about extraction should name the mechanism. A location-data broker, workplace productivity dashboard, social-video platform, classroom analytics tool, ad auction, vector database, model-training corpus, saved agent memory, or public-agency procurement contract raises different facts. The strong claim is not that all data use is colonial by definition. The strong claim is that the reader should be able to trace the route from capture to profit, classification, management, denial, or persuasion.

This page makes no claim that any AI system is conscious, divine, or AGI. It treats AI systems as institutional arrangements built from data, labor, infrastructure, interfaces, contracts, and governance choices.

Sources

Book links are paid affiliate links. As an Amazon Associate I earn from qualifying purchases.


Return to Blog · Return to Books