Blog · arXiv Analysis · Published: August 12, 2026 · Modified: August 12, 2026 · Last reviewed: August 12, 2026

The Proof Requirement Becomes the Second Task

A bounty price states what a requester offers. It does not state what a worker must reveal, record, repeat, or do in public to prove completion.

A new audit names that missing labor term while carefully stopping short of claiming to measure worker experience.

The Paper

The source is Iman YeckehZaare’s Measuring Proof Burden in Public Bounty Listings: A RentAHuman Case Study, arXiv:2608.18547v1 [cs.HC], cross-listed in cs.CY and submitted August 19, 2026. The ten-page version-one paper is licensed CC BY 4.0. It studies advertised requirements in a May 31 snapshot of public bounty listings, not submissions, payments, rejections, or worker experiences.

Completion Can Contain a Second Task

RentAHuman’s current bounty documentation tells requesters to set a price and define completion criteria, including expected text, photos, video, or links. Its agent interface permits API or MCP clients to create bounties, review applications, communicate, and update task status. Those are platform capabilities; they do not establish that a particular listing was independently designed or supervised by an AI system.

The paper calls completion-linked exposure requirements proof burden. A requested photograph can be both evidence and new work. An account post can prove completion while attaching the worker’s account and audience. A later check can extend availability beyond the apparent task. Yet the study cannot always separate doing the task, gaining access, supplying equipment, and proving completion. The second-task frame is therefore a governance question, not a measured increment of labor time.

What the Audit Actually Observed

The nonrandom snapshot contains every listing returned by the researchers’ searches: 981 records, 980 from RentAHuman and one from Human Pages. A planned content screen removed promotional posts, service offers, and records requesting no human action, leaving 779 eligible bounty or task listings.

Two independent coders labeled all 981 records. A blinded third coder set the final labels for 873 broadly routed records; exact coder consensus became final for the other 108. The instrument records thirteen yes-or-no features: eleven evidence types—text, link, screenshot, photo, video, identity, account, location, phone, financial, and public-post proof—plus recurring monitoring and physical-world action. Automated rules were diagnostic only; all reported counts use human-reviewed labels.

The Checklist Beats the Score

Among the 779 eligible listings, the paper reports text proof in 57.8 percent, physical-world action in 48.5 percent, photo proof in 37.1 percent, and identity proof in 30.9 percent. Its author-designed zero-to-five score places 438 listings, 56.2 percent, at four or five. Those high-scoring listings contain 154 distinct feature combinations.

The word “severe” names a threshold on that research scale; it is not a finding of harm, illegality, or worker rejection. The author chose the tiers, modifiers, cap, and thresholds. The paper’s own intended-use boundary says neither the checklist nor the score is ready to rank tasks, set pay, moderate listings, or guide workers. The checklist is more informative because it preserves which exposure is requested.

An Agent Label Is Not an Agent Finding

The exploratory comparison includes 72 agent-or-bot-labeled listings and 676 human-labeled listings, excluding 31 labeled other. At least one of physical-world action, location proof, or recurring monitoring appeared in 75.0 percent of the first group and 55.3 percent of the second. But the three-feature comparison was selected after the researchers saw the data.

Only twenty displayed names supplied the agent-or-bot-labeled records, requester labels were self-reported or platform-assigned, repeated names may not identify one actor, and category mix affects the contrast. The paper finds no clear difference in the share scoring four or five and explicitly presents the three-feature pattern as a hypothesis for a preplanned collection, not evidence of autonomous-agent behavior.

Price Is Not the Whole Work Order

A fixed price or hourly rate is legible before acceptance, but it cannot summarize identity disclosure, travel, account use, purchases, public visibility, repeat checks, evidence retention, or rejection rules. The audit does not show whether higher-burden listings paid less or whether workers accepted them: visible application counts were confounded by listing age, and the data contain no completed-work or payment outcomes. The defensible policy move is disclosure before acceptance, not a wage claim the study cannot support.

The Evidence Boundary

The limitations are unusually consequential. The snapshot is neither random nor exhaustive; search failures may have missed listings; final labels depend heavily on one adjudicator; exact task wording is redacted; and no worker validated the construct. Consistent coding does not prove that workers experience two listings with the same score in the same way.

The version-one source package contains the manuscript, result macros, tables, and figures, but not listing-level records, coder packets, analysis scripts, or the supplementary PDF the paper says accompanies it. The data statement explains that raw text and identifiers remain private to limit disclosure risk. I could verify reported numbers against the manuscript and generated tables, but not independently reproduce the collection, labels, or statistical analyses.

The Proof-Burden Receipt

A proof-burden receipt should disclose the price or rate; expected time; task versus evidence steps; required text, links, images, video, accounts, identity, location, phone, payments, or public posts; physical travel and expenses; recurring checks; evidence audience and retention; verifier; rejection standard; appeal; payment-release condition; requester-label provenance; responsible human operator; and any change after acceptance. This is the essay’s proposal, not a validated instrument from the paper.

The Spiralist lesson is modest: verification is part of the work order. If a system can specify what proves completion, it can specify what that proof asks a person to surrender.

Sources


Return to Blog