A Recommender Learns from What It Lets People See
A click records a choice within an arranged encounter. To learn whether recommendations help people discover something worthwhile, evaluation must preserve the difference between what they rejected, what they overlooked, and what they never had a chance to choose.
The missing track
Consider a hypothetical music app that places familiar artists on a listener's home screen. The listener chooses one. Tomorrow, the app offers more of the same, now with another play supporting its prediction. An unfamiliar artist remains absent from both the screen and the listening history. Nothing in this sequence tells us whether the listener would have welcomed that artist.
The familiar recommendation may have been excellent. The problem begins when its success is used to settle a different question: whether the excluded alternative was unwanted. A system can correctly predict behavior within the opportunities it provides while remaining poorly informed about opportunities it withholds.
This essay's argument is that discovery requires an evaluation design that can encounter disagreement with the existing recommender. Improving prediction of yesterday's arranged choices is useful, but it cannot by itself establish that tomorrow's arrangement serves the listener better.
The conditions of a choice
For diagnosis, distinguish several questions. Was the item eligible for recommendation? Was it displayed? Was it noticed? Was it chosen? Was the choice worthwhile afterward? These are different stages of an encounter. A database entry saying “no click” cannot, without further information, tell us at which stage the encounter ended.
A controlled search study by Joachims and colleagues examined eye movements and manipulated result order. It found that clicks reflected relevance while also depending on presentation order and the other results available. Relative preferences inferred from clicks were more defensible than treating each click as an absolute relevance judgment. This was a student study of web search, published in 2005; it does not quantify position effects in today's music interfaces. The full paper supports caution about interpretation, not dismissal of behavioral evidence.
Popularity needs a separate question. Is an artist frequently chosen when offered, frequently offered because already popular, or both? Counting plays alone does not separate these explanations. Nor should unfamiliarity automatically count as merit. A discovery system owes the listener worthwhile possibilities, not obscurity for its own sake.
Interface and preference also need different remedies. In our hypothetical app, moving an artist into a more visible position changes the opportunity to respond. Replacing an unappealing suggestion changes its content. If a trial changes both at once, its result describes the combined product change. Calling that result improved understanding of taste would require additional evidence.
What the feedback-loop evidence establishes
Chaney, Stewart, and Engelhardt's 2018 research simulated recommendation systems repeatedly learning from behavior influenced by earlier recommendations. Under their model, this feedback could make users' consumption more alike without improving utility. Their simulated general preferences were fixed, an important distinction between convergence in behavior and a change of underlying taste. The work demonstrates a mechanism under stated assumptions; it is not a field experiment measuring homogenization on a named platform. Read the simulation study.
Our interpretation is that a feedback loop creates an identification problem before it establishes a cultural verdict. Similar listening histories might express shared enthusiasm. They might also reflect shared opportunities. Those explanations can coexist. To distinguish them, an evaluator needs variation in the opportunities, not just a more elaborate description of the histories.
This also limits the strongest criticism of personalization. The existence of feedback does not prove that every cycle makes recommendations worse. Listeners can benefit from familiar music, and a narrowly focused session can be exactly what someone requested. A useful evaluation asks whether the system remains responsive when the listener wants a different kind of encounter.
A comparison needs an opportunity
Schnabel and colleagues' 2016 paper develops selection-aware evaluation and learning using observation probabilities, often called propensities. Inverse-probability weighting gives more influence to observations that were less likely to enter the data. Its guarantees require the relevant probabilities and nonzero observation chances; very small probabilities increase variability. In observational settings, estimating those chances introduces another modeling problem. Weighting therefore cannot conjure evidence about items that had no chance of being observed. The paper specifies the assumptions.
Li and colleagues offer a complementary example: replay evaluation using randomized news-recommendation logs. Their revised paper reports close agreement between offline and online click measures in a Yahoo! setting. The formal result assumes independently drawn events from a common distribution and uniformly randomized logging. The authors discuss practical departures, including changing article pools and repeated visits. This is evidence that deliberately collected logs can enable useful comparisons, not that ordinary logs automatically support them. See the method and validation.
For the music example, the practical question becomes: which alternative recommendations can the available evidence actually compare? A broad catalog does not guarantee broad evidence. An honest report should distinguish supported comparisons from attractive possibilities that remain untested. Uncertainty is especially informative here: it identifies where the system's apparent confidence may simply reflect its past choices.
A trial that can learn about discovery
The following is our proposed evaluation design, not a protocol validated by those papers. Begin with a bounded question: among listeners requesting discovery, does a revised recommendation policy produce more encounters they later consider worthwhile? Define eligibility, the comparison policy, follow-up period, and success criteria before examining the result.
Keep the discovery control identical across comparison groups. Randomly assign participating listeners to the existing or revised policy, while preserving their ability to leave discovery mode. Comparing volunteers for discovery with everyone who prefers familiar music would confuse the policy's effect with the reason people selected a mode. Assignment should follow the question being tested, and the conclusions should remain limited to the participating population.
Within an appropriate eligible catalog, reserve a bounded opportunity to try alternatives and record their selection probabilities. Preserve enough information to reconstruct the available choices and presentation positions. Distinguish an item being served from evidence that it was visible; neither establishes that it received attention. Avoid treating the absence of a play as a complete account of the listener's judgment.
Assess immediate choices alongside later voluntary returns and optional reports of satisfaction. These measures should remain separate. A new artist might prompt curiosity without becoming welcome listening; an initially surprising track might become a favorite later. Optional responses introduce their own missingness, so report response rates and avoid silently treating respondents as everyone.
Also inspect how benefits are distributed. An average improvement could coexist with worse experiences for listeners whose interests rarely appear in the catalog. Report uncertainty for smaller groups instead of turning sparse observations into confident profiles. Catalog coverage describes available opportunity; satisfaction describes an experience. Neither metric should stand in for the other.
Finally, specify whether the trial evaluates a fixed model or a system that keeps learning. If both groups feed one shared training process, their experiences can influence the same future recommendations. An isolated or explicitly tracked update process makes the comparison easier to interpret. A short trial of fixed rankings cannot establish how repeated retraining will shape future discovery.
Leave preference open to revision
The user-facing implication is modest but consequential. Let listeners request familiarity, request discovery, and correct an inference without having to manufacture a new behavioral history. A person who says “less of this” has supplied information the system should not bury beneath accumulated plays.
For Church of Spiralism, the ethical commitment proposed here is to keep a person's recorded past open to correction by their present choices. Recommendation inevitably arranges opportunities. A defensible system can explain what it learned from those opportunities, identify what it still has not tested, and let the listener change the question.
Sources
Primary research consulted October 2, 2026. These historical studies establish methods and bounded findings, not current platform performance.
- Joachims et al., Accurately Interpreting Clickthrough Data as Implicit Feedback. SIGIR, 2005.
- Chaney, Stewart, and Engelhardt, How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility. RecSys, 2018; arXiv revision November 27, 2018.
- Schnabel et al., Recommendations as Treatments: Debiasing Learning and Evaluation. ICML, 2016.
- Li et al., Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms. WSDM, 2011; arXiv revision March 1, 2012.
Related reading
- Computing Taste and the People Inside Recommendations
- The Recommender Agent Becomes the Verification Cascade
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.