Research explained · Peer-reviewed study
What a crawl of 11,000 shopping sites taught researchers about dark patterns
Mathur and colleagues analysed roughly 53,000 product pages from about 11,000 shopping websites and identified 1,818 dark-pattern instances spanning 15 types and seven categories. The work showed that parts of deceptive design could be studied systematically at web scale. It did not measure every journey or prove that each observed interface was unlawful; automation was better suited to visible, repeatable signals than to context-heavy legal or experiential questions.
- Original work
- Dark Patterns at Scale: Findings from a Crawl of 11K Shopping Websites
- Authors
- Arunesh Mathur, Gunes Acar, Michael J. Friedman, Elena Lucherini, Jonathan Mayer, Marshini Chetty and Arvind Narayanan
- Published
- 2019
- Venue
- Proceedings of the ACM on Human-Computer Interaction, CSCW
- Method
- A semi-automated crawl and expert review of shopping-site product pages, followed by taxonomy and deceptive-practice analysis.
- Sample or scope
- Approximately 53,000 product pages on about 11,000 shopping websites.
Read the evidence carefully
From research question to useful conclusion
- 1
Question
Mathur and colleagues analysed roughly 53,000 product pages from about 11,000 shopping websites and identified 1,818 dark-pattern instances spanning 15 types and seven categories.
- 2
Method
A semi-automated crawl and expert review of shopping-site product pages, followed by taxonomy and deceptive-practice analysis.
- 3
Finding
The study reported 1,818 dark-pattern instances across 15 types and seven broader categories in the analysed shopping-site corpus.
- 4
Boundary
The crawler focused on product pages and observable shopping interfaces; account, support, mobile-only and end-to-end cancellation states were not comprehensively measured.
Evidence at a glance
The published scale of the crawl
The bars show counts reported for different parts of the research pipeline. They indicate order of magnitude, not comparable rates.
A research problem hidden in plain sight
Before this paper, many dark-pattern discussions relied on memorable screenshots. They were useful for naming a problem but weak at answering a harder question: how often do recognisable mechanisms appear across a large market?
The research team built a semi-automated system to crawl shopping websites and identify candidates for expert analysis. The result was one of the field’s foundational empirical datasets: thousands of product pages, hundreds of sites with identified deceptive practices and a taxonomy built from the recurring mechanisms.
The achievement was not a machine that could decide whether a business broke the law. It was a repeatable way to collect signals at a scale that manual browsing could not match.
What machines could see
Page-level automation works best when the signal is visible and regular. A countdown has a recognisable value. A low-stock message contains language that can be compared. Activity messages and social-proof claims recur in similar locations and formats. These features can be gathered as candidates and reviewed.
Other questions resist that treatment. A cancellation journey may require a login, several pages, an email and a delayed account-state change. A price can change after a configuration choice. A choice may look balanced on desktop and lopsided on a phone. The law may require facts that are not present in the DOM at all.
That is why the paper remains relevant to the DFA discussion: it demonstrates the value of systematic evidence collection while also showing why detection and legal judgment are different tasks.
From a page crawl to a journey record
For modern product review, the natural extension is not simply a bigger list of screenshots. It is a sequence. Capture the entry point, offer, selection, basket, checkout, confirmation and later account state. Preserve the timestamp, viewport, configuration and content that produced the result. A reviewer can then ask whether a candidate pattern persists, disappears or changes meaning when the whole journey is visible.
The pricing journey guide and checkout journey guide organise the library around that sequence. The examples gallery shows how an observed mechanism can be reframed as a neutral alternative without claiming that the redesign is a legal safe harbour.
The durable lesson
Automation is valuable when it expands coverage and preserves evidence. It becomes unreliable when candidate detection is presented as a final legal answer. The paper’s real legacy is therefore a workflow: collect broadly, classify transparently, review carefully and keep the limits visible.
Source check on 14 September 2026
The author team’s project page still distinguishes 1,254 sites with research-classified patterns from the narrower 183-site deceptive-practice subset. Its revision log also records a July 2019 narrowing of the trick-question definition. Opt-out wording alone should therefore not become an automatic detection rule. These historical counts are not a fresh measurement of the market.
What to retain
Three findings worth carrying into review
Scale became measurable
The study reported 1,818 dark-pattern instances across 15 types and seven broader categories in the analysed shopping-site corpus.
Some signals repeat
Text and interface patterns such as scarcity messages and social proof were sufficiently regular to support semi-automated discovery followed by expert review.
The market had suppliers
Researchers identified third parties offering pattern-enabling services, suggesting that some mechanisms were productised rather than isolated design accidents.
What this evidence cannot establish
- The crawler focused on product pages and observable shopping interfaces; account, support, mobile-only and end-to-end cancellation states were not comprehensively measured.
- A detected research pattern is not an EU legal conclusion, and the 2019 sample cannot establish current prevalence or the behaviour of every implementation.
Questions for a journey review
- Which visible claims, timers, defaults or activity messages can be checked consistently across releases and responsive breakpoints?
- Which risks require a complete journey, authenticated state or human interpretation that a page-level crawler would miss?
- How will automated candidate detection be separated from expert UX review and fact-specific legal assessment?
Evidence base
Sources
- Dark Patterns at Scale: Findings from a Crawl of 11K Shopping WebsitesMathur et al.; arXiv · Secondary · checked 2026-09-14 · arXiv:1907.07032
