Delivery Disruption Cuts Trade but Concentrates Sellers
A 20% chance of failed delivery reduced efficiency by 70% while ratings noise left the experimental market largely unchanged.
Underlying Paper
How to Disrupt a Market
Market design research in economics naturally focusses on how to improve market efficiency. Our objective here is exactly the opposite - how to design interventions that make a market less efficient. Our research is inspired by the growth of illicit markets online where reducing their efficiency may reduce societal harm. Using a web-based experiment, we find that a partial disruption to delivery is an effective method to decrease market efficiency. The decrease is borne by sellers who sell fewer goods and have lower earnings. A consequence of a disruption to delivery, however, is an increase in market concentration because it facilitates the emergence of a dominant seller. In contrast, we find that attacks on seller ratings are ineffective at reducing market efficiency. This study paves the way for evidence-based, causally driven investigations to aid policies to disrupt cybercrime and other illicit markets.
Markets with anonymous sellers depend on reputation and delivery reliability to make trade possible under asymmetric information. That matters for cybercrime and other illicit markets because enforcement sometimes aims not to improve allocation, but to make harmful trade less efficient. This paper tests that inverse market-design problem in a controlled online marketplace: if a platform or authority wants to disrupt trade, should it corrupt seller ratings or interfere with delivery?
The answer in the experiment is asymmetric. Random delivery failures sharply reduce realized surplus and seller earnings. Random rating errors do little. The same delivery intervention, however, also makes the market more concentrated by helping a dominant seller emerge.
Core Contribution
The paper’s contribution is not another model of reputation, but a causal comparison between two operationally different disruptions. A rating attack changes the public signal buyers use to judge sellers. A delivery attack changes the transaction outcome itself and may spill into reputation because buyers do not know whether non-delivery was caused by the seller or by the intervention.
That distinction is the key mechanism. If buyers can rely on repeated personal trading relationships, corrupting public ratings may not move behavior very much. Failed delivery is harder to route around: it directly removes some successful trades and indirectly makes buyers less willing to purchase.
Technical Approach
The authors ran a web-based experiment on oTree with Amazon Mechanical Turk participants. The final sample contains 392 participants, organized into 56 complete groups: 14 groups in each of four treatment conditions. Each group formed a market with three sellers and four buyers for 20 trading periods, with roles fixed throughout the session.
Sellers chose what to produce, what to advertise, and what price to set. Goods could be Regular or Super; buyers valued Regular goods at 30 points and Super goods at 150 points. Sellers could advertise a different quality or quantity than they produced, so buyers faced both quality uncertainty and seller-selection uncertainty. Buyers saw advertised quality, price, seller identifier, average rating, and the seller’s last three ratings before choosing whether to buy. Figure 1 summarizes the experimental sequence and the two disruptions.
The design is a 2 × 2 between-subjects experiment: Baseline, Rating, Delivery, and Combined. In the Rating condition, each submitted rating has a 20% probability of being replaced by a different random rating. In the Delivery condition, each purchased good has a 20% probability of not being delivered, regardless of what the seller produced. The Combined condition applies both disruptions independently.
For the main tests, the authors focus on the final 10 market rounds because several outcomes exhibit time trends. Treatment comparisons use group-level Mann–Whitney tests, with Wilcoxon signed-rank tests for trends, and the appendix checks the conclusions using random-effects models.
Results and Analysis
The delivery intervention is the result that carries the paper. In the final 10 rounds, market efficiency is 70% lower than Baseline in the Delivery treatment and 76% lower in the Combined treatment. The summary table reports market efficiency of 29.14 in Baseline, 34.92 in Rating, 8.673 in Delivery, and 6.939 in Combined for rounds 11–20. Rating noise alone does not reduce efficiency.
The loss falls mainly on sellers. Seller earnings in rounds 11–20 drop from 49.61 in Baseline to 18.40 in Delivery and 21.10 in Combined, corresponding to reductions of 63% and 57%. Buyer earnings are reported as unchanged across treatments in the main analysis. That split is useful for policy interpretation: delivery disruption makes the market worse by reducing seller revenue and trade volume, not by imposing a measured buyer-side loss in this experimental setting.
The indirect behavioral effects are also visible. Goods sold fall from 2.836 in Baseline to 2.329 in Delivery and 2.329 in Combined in rounds 11–20, about an 18% reduction. Buyer inactivity rises from 24.64% in Baseline to 40.00% in Delivery and 40.95% in Combined, a 62–66% increase. The appendix reports similar long-run effects after adjusting market efficiency for the direct mechanical effect of seized goods: adjusted efficiency remains 55% lower in Delivery and 64% lower in Combined.
The unintended effect is concentration. The dominant seller’s market share rises to 59% in Delivery and 56% in Combined, compared with 41% in Baseline. Figure 3 shows why: delivery failures are more damaging for smaller sellers because a seller with one sale has a 20% chance that all sales in the round are hit, while a seller with two sales has only a 4% chance that both are hit. The intervention weakens small sellers more often, making repeat trade gravitate toward larger sellers.
Caveats in Practice
The evidence is credible for the laboratory market, but the policy translation is narrower than the headline result might suggest. Participants traded in small, fixed groups, with only three sellers, four buyers, and 20 rounds. The delivery attack is modeled as random non-delivery, not as a real enforcement operation with strategic adaptation, retaliation, or migration to another platform.
The rating result is also specific to the tested attack. A 20% random replacement of numerical ratings may be too diffuse if buyers rely on personal trading partners. The authors themselves point to sharper alternatives: targeting top-rated sellers, degrading qualitative feedback, or testing seller-side attacks. The main lesson is therefore not that reputation attacks never work, but that this simple rating-noise intervention was weak compared with failed delivery in this market design.
Evidence Box
strongKey Claims
- •Delivery failures reduce market efficiency in anonymous online markets
- •Ratings noise is ineffective when buyers can rely on repeated trading partners
- •Delivery disruption can increase seller concentration
- •Personal trading relationships reduce reliance on public ratings
Key Results
- •392 participants across 56 complete groups, with 14 groups per treatment
- •Market efficiency 70% lower in Delivery and 76% lower in Combined than Baseline in the final 10 rounds
- •Seller earnings 63% lower in Delivery and 57% lower in Combined than Baseline
- •Dominant seller market share 59% in Delivery and 56% in Combined versus 41% in Baseline
Limitations & Caveats
- •Controlled online experiment with three sellers and four buyers per market
- •Delivery attack modeled as random 20% non-delivery rather than a field intervention
- •Rating attack limited to random replacement of numerical scores
- •Short 20-round setting may understate adaptation, migration, or retaliation in real illicit markets