A/B testing conventions
Parent: Headline craft · Published reference · snapshot 2026-09-24
↓ Facts as markdownall context files
Depth-first rabbithole dossier for A/B testing conventions; source-anchored research pack.
These notes link each claim to its source. A source may be a research report hosted on this site rather than the primary document. A published reference means the content is available; it does not certify independent review or accuracy.Read the editorial policy and follow the sources before relying on a claim.
Structure and components
- 30. Choose the metric first. Optimizing raw CTR rewards negativity (claim 26). A post-click engagement metric such as Quality Clicks (claim 10) is the documented guard against click-only winners. 31. If the goal is to pick a winner fast on short-lived content, use a bandit (claims 14–16). If the goal is to learn which headline features work, use fixed-horizon or always-valid tests, because bandit estimates are biased (claim 17) and peeking inflates errors (claims 18, 20). 32. Run an SRM check on every headline test. Caching layers in front of news pages are a documented source of broken random [source]
How it works
- IN: how the conventions of A/B testing (and headline A/B testing specifically) arose and changed — who ran tests, how many variants, what metric, how a winner was picked, when to stop — and the primary sources that record each step. OUT: sibling headline-craft concepts (clickbait taxonomy, headline formulas, SEO titles), the parent domain, and general experimentation-platform engineering. Statistical methods appear only where they changed a convention. [source]
- 28. Hagar & Diakopoulos (*Media and Communication* 7(1), published 2019-02-19) interviewed 10 practitioners in August–September 2018. https://www.cogitatiopress.com/mediaandcommunication/article/view/1801 29. They found tests typically used two to five headline options, written by a writer or editor. https://www.cogitatiopress.com/mediaandcommunication/article/download/1801/1017 30. They found tests often ran "until the test reaches statistical significance," usually on CTR. Some systems then served the winner automatically, and others left the choice to editors. https://www.cogitatiopress.com [source]
Measurements and reference values
- 18. A bandit treats each headline variant as an arm and aims to maximize total clicks during the test, not just after it. https://arxiv.org/pdf/1908.06256 19. Thompson Sampling gives each arm a Beta(1,1) prior, draws one sample per arm from its Beta posterior, and shows the arm with the largest sample. https://arxiv.org/pdf/1908.06256 20. After a click the arm's posterior becomes Beta(α+1, β); after a non-click it becomes Beta(α, β+1). https://arxiv.org/pdf/1908.06256 21. At high traffic, Yahoo updated posteriors in batches ("batched Thompson Sampling", bTS), not per event. https://arxiv.org/p [source]
- 1. Digital newsrooms may test up to about a dozen headlines per article and let an algorithm converge on the best one against a metric such as click-through rate (CTR). https://www.cogitatiopress.com/mediaandcommunication/article/view/1801 2. Hagar & Diakopoulos (2019) interviewed people in US newsrooms. They found that a group of "optimizers" runs the tests, reads the results, and advises editors. This group includes social media, audience-engagement, and analytics staff. https://par.nsf.gov/biblio/10096346-optimizing-content-headline-testing-changing-newsroom-practices 3. Upworthy ran a rand [source]
- 56. Kohavi et al. (2007) trace the Control/Treatment design back to an 18th-century scurvy trial run by "a British ship's captain". ⚠ [H] https://ai.stanford.edu/~ronnyk/2007GuideControlledExperiments.pdf 57. By 2007 the method went by several names: randomized experiments, A/B tests, split tests, Control/Treatment tests and parallel flights. [H] (same source) 58. The 2007 guide cites Hopkins's *Scientific Advertising* (1923), which links web split tests to mail-order advertising tests. [H] https://www.researchgate.net/publication/221654095_Practical_guide_to_controlled_experiments_on_the_web_ [source]
- **Multi-armed bandit (Thompson Sampling)** 29. A bandit treats each variant as an arm and tries to maximize clicks during the test, not only after it. [M] https://arxiv.org/pdf/1908.06256 30. Each arm starts with a Beta(1,1) prior. The bandit draws one sample per arm and shows the arm with the largest draw. [M] (same source) 31. A click updates the arm's posterior to Beta(α+1, β); a non-click updates it to Beta(α, β+1). [M] (same source) 32. At high traffic, Yahoo updated the posteriors in 5-minute batches. Updating more often gave only a marginal gain in clicks. [M] (same source) 33. Summing [source]
- 13. Mao et al. (Washington Post, 2019) found that testing-period impressions reach "as high as 24.36% of the total impressions across all articles." The opportunity cost of showing losing variants is therefore large for news. — https://ar5iv.arxiv.org/html/1908.06256 14. The same paper measured a 12% CTR discrepancy between the testing and post-testing periods. A winner picked early can underperform later. — https://ar5iv.arxiv.org/html/1908.06256 15. The authors acknowledge that bandits break down if the true CTRs change so fast that "the order among θₖ flips multiple times during the article [source]
- 18. Chartbeat states that "homepage audiences are usually loyal and visitors from social tend to be new, so it's likely they'll prefer different headlines." A winner from a front-page test is therefore not evidence for social or search placements. — https://chartbeat.com/resources/product/headline-testing-101/ 19. Hagar and Diakopoulos found that "no single feature of a headline's writing style makes much of a difference in forecasting success." They advise testing for the moment and context, not deriving general style rules. (Search-index excerpt; niemanlab.org returned 403 to direct fetch.) [source]
- 18. On 2010-04-18, Evan Miller showed that if you stop a test as soon as it looks significant, "all the reported significance levels become meaningless." https://www.evanmiller.org/how-not-to-run-an-ab-test.html 19. In Miller's worst case (checking after every observation, up to 150 observations), a nominal 5% test produced false positives 26.1% of the time. https://www.evanmiller.org/how-not-to-run-an-ab-test.html 20. Miller's fixes were a sample size fixed in advance, a sequential design with pre-set checkpoints, or a Bayesian design. https://www.evanmiller.org/how-not-to-run-an-ab-test.html [source]
- 24. On 2014-08-25, Facebook began demoting "click-baiting" headlines, which it defined as headlines that "encourage people to click to see more, without telling them much information." https://about.fb.com/news/2014/08/news-feed-fyi-click-baiting/ 25. Facebook's clickbait signals were time spent on the article and the ratio of clicks to likes, comments and shares. This made raw CTR a weaker winner criterion for headlines distributed on Facebook. https://about.fb.com/news/2014/08/news-feed-fyi-click-baiting/ 26. On 2016-02-09, Nieman Lab reported The Washington Post's "Bandito" tool. Bandito te [source]
- - **Does positive wording lower CTR?** Robertson et al. (2023) say yes: https://www.nature.com/articles/s41562-023-01538-4. Gligorić et al. (2023) found no significant positive-emotion effect (β = −0.04, p = 0.071): https://pmc.ncbi.nlm.nih.gov/articles/PMC10038272/. A 2025 reanalysis (Reiss & Roggenkamp) supports the negativity effect but not the positive-word penalty when positivity is measured semantically: https://www.journalofrobustnessreports.org/does-negativity-drive-online-news-consumption/ - **Can tests find general rules?** Practitioners say tests reveal stable best practices that ev [source]
- 37. At Yahoo, the fixed test window sent (K−1)/K of test traffic to inferior headlines, and test traffic was 24.36% of all impressions. https://arxiv.org/pdf/1908.06256 38. At Yahoo, CTR in the testing period and CTR after it differed by 12% on average. So the one-hour winner can be the wrong choice for the article's full life. https://arxiv.org/pdf/1908.06256 39. Most newsrooms could not test across platforms. The CMS often allowed only one headline everywhere, or one extra search or social headline. https://www.cogitatiopress.com/mediaandcommunication/article/download/1801/1017 40. No interv [source]
- 13. Yahoo Front Page's baseline convention was "test-rollout": split traffic equally for a set test period, then send all traffic to the winner. https://arxiv.org/abs/1908.06256 14. On Yahoo data, batched Thompson sampling beat test-rollout by 3.69% in clicks, because it shifts traffic to better headlines while the test is still running. https://arxiv.org/abs/1908.06256 15. Chartbeat's headline testing is a multi-armed bandit. It serves winning variants more often as evidence builds, then shows the winner to everyone. https://help.chartbeat.com/hc/en-us/articles/360023929453-Guide-to-Headline- [source]
Problems, failure modes and limitations
- **In scope:** how publishers A/B-test headlines. That covers: - the test unit and variants; - the metric; - how traffic is split; - how tests stop and how a winner is chosen; - the checks that keep a result valid; - how the conventions changed over time; - failure modes and disagreements. [source]
Comparisons and alternatives
- 9. By October 2009, The Huffington Post was running real-time headline A/B tests. Readers saw one of two headlines at random. After five minutes, the headline with more clicks was shown to everyone. https://thenoisychannel.com/2009/10/15/innovation-at-huffington-post-data-driven-headlines/ (summarizing Nieman Lab, 2009-10-14: https://www.niemanlab.org/2009/10/how-the-huffington-post-uses-real-time-testing-to-write-better-headlines/) 10. HuffPost's CTO described the practice at the Online News Association conference in 2009. https://www.mediapost.com/publications/article/115675/ 11. Conventions [source]
- IN: the conventions editors and platforms use to A/B test *headlines* (and headline+image packages). That covers the choice of metric, stopping rules, confidence thresholds, variant design, randomization integrity, how winners are interpreted, and the fixed-split vs bandit choice. OUT: headline writing technique itself, the parent domain (headline craft), general CRO/landing-page testing, and sibling concepts (e.g., SEO titles, social-card optimization). Where a general experimentation source is cited, it is cited only for a failure mode that headline tests demonstrably share. [source]
- 1. Kohavi, Henne & Sommerfield open their 2007 KDD guide with an 18th-century scurvy trial as the ancestor of the Control/Treatment design ("a British ship's captain" gave limes to half his crew). https://ai.stanford.edu/~ronnyk/2007GuideControlledExperiments.pdf 2. The same paper lists the synonyms in use by 2007: "randomized experiments (single-factor or factorial designs), A/B tests (and their generalizations), split tests, Control/Treatment tests, and parallel flights." https://ai.stanford.edu/~ronnyk/2007GuideControlledExperiments.pdf 3. The paper cites Claude Hopkins's *Scientific Advert [source]
- | # | Point | Side A | Side B | |---|---|---|---| | C1 | Who wrote arXiv 1908.06256 | [M][H][P]: Mao et al., Yahoo Front Page. [M] read the PDF. [P] says a search snippet wrongly credited it to the Washington Post. | [E] claims 13–16: "Mao et al. (Washington Post, 2019)", and it calls the baseline "standard A/B testing". | | C2 | Chartbeat's recommended number of variants (same URL) | [M] claim 7, [P] claim 9: "at least 4 or 5". | [E] claim 29: "three or four". | | C3 | How Upworthy's "significance" column was computed | [M] claim 17: "NOT an estimate of statistical significance", and former s [source]
- - **D1. Bandit or fixed/always-valid test.** - For bandits: batched Thompson Sampling gives +3.69% clicks, and short-lived news favours bandits. https://arxiv.org/pdf/1908.06256 - Kohavi agrees that bandits suit headlines, but says they need an equal or larger sample and give weaker causal inference. https://www.linkedin.com/pulse/multi-armed-bandits-thompson-sampling-ab-testing-you-headlines-ronny - Bandit means are biased (claim 53), and always-valid testing keeps error control. https://arxiv.org/abs/1512.04922 - What decides it: optimizing this one story versus learning something reusable. [source]
- - **Size of the headline-testing payoff.** Chartbeat, the vendor, reports "roughly 45% lift" in winning tests, holding steady over two years across more than 1M tests ( https://chartbeat.com/resources/research/headline-experimentation-time-value/ ). Against that: the selection bias above (claims 30–31), the 12% drop from the testing to the post-testing period (claim 14), and the 3.69% gain the Washington Post measured for its method over A/B testing (claim 16). The vendor figure does not say it corrects for winner's curse. - **Whether tests teach writing rules.** Registered-report evidence fin [source]
- - **How much testing helps.** Chartbeat reports large lifts (62% win rate, 71–78% lift; claim 25). The Yahoo bandit study reports a 3.69% gain over a test-rollout baseline (claim 14). Neither side reconciles the two. They measure different things: an alternative versus the original headline, and one allocation method versus another. - **Positive words.** Robertson et al. found that positive words lower CTR (claim 26). A later reanalysis (Reiss & Roggenkamp 2025) used semantic rather than lexical positivity and found that the negativity effect persists but the positive-word penalty disappears. [source]
Facts and statements
- Run: /rabbithole, 2026-09-24. Concept: A/B testing conventions as applied to headline craft. [source]
- 12. At Yahoo Front Page, the testing period was the first hour after publication. https://arxiv.org/pdf/1908.06256 13. During that hour, Yahoo assigned each view request to a headline variant with equal probability. https://arxiv.org/pdf/1908.06256 14. After the hour, Yahoo showed the variant with the highest CTR to all remaining traffic. https://arxiv.org/pdf/1908.06256 15. Newsroom tests often continue until one headline "can be confidently declared better" on the optimized metric. https://www.cogitatiopress.com/mediaandcommunication/article/download/1801/1017 16. Some systems switch all rea [source]
- These are separate conventions, and no single one dominates. [source]
- **Audience and platform** 77. Most newsroom CMSs allowed only one headline everywhere, or one extra search or social headline. No interviewee ran fully separate tests per platform, and only one ran native tests on Facebook. [M][H] https://www.cogitatiopress.com/mediaandcommunication/article/download/1801/1017 78. Homepage visitors tend to be loyal and social visitors tend to be new, so the traffic source changes which headline wins. [M][E][P] https://chartbeat.com/resources/product/headline-testing-101/ 79. Smaller sites reach significance more slowly, and some stopped testing for that reason. [source]
- **Passes most likely to add claims next**, collected from the four reports: 1. Primary Chartbeat methodology pages (these returned 403): https://help.chartbeat.com/hc/en-us/articles/360050302434-Thompson-Sampling-Methodology and the testing guide. These would resolve C2 and confirm claims 24–25 and 38. 2. The full text of Hagar & Diakopoulos (2019), and Hagar, Diakopoulos & DeWilde (2021). 3. Novelty and position effects inside the homepage headline module. 4. Sequential or always-valid inference inside commercial headline tools (this would resolve C4). 5. Bandito's allocation algorithm (C5), [source]
- - Headline linguistics: negativity bias, positive words, curiosity gap and concreteness (the concreteness study is https://pmc.ncbi.nlm.nih.gov/articles/PMC11704130/) - Clickbait detection and platform demotion; clickbait and trust - Multi-armed bandits for content optimization - Sequential testing and always-valid inference - General SRM and trustworthy-experimentation practice [source]
- - Clickbait detection and platform demotion - Multi-armed bandits for content testing - Sequential testing and always-valid inference - The headline-negativity bias [source]
- - https://www.cogitatiopress.com/mediaandcommunication/article/download/1801/1017 — Hagar & Diakopoulos, *Media and Communication* 7(1), 2019-02-19 - https://arxiv.org/pdf/1908.06256 — Mao et al. (Oath/Yahoo), "A Batched Multi-Armed Bandit Approach to News Headline Testing", arXiv v2 2019-08-25 - https://upworthy.natematias.com/about-the-archive.html — Upworthy Research Archive data documentation - https://upworthy.natematias.com/2024-06-upworthy-archive-update.html — archive randomization update, 2024-06 - https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0281682 — Gligorić et [source]
- **Out of scope:** general experimentation platforms, the wider headline-writing craft, clickbait as a genre, SEO titles, and email subject-line testing. Each of those is a separate frontier item. [source]
Related concepts
- testing — is a part of A/B testing conventions
- conventions — is a part of A/B testing conventions
Children
- No children recorded.