Gmail launched in April 2004 carrying a Beta tag, and Google kept that label for five years, through the point where hundreds of millions of people had made it their primary email account and businesses had built processes around it running correctly every day. The label outlasted any reasonable definition of experimental; it just kept being useful to Google as a hedge.
Staying in perpetual preview gives a company real advantages: lower liability if something breaks, room to change core behavior without the version-control discipline a finished product requires, and cover for rough edges that would otherwise draw harder scrutiny. None of those advantages disappear once users start depending on the product for real work — only the honesty of the label does.
Google's AI Overviews followed a near-identical arc two decades later: it launched through the opt-in Search Labs experimental program, then rolled out as the default result for a large share of search queries, including during a period in May 2024 when it generated widely reported, verifiably wrong answers — recommending, in one circulated example, adding glue to pizza sauce to help cheese stick. The feature had graduated from experiment to default while still producing that category of error in production.
The distinction that actually matters isn't whether a label says Beta. It's whether the company communicates breaking changes in advance, maintains a visible changelog, and treats a feature's reliability as something owed to the people depending on it rather than a fact to be disclaimed by badge alone.
Gmail eventually dropped its label in 2009, quietly, long after the label had stopped describing anything true about the risk of relying on it. AI Overviews has not dropped its own caveats, and the honest question for any product wearing a preview badge years into real dependency is not when the label comes off, but whether the company is willing to take on the reliability commitment the label was always supposed to be standing in for.
