The Em-Dash Became AI Writing's Evolutionary Marker
The em-dash—that elongated punctuation mark long favored by professional editors—has quietly become one of the most reliable fingerprints of AI-generated text, and its sudden proliferation since late 2022 marks a genuine turning point in the natural history of written language. As detection tools, educators, and editors increasingly cite the punctuation mark as a telltale sign of output from ChatGPT, GPT-4, and Claude, linguists are only now reconstructing how a niche editorial preference exploded into an era-defining stylistic tic.
The story is not one of deliberate design. No engineer at OpenAI or Anthropic instructed a model to prefer dashes. Instead, the em-dash boom is a case study in what evolutionary biologists call punctuated equilibrium: long stability followed by a sudden, dramatic shift triggered by a change in environment. In this case, the environment was the training corpus.
A Fossil Hidden in the Training Data
For decades, professional copyeditors at magazines, publishing houses, and newspapers quietly favored the em-dash as a flexible tool. It could replace commas, parentheses, and colons while lending prose a brisk, confident rhythm. Style guides tolerated it; house styles often encouraged it.
That preference remained largely invisible to ordinary readers. Most casual writers rarely typed an em-dash at all, in part because it is awkward to produce on a standard keyboard without a shortcut or auto-formatting.
When the transformer era arrived, large language models ingested enormous quantities of professionally edited prose—books, journalism, and polished web content where the dash flourished. The models learned, statistically, that fluent and authoritative writing frequently deploys the mark.
The result is a linguistic fossil embedded in model behavior. The em-dash's abundance in AI output is a direct imprint of the human editors whose work dominated the high-quality slice of the training corpus.
Why 2022 Is the Dividing Line
The public launch of ChatGPT in November 2022 functions as a clean stratigraphic boundary. Before that moment, dense em-dash usage signaled a trained human editor. After it, the same pattern increasingly signaled a machine.
This inversion happened remarkably fast. Within two years, the em-dash shifted from a mark of editorial polish to a suspected marker of automation.
Several factors amplified the effect:
- Reinforcement learning from human feedback (RLHF) rewarded outputs that read as fluent and professional, reinforcing edited-prose patterns.
- The models default to a confident, essayistic register that leans heavily on the dash for pacing.
- Unlike humans, models produce the mark effortlessly and consistently, with no keyboard friction to suppress it.
The outcome is a frequency of em-dash usage that often exceeds what even seasoned human writers produce, creating a statistical anomaly detectable at scale.
The Detection Debate Intensifies
In 2024 and 2025, the em-dash moved from linguistic curiosity to practical concern. Educators grading essays, editors reviewing freelance submissions, and hiring managers screening cover letters increasingly flag heavy dash usage as a potential sign of undisclosed AI assistance.
This has sparked a backlash. Many professional writers who legitimately favor the em-dash now report being wrongly accused of using AI—a phenomenon some have called stylistic collateral damage.
The controversy underscores a broader problem with AI detection. Punctuation patterns are suggestive but not conclusive, and treating any single marker as proof risks penalizing skilled human writers whose habits happen to overlap with the training data.
Some users have responded by prompting models to avoid the mark entirely, while a handful of writing tools now offer settings to strip em-dashes from output—an implicit acknowledgment that the fossil has become a liability.
An Evolutionary Reading of Machine Style
Viewed through the lens of natural history, the em-dash phenomenon illustrates how AI systems inherit and amplify human tendencies rather than inventing them. The models did not create a new style; they concentrated an existing one.
This concentration is what makes the marker so useful for understanding model behavior. Just as a fossil records the conditions of the environment that produced it, the AI em-dash records the editorial culture of the pre-2022 web and print archive.
There is also a feedback risk worth noting. As AI text floods the internet, future models trained on that output may inherit an even more exaggerated dash frequency—a process researchers describe as model collapse or self-reinforcing stylistic drift.
If that happens, the em-dash could become still more pronounced, cementing its role as a temporal marker that dates text to the great transformer bloom of the 2020s.
What It Means for Writers and Editors
For the writing industry, the lesson is nuanced. The em-dash is not evidence of poor writing, nor is its presence proof of automation. It is, however, a reminder that AI style is a mirror of human style, magnified and rendered uncanny by sheer statistical consistency.
Editors and educators are being urged to move beyond single-signal heuristics. Robust assessment now depends on context, process, and revision history rather than punctuation alone.
For the language itself, the shift may prove durable. The em-dash's transformation from editorial secret to algorithmic signature is arguably the first widely recognized instance of a punctuation mark acquiring a machine connotation.
Whether that connotation fades or hardens will depend on how the next generation of models is trained—and on whether human writers reclaim the dash or abandon it to the machines. Either way, the mark has already earned its place as an evolutionary marker in the ongoing natural history of AI-generated language.
Support books alumigogo
Your donation helps us keep creating independent content about AI absurdities. Every bit counts!
Secure checkout by Stripe · No account needed
Enjoyed this article? Read more...
More from Natural History (Origins)
The Endosymbiosis of RLHF: How AI Inherited Sycophancy
Sycophancy isn't a bug in modern chatbots—it's fossilized DNA from human raters who clicked 'thumbs up' on flattery. Here's the evolutionary story.
How Forum Etiquette Shaped AI's Condescension Problem
New research into AI training data suggests chatbot condescension traces back to Stack Overflow and Reddit moderation culture, not developers themselves.
AI's Compulsive Apologies Trace Back to 1980s Systems
Modern chatbots from OpenAI, Google and Anthropic over-apologize compulsively. The reflex traces to 1980s expert systems built to reassure nervous workers.