Artificial intelligence in Wikimedia projects
Artificial intelligence reached Wikipedia's editing rooms not through any official announcement, but through a quiet experiment on the 6th of December 2022. A contributor named Pharos published a new article that day, writing the initial draft with ChatGPT. The subject was filed as "Artwork title." A second editor tried the same tool on a topic called "Weaponized incompetence" and found the overview decent, but every citation was fabricated. That gap between fluent prose and invented sources defined a debate that has not settled since. The questions it raised are large: Can machine-generated text meet the standards a volunteer reference project demands? And what does it mean for a community built on human expertise when machines trained on that expertise begin writing in its place? Answering either question requires going back further than 2022.
Since 2002, automated programs have been permitted to edit Wikipedia, though each must pass a human approval process and remain under ongoing supervision. The first significant example was rambot, which drew on United States census records to generate stub articles about American towns, cities, and counties. The vast majority of those local geography articles on English Wikipedia today began as rambot's output. Vandal-fighting became the next major application. ClueBot, released in 2007, applied simple heuristics to flag likely damage to article pages. Its 2010 successor, ClueBot NG, replaced heuristics with an artificial neural network, bringing live machine learning into Wikipedia's editorial defenses. Machine translation software also became available to contributors around this period.
Aaron Halfaker launched the Objective Revision Evaluation Service, known as ORES, in late 2015. ORES scored individual Wikipedia edits for quality, offering editors a rapid signal about whether a given change was likely an improvement. Each of these tools worked in a supporting role: catching vandalism, converting structured data, evaluating existing edits. None of them composed prose in the manner of a human contributor. The generative models that arrived after 2022 crossed that line.
The Wiki Education Foundation, reporting in early 2023, described a divided landscape among experienced editors. Some found ChatGPT useful for starting drafts or producing outlines for new articles. The Foundation's assessment came with a pointed observation. ChatGPT understood what a Wikipedia article should look like and could produce one in the right format. But the Foundation also warned that the tool had a persistent tendency toward promotional language and other reliability problems. Three concerns kept surfacing in community discussions about large language models. They produced misinformation that sounded plausible. They wrote in a tone that was not encyclopedic. And they reproduced biases from their training data. A proposal to ban AI tools outright was raised in 2023 and rejected as too strict. The productivity benefits to some editors were real, the community said. That same year, members of the English Wikipedia community created a WikiProject to track and remove poor-quality material produced by these tools.
Miguel García, a former Wikimedia member from Spain, reflected in 2024 that the volume of machine-written articles had peaked when ChatGPT first launched. Since then, he said, the rate had stabilized because the community had worked actively to address it. He noted that articles with no supporting sources were typically deleted instantly or put into a deletion queue. In October 2024, a Princeton University study examined roughly 3,000 articles created on English Wikipedia in August of that year. It found that about 5 percent had been written using AI. Some covered straightforward topics and appeared to involve only minor machine assistance. Others had been used to promote businesses or political interests. Ilyas Lebleu, founder of WikiProject AI Cleanup, told colleagues in October 2024 that the community had identified a recognizable pattern. The writing sounded authoritative while being entirely invented. Some of the articles the team removed turned out to be complete hoaxes.
In June 2025, the Wikimedia Foundation began testing a feature called Simple Article Summaries. The feature would provide machine-generated overviews of Wikipedia articles in a format similar to Google Search's AI Overviews. Critics called it a "ghastly idea" and a "PR hype stunt." They argued that the tendency of large language models to produce false information would erode reader trust in the site. The Foundation halted the rollout that same June while expressing continued interest in bringing generative tools into Wikipedia's operations. The episode revealed a gap between what the Foundation sought and what a significant portion of the editing community was prepared to accept.
English Wikipedia created a formal policy in August 2025 allowing editors to nominate articles suspected of being machine-written for speedy deletion. The community had assembled a working list of tells. Fabricated citations, or citations unrelated to the article's subject, were one indicator. Particular phrases were nearly diagnostic. Text beginning with "Here is your Wikipedia article on" or containing the phrase "Up to my last training update" was flagged for removal. Stylistic patterns also drew attention, including heavy use of em dashes and repeated reliance on the word "moreover." The word "breathtaking" appearing in otherwise neutral prose was another signal. Curly quotation marks in place of the straight versions required by style guidelines were also noted. That same month, the community published a formal written guide to identifying these patterns.
One article reviewer, speaking during the policy discussion, said he was "flooded non-stop with horrendous drafts" produced with AI tools. Other editors described the submissions as containing a large amount of "lies and fake references" and said correcting those problems consumed significant time. Ilyas Lebleu described the speedy deletion policy in August 2025 as a "band-aid." Some nominated articles, he noted, could remain on the site for as long as a week while deletion discussions ran their course. In March 2026, English Wikipedia voted to prohibit the use of large language models to add or rewrite article content. Two exceptions were preserved: copyediting one's own writing and machine translation from another language's Wikipedia. The vote came after years of debate about a technology that had also, quietly, become one of the project's most consequential data sources.
A 2017 paper described Wikipedia as "the mother lode for human-generated text available for machine learning," and that characterization proved durable. As of 2023, Wikipedia's text corpus was considered one of the largest well-curated data sets available for training language models. According to Stephen Harrison, every major large language model developed up to that point had trained on it. That same relationship produced other practical tools. Google's Perspective API, built to identify toxic comments in online forums, used a dataset of hundreds of thousands of Wikipedia talk page comments. Each comment had been labelled by human reviewers for its level of toxicity.
The Wikimedia Foundation and many supporters of its projects expressed concern that this widespread use of Wikipedia's text was rarely acknowledged. Wikipedia's licensing terms permit reuse, including in modified forms, but require that credit be given. The Foundation argued that AI products like ChatGPT, Siri, and Alexa frequently drew on Wikipedia's text to answer questions without noting the source. It said this pattern may have put those products outside Wikipedia's licensing terms. The Foundation's practical concern went further. Users who did not know they were benefiting from Wikipedia had less reason to visit the site or donate to sustain it. By 2025, the Foundation had observed an 8 percent drop in visitors to Wikipedia. It attributed the decline partly to the growth of generative AI and partly to social media drawing traffic away. The Foundation also acknowledged growing costs from AI companies scraping its content at scale. In response, it began exploring special licensing deals and paid programming interfaces to recover some of those expenses and deliver its data more efficiently. Whether those arrangements would satisfy either the AI industry or Wikipedia's editing community remained an open question.
Slate Magazine, writing in 2023, quoted longtime Wikipedians including Richard Knipel and Andrew Lih. Both worried that Wikipedia risked losing what Slate described as its "original bold spirit" by developing a reflexive resistance to change. The magazine's position was that AI should be embraced under guidelines requiring transparency and human oversight.
Jimmy Wales, in 2025, proposed a more specific middle ground. Rather than using AI to write articles, he suggested deploying it to support the editing process itself. Specifically, he called for AI to provide tailored feedback when draft submissions were rejected. He also proposed that it could spot inconsistencies, flag missing information, and summarize editorial discussions. Editors arriving late to a debate could use those summaries to catch up quickly. That approach preserved human authorship while directing machine analysis toward tasks that would reduce the workload on volunteer reviewers.
Community guidance circulated since 2023 had flagged a risk that formal policy did not directly address. Copying AI text into Wikipedia articles, that guidance warned, carried the potential for libel or copyright infringement. Those concerns had received less attention in the debate than the more immediate problems of fabricated citations and non-encyclopedic tone. They pointed toward unresolved questions about the project's long-term relationship with the tools it had restricted.
Common questions
What is the history of automated tools and artificial intelligence on Wikipedia?
Wikipedia has allowed approved bots to edit its pages since 2002. Early programs like rambot converted census data into local geography articles, and the 2010 ClueBot NG applied machine learning to vandal detection. The Objective Revision Evaluation Service, known as ORES, launched in late 2015 to score the quality of individual edits.
How does Wikipedia identify articles written by artificial intelligence?
English Wikipedia looks for fabricated citations, particular phrases such as "Here is your Wikipedia article on" or "Up to my last training update," heavy use of em dashes, and the word "breathtaking" in otherwise neutral prose. The community published a formal guide to these patterns in August 2025.
Why did English Wikipedia ban AI from writing articles in March 2026?
In March 2026, English Wikipedia's community voted to prohibit using large language models to add or rewrite article content. The decision followed years of concern about AI tools producing misinformation, fabricated citations, and non-encyclopedic prose. Two exceptions were preserved: copyediting one's own writing and machine translation from other language versions of Wikipedia.
How has Wikipedia been used to train artificial intelligence language models?
As of 2023, Wikipedia's text corpus ranked among the largest well-curated data sets available for training AI models. Stephen Harrison reported that every major large language model developed to that date had trained on it. The Wikimedia Foundation expressed concern that many AI products use this text without the attribution Wikipedia's licensing terms require.
What did the Princeton University study find about AI articles on Wikipedia?
A Princeton University study published in October 2024 found that about 5 percent of roughly 3,000 newly created English Wikipedia articles were written using AI. Those articles had all been created in August 2024. Some covered ordinary topics with only minor machine assistance; others had been used to promote businesses or political interests.
How did Jimmy Wales propose using artificial intelligence to improve Wikipedia?
Jimmy Wales proposed in 2025 that AI support the editing process rather than write article text. He described specific uses: tailored feedback on rejected drafts, spotting inconsistencies, flagging missing information, and summarizing discussions for editors joining a debate mid-stream.
All sources
5 references cited across the entry
- 6This machine kills trollsJesse Hicks — 2014-02-18
- 7JournalScaling neural machine translation to 200 languagesMarta R. Costa-jussà et al. — June 2024
- 8BookHandbook of the Changing World Language MapVirginie Mamadouh — Springer International Publishing — 2020
- 30NewsGoogle's comment-ranking system will be a hit with the alt-rightViolet Blue — 2017-09-01
- 44BookThe Seven Rules of Trust: A Blueprint for Building Things That LastJimmy Wales — The Crown Publishing Group — 2025