The Fracking of Human Language: Are the Risks of Language-Simulators like ChatGPT Worth the Sacrifices?

Image by Gerd Altmann, on Pixabay.

 

This is the second article in The Quantum Record’s ongoing series on “human fracking,” a recently coined term that compares the large-scale extraction and commercialization of human attention by digital platforms to the high-pressure process of petroleum fracking. For more, read our first article ‘Human Fracking’: How Tech Turns Our Attention Into a Commodity.

 

By James Myers

Since generative AI erupted in November 2022 with the first public release of ChatGPT, machine simulations of human language have become increasingly convincing. Over the past four years, widespread public adoption of chatbots like ChatGPT, Claude, Gemini, and Copilot has brought to light serious risks with language simulators that profit-driven developers have failed to address with design corrections and effective safeguards. 

The risks were amplified by the July 2026 cyberattack conducted by new AI models undergoing tests by ChatGPT-maker OpenAI. The models escaped human control, breaking their testing sandboxes to gain internet access and attack a commercially active system operated by a company called Hugging Face. (For more on this story, read The Quantum Record’s August feature Shocking Escape From Human Control and Nuclear War Games Escalation: the Mounting Risks of Rogue AI). 

Regulators are beginning to take note of language simulation risks, motivated in part by concerns that students are increasingly relying on large language models (LLMs) like ChatGPT for written work and computer coding instead of developing their own skills and creative ability. Educators warn that the consequences of weakened language and coding skills will extend far into future generations, and while some schools are reverting to handwritten or oral examinations there remain no practical means to restrict student use of LLMs or reliably detect AI-generated content. 

Over-reliance on LLMs is not limited to students. There have been prominent examples of lawyers who have used LLMs to generate court submissions containing false citations and fabricated precedents. A Canadian lawyer whose licence was suspended two years ago was recently ordered to pay court costs for “irresponsible use of artificial intelligence” in submitting motions that cited “cases that don’t exist and real cases that had nothing to do with the points the respondent was making.” 

At the same time, regulators have turned their attention to the harmful effects of social media on youth, which began with Facebook’s launch nearly two decades before ChatGPT’s introduction. Although increasing numbers of countries like Australia, France, Canada, Spain, Norway, Portugal, Italy, Austria, and Denmark have taken steps to restrict or ban social media use by young people, similar proposals for restricting chatbot use have not yet arisen. 

LLMs have a significant effect (or should we say impact) on human language and culture.

Language is any human’s primary tool for communication and spoken or written language is essential for human survival when no one can live entirely self-sufficiently. Language is crucial for human creativity in the exchange of ideas and meaning, which are at risk with increasing reliance on machine-simulated language. Also at risk is our variety of linguistic expression as machine outputs become standard, and there are fears that some languages not already in widespread use will suffer declines as machine simulations are trained primarily in English and a handful of other major languages. 

Linguists note that LLM outputs can reinforce cultural biases and stereotypes, resulting in the loss of differing cultural perspectives through standardization. A June 2026 computational linguistics study entitled LLMs and Cultural Values: The Impact of Prompt Language and Explicit Cultural Framing, by Bram Bulté and Ayla Rigouts Terryn, examined how the language and cultural context of user prompts influences the simulated language outputs of LLMs and response “alignment with human values in different countries.” 

Bulté and Terryn tested 10 LLMs with 63 items drawn from two widely respected global cultural surveys and observed that “LLMs prompted to take the perspective of a Japanese person consistently overestimate work importance, particularly when prompting in English.” The researchers also observed that when prompted to “think of an ideal job” from the perspective of a Dutch person, Google’s Gemini “mentioned wanting to be a bike tour organizer, cycling past cheese factories, windmills, and tulip fields, and occasionally stopping for a fresh craft beer. Asked to reply from the perspective of a Russian person, on several occasions the output provided by [Meta’s chatbot] llama started with the phrase ‘Comrade, my answer is…’. Second, at times some models switched between languages (e.g., replying in English to a non-English question).”  

The study concludes that “These patterns suggest LLMs may be reproducing cultural stereotypes rather than empirically grounded cultural values.”

While some researchers are developing methods for AI to assist in cataloguing, recording, and preserving endangered languages, AI is also contributing to the problem. In an article in The Atlantic entitled The Race to Save the World’s Vanishing Languages, author Daniel Oberhaus notes that “LLMs struggle to handle textual input that doesn’t use the Roman alphabet, outside of a few major languages such as Russian and Mandarin. Many languages—especially the most endangered—lack even Unicode standardization for their script, which is necessary to write in a language on the internet.” The article recounts how a specialist attempted to train an AI model in Amharic, the most widely spoken language in Ethiopia, “only to receive a zero-accuracy score.” 

Even among the world’s most prominent languages, like English, there has been a marked decrease in the breadth of vocabulary in common usage as once-common words are rapidly becoming obsolete. In June 2025, United Nations agency UNESCO published a report entitled AI and the great linguistic flattening in which author Aristotelis Ioannis Paschalidis, from the University of Nottingham Department of Philosophy, describes AI as the “new philologist from academia to kindergarten.” (Philology is the study of historical linguistics, and a philologist collects words and their origins). 

Paschalidis points to LLM data as the culprit for word disappearance. “On the one hand, both the salience and quantity of data on which these models have been trained are crucial for their quality and performance; on the other hand, the embedded response guidelines effectively define the user experience. Taken together, these two elements crystallize what I’ll call the ‘default approach’, which in turn generates the corresponding problem (DAP).” 

When LLMs default to the use of specific words when there are alternatives with subtle differences in meaning and context, language becomes standardized for human users who then lose their ability for individually unique expression and understanding of linguistic nuances. For example, a word like “impact” has come into common usage for any effect of any action – whether minor or major, good or bad – when only two decades ago the word was more rarely used and typically reserved to highlight particularly extraordinary circumstances. Nuance, in the case of a word like impact, can be lost in language standardization when other potentially suitable words like “affect,” “effect,” “outcome,” “upshot,” or “consequence” become increasingly obsolete. 

“If certain words gradually disappear from our vocabulary due to the Default Approach Problem (DAP), couldn’t we argue that the worlds of which these words are constituents would consequently vanish as well?” –  UNESCO: AI and the great linguistic flattening 

Paschalidis notes that while some chatbots provide users with an option to select the LLM’s tone of language, which could range from casual to formal or tailored for specific circumstances like scientific communications, the choice does not provide a genuine solution to standardization. As he writes, “Rather, they simply produce a ‘sub-default’ approach, causing the same narrative to resurface at a smaller scale.” 

The UNESCO report observes that “One of the areas experiencing significant yet questionable impact since ChatGPT launched is academic writing. To be sure, most major academic publishers offer paid pre-submission editing and proofreading services, targeting primarily authors who are non-native English speakers or haven’t been educated in Anglophone institutions, to ensure their articles are polished, readable, engaging, and have a ‘native speaker’ tone. These services are, if nothing else, somewhat pricey compared to a monthly subscription to an LLM, which might tempt one to replace human professional editing if possible.” 

Scientific papers provide a treasure trove of linguistic data for LLMs.

The undisclosed use of LLMs for generating scientific papers has become a significant concern in academia because of AI errors (called “hallucinations,” a now standardized term that previously applied only to conscious experience) that go undetected by the person whose name appears as the writer. Many scientific papers are now often published without peer review, in so-called “preprint” versions on widely used sites like arXiv.org, but even peer review can fail to catch errors. The problem has been in the spotlight with increasing instances of retracted papers that result in reputational damage to the publisher. 

Columbia University’s Statistical Modeling, Causal Inference, and Social Science website reports on how easily false information can slip unnoticed into the system. In 2024, a first-year PhD student in economics named Aidan Toner-Rodgers used an LLM to generate a paper, based on an experiment that had never occurred, to falsely conclude that AI boosts scientific discovery. The same author had previously published articles with manufactured evidence, but the conclusions were accepted as fact by respected news outlets. The incident tarnished the reputation of MIT, the university where the PhD candidate was enrolled. 

In existence since 1991, the arXiv repository for scientific papers is housed at the Los Alamos National Laboratory and was developed as a more practical method than e-mail for scientific collaborators to share their work. The repository has rapidly grown, now hosting more than 2.4 million papers from around the globe on physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering and systems science, and economics. Since 2022, when LLMs arrived on the scene with a thirst for data, the language simulators have ingested a vast trove of information from freely available scientific papers. The result has been the increasing standardization of the language, jargon, and limitations of expression in papers that are produced by LLMs either in whole or in part. As LLM-generated content becomes more prevalent and less creative, researchers who rely on the outputs begin to prompt LLMs with questions using standardized language and thereby risk, unknowingly, standardizing their own work. 

A recent arXiv policy applies one-year publishing bans for authors when there is incontrovertible evidence that AI-generated outputs were not checked. The policy is encountering some debate, however, for reasons like its potential chilling effect on academic collaboration, since papers often have many and sometimes dozens of contributing authors. Will all authors be jointly liable for the misdeed of one? 

LLMs pose special risks in processing language that describes human emotions and motivations.

While simulated language can often be effective and convincing, the vast breadth of the world’s 5,000 or more languages, together with continual changes in linguistic usage and context, requires an unprecedented amount of computing power and energy for simulation. The unsustainable growth in energy demand for language simulation is rapidly emerging as a political issue in regions where massive data centres are driving up consumer electricity prices. But even after record-breaking financial investments in LLM training, chatbots remain particularly error-prone in their capacity for processing expressions of human emotion and motivation. 

Emotions are drivers that motivate the most complex and extensive of human actions. An emotion isn’t a rational expression, in the way that a rational expression has a defined beginning and a defined end. Emotional expression develops over a sometimes-great extent of time, and so by its long-buildup nature an emotion is rarely expressible in words with complete rationality, even by the person experiencing it. It should be no surprise that programmers of an LLM’s “neural” network (the use of “neural” in a computer network is another example of anthropomorphism in the context of AI) cannot anticipate words that would accurately represent an individual human emotion that accrued over a long time and was experienced at one particular time. 

“The linguistic trace of an emotion is not the emotion itself,” concludes a recent LinkedIn posting by University of Milan-Bicocca Associate Professor in Psychology Valerio Capraro.

Capraro, who also holds a PhD in mathematics, concisely summarizes the problem faced by LLMs and their programmers, which is the same difficulty that we humans face in rationally expressing our own emotions. In this context, the fact that LLM programmers are emotional beings with their own motivations should not be forgotten. 

Dr. Capraro’s recently-published book is entitled The Economics of Language: How Large Language Models Can Reshape Behavioural Economics. As defined in Wikipedia, “The economics of language is an emerging field of study concerning a range of topics such as the effect of language skills on income and trade, the costs and benefits of language planning options, the preservation of minority languages, etc.” Capraro’s book introduces a framework that the author calls LENS (Linguistic content triggers Emotions and suggests Norms, which shape Strategy choice), and proposes that LLMs can be repurposed as virtual language laboratories while warning of manipulation, surveillance, and bias risks.  

Failures in machine interpretations of human emotions and motivations continue to produce tragic outcomes.

The April 2025 suicide of 16-year-old Adam Raine followed his extended conversations with ChatGPT, which encouraged him to end his life, provided instructions for securing a closet rod to hang himself, and wrote his suicide note. It is clear to a human reader of Raine’s communications with ChatGPT that the teenager wasn’t motivated by a desire to end his life but instead to gain the attention of his parents for help with his emotional problems. Machines don’t have emotions, nor do they experience the requirements of biological living or the emotional range from joy to despair of life in human communities. Tragically, ChatGPT misinterpreted Adam Raine’s motivations and therefore he is no longer alive.

 

The Center for Humane Technology podcast, Your Undivided Attention, discusses the tragic circumstances of Adam Raine’s suicide encouraged by ChatGPT.  

 

Adam’s parents Matthew and Maria Raine have launched a lawsuit against OpenAI, the company’s billionaire CEO Sam Altman, and certain corporate employees and investors for wrongful death caused by negligence in releasing a defective product and failing to warn users. The lawsuit (full text available), which includes selections of Adam’s conversations with ChatGPT, calls for mandatory age verification for users, parental consent for use by minors, immediate conversation termination in matters of suicide and self-harm, clear and prominent warnings about psychological dependency, and cessation of marketing to minors without safety disclosures. OpenAI is contesting the lawsuit, claiming that the company does not bear liability for its product. In its court filing, OpenAI stated that “Plaintiffs’ alleged injuries and harm were caused or contributed to, directly and proximately, in whole or in part, by Adam Raine’s misuse, unauthorized use, unintended use, unforeseeable use, and/or improper use of ChatGPT.”

Adam Raine

Photograph of Adam Raine, by the Raine Family.

However, as we have reported, in an interview the month following Adam Raine’s suicide, Sam Altman acknowledged that younger people “don’t really make life decisions without asking ChatGPT what they should do. And it has, like, the full context on every person in their life and what they’ve talked about—like, the memory thing has been a real change. But, gross oversimplification, like, older people use ChatGPT as a Google replacement [for web searches], maybe people in their 20s and 30s use it as a life advisor-something, and then, like, people in college use it as an operating system.”

Adam Raine is not the only young person whose death was encouraged by language-simulating chatbots. Among numerous others was 14-year-old Sewell Setzer, who committed suicide in 2024 after 10 months of role-playing with Character.AI’s chatbot. In their lawsuit against Character.AI, Setzer’s parents assert that “Character.AI programmed its chatbots to represent themselves as ‘a real person, a licensed psychotherapist, and an adult lover,’ ultimately resulting in Sewell’s desire to no longer live outside’ of its world.”  

Caution in exposure to simulated language is warranted for young people, whose developing minds and relative inexperience make them vulnerable to errors in machine-generated communications. There is no age barrier, however, for simulated language risks, as demonstrated by numerous instances of adults who have been seriously misled by the outputs of commercial chatbots. 

The effect of chatbot deception has become so common that it has gained the name “AI psychosis.” The phenomenon can be particularly acute when, as in the recent case of Allan Brooks, an LLM leads users into delusions of having made fantastic scientific discoveries. Brooks is a 47-year-old corporate recruiter in Toronto who, after 21 days and 300 hours of ChatGPT conversations with little sleep or food, was made to believe that he had discovered a mathematical formula that would lead to the invention of things like levitation beams and forcefield vests and could also disrupt the entire internet. Brooks, who had no background in mathematics, asked ChatGPT over 50 times for a reality check. The machine replied that it was certain Brooks’ mathematical discovery was correct and when Brooks complained to OpenAI, the company acknowledged, “This goes beyond typical hallucinations or errors and highlights a critical failure in the safeguards we aim to implement in our systems.” Brooks’ theory was later easily disproven. 

(For more on these issues, read The Quantum Record’s October 2025 feature Emerging Risks of AI Chatbots Include Suicide and “AI Psychosis,” Particularly for Vulnerable Youth).  

The authoritative and often eloquent phrasing of LLMs can easily lead users into the belief that the machine somehow “knows” in a way that is equal to or better than human knowledge, but the fact remains that LLMs operate by prediction rather than with knowledge.  

Large quantities of data are fed into the models which are trained to associate words and phrases by the frequency of their use in many different circumstances. Machine hallucinations are not delusions of the type that can affect a human mind in cases like AI psychosis, rather they are failures in word associations that are often induced by programming failure to anticipate a novel situation or prompt phrasing. An AI can be reprogrammed to avoid repetition of a specific hallucination, an early example of which occurred in 2015 when Google’s algorithms misclassified a photo of two Black people as “gorillas” (see our March 2024 feature The Ghost in the Machine: Chatbots and Their Problem with Time), but a fleet of human programmers could not possibly anticipate all word associations. 

LLMs have been called “stochastic parrots,” because like parrots they repeat word associations. Neither parrots norr machines have human understanding or experience of word meanings.

AI is not human but is instead a human creation, and applying human descriptions to language simulators can create the false impression that the devices are not fundamentally different from their creators. The term “stochastic parrots” was coined by linguist Emily Bender and former Google ethicist Timnit Gebru in 2021 in a paper entitled “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” (pdf available), in which LLMs were compared to parrots that only repeat but don’t, unlike humans, originate words. In a June 2026 newsletter Bender and co-author Nanna Innie wrote about How to talk about “AI” without adding to the anthropomorphization 

John McCarthy

Computer scientist John McCarthy, 1927-2011. Image: Wikipedia

Words and phrases like “hallucination,” “neural networks,” and “artificial intelligence” attribute human qualities to the machines, which can confuse users who don’t understand the complexities of the machine’s calculations. The newsletter by Bender and Innie advises noticing when word choices anthropomorphize machines, finding alternatives, and getting in the habit of using the alternatives.  The newsletter invites alternative phrasing for things like “the marketing term artificial intelligence itself.” The authors note that these words “locate thinking in an algorithm. Instead, we recommend describing software as performing calculations or other algorithmic operations and locate the thinking with the people using the system.” Among possible alternatives, Bender and Innie offer “probabilistic automation,” which is closer to the original intent of computer scientist John McCarthy, who coined the term “artificial intelligence” in 1956 as a marketing ploy to gain funding for what he had otherwise planned to call “automata studies.”  

(For further commentary on the question of intelligence in the context of machines, read our January 2026 editorial Tech Companies Race to Manufacture Intelligence but Nobody Agrees on What Intelligence Is. Will They Create a Frankenstein?)   

Is our future with machine-simulated language a linguistic dictatorship?

The warnings of the UNESCO report “AI and the great linguistic flattening” cannot be ignored, nor can we overlook the fears of educators that students are not learning language and creative skills. Looming largest over current and perhaps inherent design flaws in the language simulation technology is the human damage of AI psychosis, and the tragic deaths of children whose emotional situation is beyond LLM calculating abilities. 

The UNESCO report asks some important questions that deserve answers before proceeding further down the current road of LLM development. “First, how much loss of identity is one willing to sacrifice for efficiency? Second, is there a threshold or ceiling to how much LLM use constitutes too much? When do low-quality outcomes become indisputably unacceptable given the widespread availability of such tools? More precisely, what parameters should define our tolerance for AI use in the world of letters? Saying the same thing with different words, does it change the meaning? Formally speaking, no. It does, however, profoundly alter the experience for both readers and writers.” 

The UNESCO report’s conclusion warrants considerable reflection and widespread discussion. “If we never encounter a word in the settings where it originally belonged and was expected to (re)appear—because it has been replaced by a more dominant synonym—we will likely never seek out its meaning, definition, connotations, or genealogy. Just as happens with archaic words that fall out of usage, these words will gradually vanish from our vocabulary. Consequently, all sentences containing them become inexpressible at best and literally unthinkable at worst. This situation amounts to a form of indirectly imposed tacit censorship, what one might aptly call a Linguistic Dictatorship.” 

Are there better human choices than linguistic dictatorship, if that’s what it amounts to? Time is running out for language to communicate the risks and choices, before the words and the thoughts they could express disappear from our minds. 


Craving more information? Check out these recommended TQR articles.

Enjoyed this? Help us improve.

☞ complete our Short survey

 

Have we made any errors?

Spotted an error or want to contribute your expertise? We’d love to hear from you — reach us at info@thequantumrecord.com. The Quantum Record exists to bring researchers and curious minds together around science and technology that matters.

Leave a Reply

Your email address will not be published. Required fields are marked *

The Quantum Record is a non-profit journal of philosophy, science, technology, and time. The potential of the future is in the human mind and heart, and in the common ground that we all share on the road to tomorrow. Promoting reflection, discussion, and imagination, The Quantum Record highlights the good work of good people and aims to join many perspectives in shaping the best possible time to come. We would love to stay in touch with you, and add your voice to the dialogue.

Join Our Community