Imagine if computers were creating new words in languages they’re learning! That’s exactly what’s happening with Irish and artificial intelligence. Researchers discovered that when Large Language Models (or LLMs) try to translate into Irish, they often make up entirely new words that don’t exist in any dictionary. It’s not just a glitch; it happens more often with simplified versions of these language models!
This study looked closely at how these models generate words, classifying them into verbs and nouns, and found six unique patterns in how nouns are being miscreated—or ‘hallucinated,’ as the experts put it. Even more interesting, some of these made-up words actually follow the rules of the Irish language, while others are completely off-the-wall. The Mini version of the GPT-4.o model hilariously churns out more of these oddities than its larger counterpart.
So, why should you care? This is more than just a tech quirk—it could change the future of the Irish language! As these A.I. models become more ingrained in everyday life, they might start influencing how Irish speakers use the language. Imagine students using these models to learn Irish; they could accidentally pick up and spread these new, fictional words! Who’s to say some of these might not become part of the language someday? It’s a fascinating intersection of technology, culture, and language evolution that shows how dynamic and unpredictable language can be.
Did you know? Some of the gibberish words created by A.I. translations follow the rules of the Irish language perfectly, despite not being real words!
FAQs
What are hallucinations in Large Language Models?
Hallucinations in Large Language Models refer to instances where these models generate non-existent words or phrases, particularly when translating into other languages like Irish.
Why does this happen more with the GPT-4.o Mini model?
The GPT-4.o Mini model is a simplified version of the larger model, which seems to produce creative errors or ‘hallucinations’ more frequently, likely due to its limited training capacity.
How might these hallucinations affect the Irish language?
These generated words could influence speakers, especially those learning the language, potentially introducing new words into the modern lexicon and affecting language evolution over time.
Can LLMs create words that follow Irish grammar rules?
Yes, some made-up words generated by LLMs still adhere to Irish linguistic rules, making them appear more plausible to native speakers despite being fabricated.
What implications does this have for other languages?
For languages with fewer resources, like Irish, LLMs could play a role in shaping vocabulary and linguistic trends, illustrating the broader impact of technology on language dynamics.
Background
Large Language Models (LLMs) are a type of artificial intelligence designed to understand and generate human language. They learn by being fed vast amounts of text data, which allows them to predict and construct sentences. However, when translating into languages with fewer resources, like Irish, they sometimes ‘hallucinate’—or make up words that don’t exist, often due to gaps in their language data or over-generalization of rules.
History
The development of language models has evolved substantially over time, from simple text processors to complex neural networks capable of understanding context and nuances. Early models struggled with basic translation tasks, but innovations like GPT-3 and its successors have made significant strides. Nonetheless, challenges remain, particularly in translating and working with languages that aren’t as widely represented in training datasets, leading to phenomena like hallucination.
Based on “Synthetic Fluency: Hallucinations, Confabulations, and the Creation of Irish Words in LLM-Generated Translations” by Sheila Castilho, Zoe Fitzsimmons, Claire Holton, Aoife Mc Donagh, available on arXiv (arxiv.org/abs/2504.07680), used under CC BY 4.0 (creativecommons.org/licenses/by/4.0/).





































































