Imagine asking a super-smart robot for health advice, like a doctor’s assistant, who can instantly answer questions. Sounds amazing, right? But what if that robot only understands English really well and struggles with other languages like Swahili or Mandarin? That’s what this research uncovers—AI, while brilliant in English, loses its magic when it comes to understanding health topics in different languages.
In this study, researchers tested AI on over 9,000 health claims from various topics such as COVID-19 and political health issues. They wanted to see how these smart systems, trained using language from textbooks and online sources, responded to health statements in 21 different languages. Shockingly, even the top AI systems stumbled when dealing with non-European languages, and their accuracy varied depending on the topic and source of information.
This is a big deal. Imagine you’re in a country where English isn’t the main language, and AI is being used to share health advice. If this AI isn’t spot-on due to language barriers, there’s a risk of spreading misinformation. To avoid this, the study emphasizes the need for thorough testing of AI in multiple languages before it becomes widely used in global health communication. After all, no one wants crucial health information lost in translation!
The most advanced AI systems still trip over non-European languages, showing how complex language truly is!
FAQs
Why is AI struggling with non-European languages in health communication?
AI models are often trained primarily on data from English and other European languages, which means they don’t have enough diverse language data to understand topics accurately in non-European languages.
What topics were AI models tested on for this research?
AI models were evaluated using health assertions related to a range of topics including abortion, COVID-19, and politics. These were sourced from journals, government advisories, social media, and news.
What is the risk of using unvalidated AI for global health communication?
Using AI without comprehensive validation could lead to misinformation or misinterpretation of health advice, especially in countries where the primary language is not well-represented in AI training data.
How many languages were AI models tested in during this study?
Researchers tested the AI models using health claims in 21 different languages to assess their multilingual performance.
Why is it important to validate AI across various languages and topics before use?
Validating AI ensures accurate and reliable communication in global health initiatives by accounting for different languages and diverse informational needs across topics.
Background
Artificial Intelligence (AI) is often trained using large datasets from various sources, like textbooks or online articles. These datasets are usually biased toward English and other major European languages. When AI is used to process and understand languages it hasn’t been trained on extensively, its accuracy can degrade substantially, especially in specialized fields like health.
History
The development of AI in health communication has improved significantly due to advancements in machine learning and natural language processing. However, most of this progress has been centered around languages with extensive digital resources, leaving many languages – particularly those outside Europe – underserved. This study builds on previous research highlighting these gaps by specifically testing AI across a diverse range of languages and health topics.
Based on “Artificial Intelligence health advice accuracy varies across languages and contexts” by Prashant Garg, Thiemo Fetzer, available on arXiv (arxiv.org/abs/2504.18310), used under CC BY 4.0 (creativecommons.org/licenses/by/4.0/).





































































