Zdravotnictví
AI Recommends Medications That No Longer Exist
mlj.solutions

mlj.solutions Vědci to otestovali na 103 podobných otázkách. Modely bez kontroly opakovaně doporučovaly léky, které už roky neměly být v odpovědi. Pak přidali kontrolní vrstvu – systém, který každou odpověď porovná s aktuálními daty, než se odešle dál. Nebezpečné odpovědi klesly o polovinu. 👉 Jak máte vyřešenou auditní vrstvu u AI systémů ve vaší firmě?
View on Instagrammlj.solutions

mlj.solutions Jsi nejistý a zeptáš se AI, co by ti doporučila. Jenže co když ti poradí něco, co už dávno neexistuje? Třeba lék, co se bral před lety? 💊 Odpověď mohla být klidně kdysi správná, jenže mezitím se léky mohly stáhnout z trhu nebo se dokonce zakázat. AI totiž netuší, co se stalo později. Zůstala u toho, co se naučila, a tak odpovídá se samozřejmostí, i když to dávno může být jinak. V medicíně, financích nebo jiných citlivých oblastech to bohužel nestačí. Tam AI musí vědět, co platí dnes, a ne co platilo v den, kdy se učila. ⚠️ Řešíte tohle u AI systémů ve své firmě? Celý příběh ve čtvrtek!
View on InstagramWhat the Study Is About
Language models deployed in medicine hallucinate and recommend drugs that are now banned. The study tests whether this can be fixed with a multi-agent system using real regulatory verification.
53%
reduction in hallucination error rate
A multi-agent system with real regulatory verification cut the number of dangerous recommendations across all tested models.
99%
of models recommended a banned drug
By default, virtually every tested model consistently recommended withdrawn or banned medications.
What This Means
- High accuracy doesn't mean safety, a model may correctly identify a drug but miss that it's been pulled from the market.
- RAG alone isn't enough, the model finds the correct information but fails to carry it through to the final recommendation.
- User expertise doesn't matter, even experts miss the error when the model answers confidently.
- A solution exists, an agentic layer with real regulatory verification reduces risk without retraining the model.
Náš pohled
An AI system that answers fluently and logically doesn't necessarily mean it answers safely.
Michal Dobrovolný
Zdroj
Osama, M., Amjad, M., Mustansar, Z., Shaukat, A., & Khan, M. U. S. (2026). Trust but Verify: Mitigating Medical Hallucinations via Post-Hoc Adversarial Auditing and Multi-Agent Feedback Loops. arXiv:2606.14149. https://arxiv.org/abs/2606.14149
Keep reading
Don't miss these posts
Want to implement AI in your company?
We'll go through your needs together and propose an approach that delivers measurable results.
Write to us