Security
AI Agents That Keep Each Other in Check
mlj.solutions

mlj.solutions AI agenti ve firmě mohou řešit složité procesy. Ale co když jeden selže? 🕵️♂️ POIROT: místo externího hodnotitele se agenti vyslýchají navzájem. Každý má jinou roli a jiný pohled - výsledkem je lepší fault detection než centrální LLM soudce. 🔍 Jak máte řešenou auditovatelnost ve vašich agentních systémech?
View on Instagrammlj.solutions

mlj.solutions POIROT testovali na 122 “zločinech” 🔍 Výsledek: o 12–28 % lepší odhalení viníka než klasický jeden vyšetřovatel. 🕵️ Nejlepší detektiv AI selhání nakonec nebyl člověk. Ani externí model. Byl to sám systém, který se dokázal ohlídat. 👉 Řešíte podobné “zločiny” ve svých AI systémech? Než na to pošlete detektiva zvenku, možná stačí, aby se agenti začali ptát sami sebe. Pokud tě zajímá víc, sleduj nás, nebo nám rovnou napiš!📲
View on InstagramWhat the Study Is About
POIROT is a protocol for multi-agent systems in which the agents diagnose failures themselves and determine which agent went wrong — replacing an external evaluator with the system's own collective intelligence.
1.60×
higher accuracy on complex tasks
POIROT outperforms a simple LLM evaluator, with the advantage growing as problem complexity increases.
42%
more correctly identified failures
In a clinical system, POIROT caught 42 more cases out of 100 than an approach using a single evaluator model.
32,768
possible failure combinations in a 15-agent system
In a test financial system with 12 agents, POIROT correctly identified the source of the error despite extreme complexity.
What This Means
- Centralized evaluation has blind spots, a single model acting as evaluator gets too much information at once and starts making mistakes or ignoring parts of the input.
- Agents themselves carry enough collective intelligence, to detect where the system failed.
- POIROT's advantage grows with complexity, the benefit is modest for simple tasks and significant for complex ones.
- No single model dominates universally, the right choice depends on the type of failure, not just size.
- Regulatory pressure will accelerate this — the EU AI Act requires auditability and traceability, exactly what POIROT addresses.
Náš pohled
In agentic systems, checking the result isn't enough. You need to know which agent decided, when, and why. Without that, a correct output doesn't mean a correct process.
Michal Dobrovolný
Founder & CTO
Zdroj
Dellibarda Varela, I., Sendra-Arranz, R., Romero-Sorozabal, P., Valverde-García, J. M., Laudanski, A. F., Gutiérrez, Á., Rocon, E., & Cebrian, M. (2026). POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems. arXiv:2606.02282. https://arxiv.org/abs/2606.02282
Keep reading
Don't miss these posts
Want to implement AI in your company?
We'll go through your needs together and propose an approach that delivers measurable results.
Write to us