True Intelligence, or Just a Smart Search Engine?
mlj.solutions

mlj.solutions AI složila testy s 99 % úspěšností. Pak dostala otázku, kterou nikdo nikdy nepublikoval. Projekt First Proof dal AI 10 matematických problémů přímo z aktivního výzkumu - ne školní úlohy, ne veřejné benchmarky. Otázky, které v den testování neexistovaly nikde na internetu. AI dokázala vysvětlit zadání a navrhnout postup. Sestavit nový důkaz od nuly - to je jiná liga. Rozdíl mezi rozpoznáváním vzoru a skutečně novým myšlením je větší, než veřejné žebříčky napovídají. 👉 Testuješ AI na vlastních datech - nebo věříš cizím žebříčkům?
Zobrazit na Instagramumlj.solutions

mlj.solutions AI zvládne každý test. 🧠 Ale co když dostane otázku, kterou nikdo nikdy nepublikoval? Výzkumníci dali AI 10 skutečných matematických problémů - čerstvých, neviděných, nikde na internetu. A výsledek? Ten se dozvíš brzy! 👉 Co si myslíš ty? Celý příběh ve čtvrtek. 🎬
Zobrazit na InstagramuWhat the study is about
Leading mathematicians compiled 10 unpublished questions from their own active research and presented them to the best AI models. No answer was available online. The goal was simple: to find out whether artificial intelligence can truly think or just search.
Key numbers
10
unpublished research questions
Each question arose naturally from the author's own research, was never made public, and the answer was encrypted before testing to prevent data leakage.
0
questions solved autonomously and correctly
Neither GPT-5.2 Pro nor Gemini 3.0 Deep Think were able to autonomously and correctly prove any of the ten questions in a single attempt.
20+
questions tested before selecting the final ten
The authors went through more than twenty candidate questions to find ones that AI could understand, yet were still beyond its capabilities.
What this means
- AI handles competition-style math, not research-level math, existing benchmarks test a different type of problem than what mathematicians actually work on.
- Models confabulate with confidence, in several cases AI cited a non-existent or incomplete proof as a finished result.
- A stronger model does not mean a correct proof, GPT-5.2 Pro and Gemini 3.0 failed on the same types of errors.
- Data contamination is a real problem, even with data sharing disabled, models may retain relevant context from training.
- A benchmark without automatic grading is hard to scale, the correctness of proofs must be assessed by a human expert, which slows down evaluation.
Our take
AI is good at mathematics, but not as good as it thinks it is. The gap between a confident answer and a correct proof is exactly what we need to watch out for — not just in mathematics.
Michal Dobrovolný
Zdroj
Source: Abouzaid, M., Blumberg, A. J., Hairer, M., Kileel, J., Kolda, T. G., Nelson, P. D., Spielman, D., Srivastava, N., Ward, R., Weinberger, S., & Williams, L. (2026). First Proof. arXiv:2602.05192. https://arxiv.org/abs/2602.05192
Don't miss these posts
Want to implement AI in your company?
We'll go through your needs together and propose an approach that delivers measurable results.
Write to us


