DEAR EDITOR,
In a recent contribution,1 the authors compared the quality of replies provided by different large language-based chatbots (LLMs) to questions related to diagnosis, treatment, and prognosis of pulmonary thromboembolism. The study design was based on 10 questions, all submitted to four different chatbots and their replies were scored by specialists. This study is composed of blocks of 4 dependent answers always related to one question. Therefore, the Kruskal-Wallis test, which has been created for independent data, is not adequate. When incorrectly applied for nested data, this test will provoke pseudo-replication.2 The test power will be distorted, thus provoking unreliable results. Adequate solutions would be tests for repeated measures, such as mixed models, Friedman tests with post-hoc tests and alpha-error corrections. Eventually, ANOVA tests for dependent data, when assumptions were met.
Here we tried to find out whether freely available LLM-based chatbots could help to select the adequate statistical tests. The published abstract was sent to 14 different chatbots with the following prompt:
“Please analyze this text and look for major problems, such as methodological errors, inconsistencies, or contradictions. Pay special attention to the statistical tests and examine whether they are well aligned with the study design. Please suggest corrections. Should we change a test by another one?“
The following 12 algorithms recognized the nested study design and suggested substitution by the above-mentioned adequate test. appointing it as the most important error: ChatGPT Free, Claude Sonnet 4.6, Consensus Corpus All, Copilot Smart, Deep Seek Fast DeepThink, Gemini 3.5 Flash, Grok Fast, Julius 1.2 Lite, Meta AI Instant, Notebook LM Free, Perplexity Learn Tutor. Furthermore, other questions regarding test power, data distribution, counting procedures, post-hoc tests, alpha-error correction, and inter-rater variability were also correctly discussed.
The remaining two bots (Bohrium Expert and Mistral Le Chat Balanced) did not recognize the nested study design and therefore did not indicate correction of the most important error.
Since replies of different LLMs to the same prompt or repeated submissions may lead to diverging results,3-5 we always recommend to consult several LLMs and eventually, to discuss conflicting answers with the chatbots until a final solution can be reached.6


