Preview

SRISA Proceedings

Advanced search

Comparative analysis of multi-agent debate architectures of large language models in decision support

https://doi.org/10.25682/NIISI.2026.2.0003

Abstract

The paper presents a comparative analysis of five multi-agent debate architectures of large language models in a decision-support task related to evaluating the quality of educational tests. The study aims to identify how communication topology affects decision quality, computational cost, latency, and robustness. A unified experimental setting was used to compare a baseline, linear, cyclic, cross, and model-elimination architecture. The corpus comprised 60 tests from four school disciplines split into correct and errors subsets, while the model ensemble included Qwen 2.5 72B Instruct, Grok 4.1 Fast, and Gemma 4 26B-A4B IT. The main evaluation metrics were macro-F1 for detecting problematic questions, macro-MAE for the numerical assessment of test quality, and token usage, cost per test, and response time for operational comparison. The results show that the model-elimination architecture achieved the highest observed macro-F1 within the experiment, while the baseline architecture remained superior in macro-MAE and cost-efficiency. The study also demonstrates that some debate architectures are highly sensitive to the type of test and that increasing the number of rounds does not guarantee sustained quality improvement. The findings support treating debate architecture as a controllable design parameter of decision-support systems rather than as a neutral implementation detail.

About the Author

Sh. M. Magomedov
Дагестанский государственный технический университет, Махачкала
Russian Federation


References

1. Du Y., Li S., Torralba A., Tenenbaum J. B., Mordatch I. Improving factuality and reasoning in language models through multiagent debate [Электронный ресурс] // arXiv.org. – 2023. – arXiv: 2305.14325. – URL: https://arxiv.org/abs/2305.14325 (дата обращения: 20.02.2026). – Текст: электронный.

2. Li Y., Du Y., Zhang J. [et al.]. Improving multi-agent debate with sparse communication topology // Findings of the Association for Computational Linguistics: EMNLP 2024. – 2024. – P. 7281–7294. – URL: https://arxiv.org/abs/2406.11776 (дата обращения: 30.06.2026).

3. Zheng L., Chiang W.-L., Sheng Y. [et al.]. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena [Электронный ресурс] // arXiv.org. – 2023. – arXiv: 2306.05685. – URL: https://arxiv.org/abs/2306.05685 (дата обращения: 22.02.2026). – Текст: электронный.

4. Zhu L., Wang X., Wang X. JudgeLM: fine-tuned large language models are scalable judges [Электронный ресурс] // arXiv.org. – 2023. – arXiv: 2310.17631. – URL: https://arxiv.org/abs/2310.17631 (дата обращения: 25.02.2026). – Текст: электронный.

5. Guo T., Chen X., Wang Y. [et al.]. Large language model based multi-agents: a survey of progress and challenges // Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. – 2024. – URL: https://arxiv.org/pdf/2402.01680 (дата обращения: 30.06.2026).

6. Shi Y., Liang R., Xu Y. EducationQ: evaluating LLMs’ teaching capabilities through multi-agent dialogue framework // Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. – 2025. – P. 32799–32828. – URL: https://arxiv.org/html/2504.14928 (дата обращения: 30.06.2026).

7. Аванесов В. С. Форма тестовых заданий: учебное пособие для учителей школ, лицеев, преподавателей вузов и колледжей. – 2-е изд., перераб. и расшир. – Москва: Центр тестирования, 2005. – 156 с.

8. Челышкова М. Б. Теория и практика конструирования педагогических тестов. – Москва: Логос, 2002. – 432 с.

9. Приказ Министерства просвещения Российской Федерации от 31.05.2021 № 287 «Об утверждении федерального государственного образовательного стандарта основного общего образования» [Электронный ресурс]. URL: https://publication.pravo.gov.ru/Document/View/0001202107050027 (дата обращения: 13.03.2026). – Текст: электронный.

10. Федеральный институт педагогических измерений. Аналитические и методические материалы по разработке и оцениванию контрольных измерительных материалов [Электронный ресурс]. – URL: https://fipi.ru/ege/analiticheskie-i-metodicheskie-materialy (дата обращения: 13.03.2026). – Текст: электронный.


Review

For citations:


Magomedov Sh.M. Comparative analysis of multi-agent debate architectures of large language models in decision support. SRISA Proceedings. 2026;16(2):26-31. (In Russ.) https://doi.org/10.25682/NIISI.2026.2.0003

Views: 97

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2225-7349 (Print)
ISSN 3033-6422 (Online)