Agents that argue about architecture
My master's thesis asks whether LLM agents that debate a software architecture judge it better than one model on its own.
The problem
Evaluating a software architecture properly, with a method like ATAM, is slow expert work. Language models can help with parts of it, and letting several of them debate has been reported to improve their reasoning. The catch is that much of that benefit disappears once the single model gets the same compute budget.
What I built
A multi-agent system in which agents with different roles argue about architecture alternatives for microservice systems and produce what an ATAM evaluation produces, risks, sensitivity points, trade-offs and a recommendation. Before building it I reviewed 871 studies. None of them applied a debate to architecture evaluation, and only 45 of 360 debate studies compared against a baseline with the same budget. So the experiment compares the debate with a single model and with repeated sampling at equal cost, measured against published evaluations by human experts.
What came out
The literature review and the evaluation pipeline are done. The experiment runs in October and November 2026, submission is in early 2027. Every number in the thesis will trace back to a file, a script and a commit.
Built with
Python, LLM APIs, uv, pytest, a systematic literature review
I’m writing it at the Herman Hollerith Zentrum of Reutlingen University. This page gets the results once they exist.