DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate

Conference Paper

Abstract: We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task and examine whether agreement among multiple LLM evaluators can predict official evaluation performance. Frontier models generate strong debate responses and show high within-family agreement as evaluators, but their consensus does not reliably match the official target, particularly for the Quality maxim.

This paper extends our first-place submission to the Touché 2025 Retrieval-Augmented Debate task, exploring whether agreement among multiple LLM evaluators is a reliable proxy for official evaluation performance.

The paper was accepted for publication in the CLEF 2026 Best of Labs proceedings. The original submission code is available in the retrieval-augmented debate repository, with the follow-up analysis in the paper analysis repository.