SIA Lab

Knowledge Management and AI · Knowledge management; human–AI innovation

Preserving Disagreement in Human–AI Innovation: An LLM Multi-Agent Replay Study

Minority concerns that consensus discards validate at nearly the same rate as the ones it keeps.

Stage
Accepted · OUI 2026
Last activity
2026-06
Target venue
Open & User Innovation Conference 2026 (Harvard Business School)
Collaborators
Shengsheng Huang (Texas A&M)

Consensus is the default stopping rule in multi-agent review: when several agents inspect the same artifact, the majority position becomes the output and the rest is discarded. This project asks what is inside the discarded pile. We replayed real merged Kubernetes pull requests through a panel of specialized LLM review agents, recorded every concern raised, and then followed each concern forward into three months of independent human activity across the wider repository ecosystem.

The dismissed concerns did not behave like noise. They were re-raised and acted on by human developers at almost exactly the rate of the concerns the panel adopted, which means the consensus rule was not separating signal from noise so much as choosing arbitrarily between two comparably useful sets. The practical implication is architectural: coordination residue should be preserved and made retrievable, not thrown away at the moment of aggregation.

The record

Idea
Treat minority concerns discarded by consensus-oriented multi-agent systems as potentially valuable coordination residue, then validate those concerns against later independent human activity.
Research Question
Does inter-agent disagreement discarded by a consensus rule contain innovation-relevant knowledge that subsequently appears in independent human activity?
Key Proposition
Consensus does not reliably separate useful concerns from noise: minority disagreement can carry downstream-validated knowledge and should be preserved, organized, and made retrievable rather than erased.
Data
Replay of 100 randomly sampled merged Kubernetes pull requests through four specialized LLM review agents; 615 distinct concerns partitioned into 204 adopted and 411 dismissed concerns; each concern traced across seven repositories over a three-month forward window and graded on a seven-level validation scale.
Analysis Result
Of 411 dismissed concerns, 149 (36.3%) were independently re-raised, 68 (16.5%) actively addressed, 14 (3.4%) reached high-value validation, and 6 (1.5%) mapped to incidents. Dismissed and adopted concerns validated at nearly identical rates through active addressing: 36.3% vs. 34.8% for re-raising and 16.5% vs. 17.2% for addressing. Consensus provided some lift only at the highest-severity tail.

← All research