Multi-agent consultation

BaRe-Mem

Bayesian Reliability Memory for Robust and Adaptive Agent Consultation

Peilin Feng1 Zhengyang Huang2 Soujanya Poria1,†

1DeCLaRe Lab, Nanyang Technological University 2Peking University

† Corresponding author

Reliability memory update: equal reliability without evidence; the posterior moves after right, right and wrong verified answers
The BaRe-Mem reliability memory update. (a) Without verified evidence, all advisors are assigned equal reliability. (b) Updating only candidate 1 with two correct outcomes shifts its estimated reliability to the red posterior, while a subsequent incorrect outcome yields the blue posterior. The annotated gains denote the corresponding Kalman gains.

Abstract

TL;DR  An online Bayesian memory of which advisors to trust, on which questions, and of when the central model is better off answering alone.

In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.

How BaRe-Mem works

BaRe-Mem estimates contextual reliability for both the central model and its advisors from verified interaction history. These estimates serve two roles: modulating the influence of advisor responses and determining whether consultation is preferable to autonomous reasoning.

01

Organise historical reliability

The frozen central model encodes each candidate answer in the context of the question. An exact Bayesian posterior over these belief representations gives every candidate's reliability, read before the answer and updated from its verified outcome after it.

02

Reliability-guided attention

Advisor reliability estimates steer the central model's attention toward the advisors the memory trusts, so a reliable advisor weighs more in the answer without any text added to the prompt.

03

Decide whether to consult

With T the highest advisor reliability and κ the central model's autonomous ability, BaRe-Mem estimates both abilities on the current question and selects the mode with higher estimated accuracy.

consult if  T·ρ + (1 − T)(κ − δ) ≥ κ

Results

Robust to misleading advisors

Across nine benchmarks in two capability regimes, a growing share of advisor responses is replaced by misleading ones: fluent, on-topic and well-formed, but verified to be incorrect.

Accuracy against the misleading information ratio for Qwen3-14B and Phi-4
Accuracy under increasing misleading-advice ratios for Qwen3-14B and Phi-4. Left: capability-supported regime; right: capability-challenging regime. Across both regimes, BaRe-Mem remains above the no-consultation baseline as misleading information increases, while other consultation methods degrade substantially in the capability-challenging regime.

It knows its own ability

Estimated autonomous ability against empirical autonomous accuracy
Estimated autonomous ability κ versus empirical autonomous accuracy for Qwen3-14B and Phi-4. Left: evolution of mean κ and empirical autonomous accuracy along the question stream. Right: empirical autonomous accuracy for groups of questions with similar κ values. The dashed diagonal indicates perfect calibration.

Useful from sparse feedback

BaRe-Mem accuracy against the share of questions with verified feedback
BaRe-Mem accuracy under sparse verified feedback for Qwen3-14B and Phi-4. The fraction of questions whose verified outcomes are written to memory varies from 0% to 100%. Insets magnify the low-feedback regime below 1%.

Choosing workers in agent teams

The same memory ranks the workers of a lead-worker team: the lead reads it to choose who handles each sub-task, and every verified report is written back.

Agent team pipeline with BaRe-Mem
Agent team pipeline with BaRe-Mem. The lead agent decomposes a task into sub-tasks and queries the reliability memory to rank candidate workers. Each returned report is verified by the lead agent or an external verifier: accepted reports are committed, while rejected reports trigger the next candidate. Verification outcomes are then written back to update advisor reliability.
Agent-team task completion on MuSiQue
Agent-team task completion with Qwen3-14B and Phi-4 as lead agents. Random routing, routing by historical success counts, and BaRe-Mem under three verification settings: no verification, verification by the lead agent, and exact verification by the dataset evaluator.

Data and models

Questions per benchmark and advisor accuracy under misleading information
The released question streams. (a) The capability-supported regime (GSM8K, SQuAD, APPS) and the capability-challenging regime (PIQA, MMLU, OpenBookQA, SciQ, BBH, SuperGLUE). (b) As the misleading information ratio grows, the share of right advisor answers (solid) and of questions with at least one right advisor (dashed).

BibTeX

@misc{feng2026baremembayesianreliabilitymemory,
      title={BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation},
      author={Peilin Feng and Zhengyang Huang and Soujanya Poria},
      year={2026},
      eprint={2609.35551},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.35551},
}