Abstract
TL;DR An online Bayesian memory of which advisors to trust, on which questions, and of when the central model is better off answering alone.
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.
How BaRe-Mem works
BaRe-Mem estimates contextual reliability for both the central model and its advisors from verified interaction history. These estimates serve two roles: modulating the influence of advisor responses and determining whether consultation is preferable to autonomous reasoning.
Organise historical reliability
The frozen central model encodes each candidate answer in the context of the question. An exact Bayesian posterior over these belief representations gives every candidate's reliability, read before the answer and updated from its verified outcome after it.
Reliability-guided attention
Advisor reliability estimates steer the central model's attention toward the advisors the memory trusts, so a reliable advisor weighs more in the answer without any text added to the prompt.
Decide whether to consult
With T the highest advisor reliability and κ the central model's autonomous ability, BaRe-Mem estimates both abilities on the current question and selects the mode with higher estimated accuracy.
consult if T·ρ + (1 − T)(κ − δ) ≥ κResults
Robust to misleading advisors
Across nine benchmarks in two capability regimes, a growing share of advisor responses is replaced by misleading ones: fluent, on-topic and well-formed, but verified to be incorrect.
It knows its own ability
Useful from sparse feedback
Choosing workers in agent teams
The same memory ranks the workers of a lead-worker team: the lead reads it to choose who handles each sub-task, and every verified report is written back.
Data and models
Six advisors
Six central models
- Qwen3-4B
- Qwen3-8B
- Qwen3-14B main text
- Ministral-8B
- Qwen2.5-7B
- Phi-4 main text
BibTeX
@misc{feng2026baremembayesianreliabilitymemory,
title={BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation},
author={Peilin Feng and Zhengyang Huang and Soujanya Poria},
year={2026},
eprint={2609.35551},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.35551},
}