Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
Abstract
Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.
Community
When several people are talking, a voice assistant needs to know not only what to say, but whether it should speak at all.
We introduce Duplex-MPE, a benchmark for selective participation in continuous multi-party conversations. It contains 2,000 paired scenarios with three or four human speakers and an assistant, contrasting explicit and implicit addressing. Models receive audio without transcripts or supplied turn boundaries. Separate metrics evaluate response initiation, answer correctness, remaining silent, and stopping when a human resolves a request.
Explore the audio examples and benchmark design: https://step-out.github.io/Duplex-MPE-Page/
We are preparing the data and code for release before December.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents (2026)
- MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant (2026)
- Same Words, Different Actions: Paired Turn-Taking Evaluation under Rewritten Dialogue Contexts (2026)
- Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents (2026)
- From Metrics to Natural Dialogue: French Full-Duplex Benchmark for Spoken Dialogue Models (2026)
- SteerDuplex: Steerable Duplex Speech Dialogue Models (2026)
- ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.31948 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper