Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published Aug 3 • 40
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders Paper • 2510.19779 • Published Oct 22, 2025 • 62
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Paper • 2506.07180 • Published Jun 8, 2025