Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research Paper • 2608.24306 • Published 26 days ago • 7
You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding Paper • 2608.29834 • Published 21 days ago • 9
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35 • 7
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published Aug 17 • 35 • 7
Power-Softmax: Towards Secure LLM Inference over Encrypted Data Paper • 2410.09457 • Published Oct 12, 2024
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation Paper • 2504.17502 • Published Apr 24, 2025 • 55
Intervention Lens: from Representation Surgery to String Counterfactuals Paper • 2402.11355 • Published Feb 17, 2024 • 2
Intervention Lens: from Representation Surgery to String Counterfactuals Paper • 2402.11355 • Published Feb 17, 2024 • 2
Evaluating D-MERIT of Partial-annotation on Information Retrieval Paper • 2406.16048 • Published Jun 23, 2024 • 36