Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment Paper • 2609.38972 • Published 10 days ago • 18
Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI Paper • 2609.24815 • Published 17 days ago • 12
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 12 days ago • 53
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 18 days ago • 56
CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies Paper • 2609.24118 • Published 19 days ago • 28
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles Paper • 2609.22220 • Published Sep 2 • 7