PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization Paper • 2608.30597 • Published 20 days ago • 26
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 12 days ago • 78
Simulated Scholar Search (S3) — Models & Datasets Collection Datasets and model checkpoints for training and studying scientific-literature search agents in the S3 local environment. • 3 items • Updated Jul 27 • 3
Simulated Scholar Search (S3) — Models & Datasets Collection Datasets and model checkpoints for training and studying scientific-literature search agents in the S3 local environment. • 3 items • Updated Jul 27 • 3
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 36
Simulated Scholar Search (S3) — Models & Datasets Collection Datasets and model checkpoints for training and studying scientific-literature search agents in the S3 local environment. • 3 items • Updated Jul 27 • 3
VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design Paper • 2605.10978 • Published May 13 • 19