Transluce Mental Health Suite Collection Data accompanying Transluce's Mental Health Behavior Report • 2 items • Updated 3 days ago • 3
CHIVE: Counterfactual Hypothesis Investigation Via Edits Collection Datasets and LoRA checkpoints for the CHIVE paper: Would this change your answer? Predicting the Effects of Prompt Edits on LLM Behavior in the Wild. • 13 items • Updated 27 days ago • 1
Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Paper • 2606.13603 • Published Jun 11
Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logic Paper • 2603.05198 • Published Mar 5
Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers Paper • 2507.07808 • Published Jul 10, 2025
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Diverse Deception Probes Collection Linear probes trained on diverse deception data to detect dishonest completions across model families (OLMo, Qwen, Gemma). • 5 items • Updated Mar 18 • 1