AI & ML interests

Open RL Environments at Scale

Recent Activity

AdithyaSKย  updated a dataset about 8 hours ago
FineEnvs/SmolDataEnvs
AdithyaSKย  updated a Space about 9 hours ago
FineEnvs/MiMo-RL-Envs-Explorer
AdithyaSKย  updated a dataset about 14 hours ago
FineEnvs/MiMo-V2.6-RL-harbor-music
View all activity

FineEnvs 's collections 8

MiMo-V2.6-RL in Harbor
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
Repo2RLEnv โ€” Verifiable RL Environments
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
Data Agent
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ€” verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
LaTeX OCR
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
Paint with Code
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
FineEnvs Academy
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
MiMo-V2.6-RL in Harbor
All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, set up and graded like Xiaomi's harness.
SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
Repo2RLEnv โ€” Verifiable RL Environments
Repo2RLEnv coding and terminal RL environments in Harbor format. Datasets include per-task quality labels, provenance and generation economics.
LaTeX OCR
LaTeX OCR environment, model, dataset, and five-run training comparison across four models, including unstable and stabilized Gemma.
Paint with Code
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
Data Agent
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ€” verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
FineEnvs Academy
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents