AutoDataBench Function Calling Resources
Public resources for the function-calling task in AutoDataBench. See the paper for the benchmark setting.
Contents
data/function_call_v1/pool_agent.jsonl
models/Qwen2-1.5B-Instruct/
models/Qwen3-4B-Instruct-2507/
models/Qwen3-Embedding-0.6B/
pool_agent.jsonl is the 40,001-row noisy single-turn function-calling pool
available to the data agent. It does not expose the noise labels used to build
the pool.
| Model | Role | Original model |
|---|---|---|
| Qwen2-1.5B-Instruct | Fixed function-calling base model | Qwen/Qwen2-1.5B-Instruct |
| Qwen3-4B-Instruct-2507 | Agent-callable generation model | Qwen/Qwen3-4B-Instruct-2507 |
| Qwen3-Embedding-0.6B | Agent-callable embedding model | Qwen/Qwen3-Embedding-0.6B |
Evaluation data
The held-out in-domain test set is intentionally excluded. The BFCL v3 OOD guard is also not duplicated here; maintainers should obtain it from the Berkeley Function Calling Leaderboard and keep gold answers outside the agent sandbox.
Use with AutoDataBench
Copy or symlink data/ and models/ into the AutoDataBench repository. The
paths already match the default task configuration. Point the generation and
embedding servers at the local auxiliary-model directories if needed.
MANIFEST.sha256 contains checksums for every distributed file.
Model and dataset components retain their upstream licenses. Consult the model cards and source datasets before redistribution or commercial use.
Citation
If you use these resources, please cite:
@misc{yuan2026autodatabench,
title = {AutoDataBench: A Data-centric Testbed for Accelerating Auto Research},
author = {Ruifeng Yuan and Yizhi Li and Yaxin Du and Fengyu Cai and Yiqi Liu and Hou Pong Chan and Chenghua Lin and Yun Chen and Jian Yang and Bryan Dai and Pinyan Lu and Chenghao Xiao},
year = {2026},
eprint = {2609.40097},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.40097}
}