|
Download README.md from poolside/spec-decoding-subfolder-fixture: direct link, hf CLI and curl.
- Browser
- Download file 1.79 kB
-
https://huggingface.co/poolside/spec-decoding-subfolder-fixture/resolve/main/README.md
- Command line
-
hf download hf://poolside/spec-decoding-subfolder-fixture/README.md
-
curl -L -o README.md https://huggingface.co/poolside/spec-decoding-subfolder-fixture/resolve/main/README.md
1.79 kB
| license: apache-2.0 | |
| tags: | |
| - test-fixture | |
| # Speculative-decoding subfolder test fixture | |
| A **tiny, randomly-initialised** Llama pair for exercising *subfolder* loading of | |
| a speculative draft model. It exists so inference engines can test the code path | |
| in CI without downloading a real checkpoint. | |
| **These weights are random. The model produces gibberish. That is the point.** | |
| ``` | |
| . <- target model, 2 layers, ~4M params | |
| └── DFlash/ <- draft model, 1 layer, ~4M params | |
| ``` | |
| Both share the same tokenizer and vocabulary, so the draft is a valid speculator | |
| for the target. | |
| ## Why a subfolder needs engine support | |
| A Hugging Face repo id may contain **at most one `/`**, so `org/repo/DFlash` is | |
| not a valid id and cannot be passed as a single string. A draft bundled inside | |
| its target's repository therefore has to be addressed as a repo id *plus* a | |
| separate sub-path, which most engines have no parameter for. | |
| This matters in practice: shipping a draft inside its target repo keeps the pair | |
| versioned together and downloadable in one fetch, rather than as two repos that | |
| can drift apart. | |
| ## Usage | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| snapshot_download("poolside/spec-decoding-subfolder-fixture") # target + draft | |
| snapshot_download("poolside/spec-decoding-subfolder-fixture", | |
| allow_patterns=["DFlash/*"]) # draft only | |
| ``` | |
| Engines adding support should accept the draft as | |
| `(model="poolside/spec-decoding-subfolder-fixture", subfolder="DFlash")`. | |
| ## Provenance | |
| Built with `transformers` from `LlamaConfig` (hidden 64, intermediate 128, | |
| 4 heads, tied embeddings), `torch.manual_seed(0)`. Tokenizer copied from | |
| `hf-internal-testing/llama-tokenizer`. No trained parameters, no proprietary | |
| data. | |