techdotus

Tokle-3M

text-generationmittransformers

huggingface.co/techdotus/Tokle-3M

Updated 2026-10-01 ·Open on Hugging Face →
transformers · safetensors · tokle · text-generation · causal-lm · custom-architecture · custom_code · slm · small-language-model · en · dataset:HuggingFaceFW/fineweb-edu · dataset:HuggingFaceTB/cosmopedia · dataset:agentlans/high-quality-english-sentences · dataset:nampdn-ai/tiny-strange-textbooks · dataset:armanc/ScienceQA · dataset:nvidia/OpenMathInstruct-2 · dataset:microsoft/orca-math-word-problems-200k · license:mit · region:us

Tokle-3M

Model Summary

Tokle-3M is a decoder-only language model with 2.91M parameters. It was first trained on 12B tokens with SPAB (Static Pairwise Attention Bias), a frozen table of 8.39M token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus, giving 11.3M parameters in total during this stage. During training, for every query-key pair, SPAB hashed the two token IDs into the table, retrieved their PMI value, scaled it by a learned per-head factor, and added it to the attention logits before softmax.

After this stage, the SPAB table was removed and the model was trained for an additional 0.5B tokens to distill the knowledge in the SPAB matrix into its own layers. As a result, Tokle-3M runs entirely on its 2.91M parameters at inference, with no SPAB table required.

Model Architecture

| Parameter | Value | |---|---| | Architecture | Decoder-only transformer (RMSNorm, RoPE, GQA, SwiGLU) | | Layers | 9 | | Hidden size (d_model) | 144 | | Attention heads | 3 | | KV heads (GQA) | 1 (multi-query attention) | | Head dim | 48 | | FFN intermediate size | 432 | | Max sequence length | 512 | | Tie word embeddings | Yes | | Precision | FP32 weights | | Parameters | 2.91M |

How to use

```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "techdotus/Tokle-3M" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()

ids = tok("The climate change", return_tensors="pt") with torch.no_grad(): out = model.generate(**ids, max_new_tokens=32, do_sample=False, repetition_penalty=1.3) # greedy print(tok.decode(out[0], skip_special_tokens=True)) ```

Benchmark Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology.

| HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 | |---|---|---|---|---| | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |

Ablation: SPAB vs. Distilled

Stage 1 model (SPAB active) vs. the released Tokle-3M (SPAB removed and distilled).

| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 | |---|---|---|---|---|---|---|---| | Tokle-SPAB-11.3M (with SPAB) | 11.3M (2.91M trainable + 8.39M frozen) | 9.16 | 27.22% | 34.68% | 24.49% | 54.95% | 41.70% | | Tokle-3M (SPAB distilled) | 2.91M | 8.92 | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |

Comparison Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology. Scores for the other models are from the Open SLM Leaderboard. Bold marks the best result in each column.

| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 | |---|---|---|---|---|---|---|---| | Tokle-3M (Tech.us) | 2.91M | 8.92 | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% | | Ember-2 (SurjoLabs) | 2.96M×2 | 7.21 | 27.28% | 33.42% | 22.01% | 55.11% | 35.90% | | BananaMind-2-Micro (BananaMind) | 2.9M | 6.01 | 28.27% | 33.12% | 21.93% | 53.21% | 34.00% | | GPT-S-1.4M (Axiomic Labs) | 1.4M | 5.40 | 26.89% | 31.57% | 21.93% | 55.17% | 30.20% |

Training Details

Tokle-3M was trained in two stages on the same data mixture.

| Stage | Tokens | SPAB | Parameters | |---|---|---|---| | 1. Pretraining | 12B | Active (frozen PMI table) | 11.3M (2.91M trainable + 8.39M frozen) | | 2. Distillation | 0.5B | Removed | 2.91M |

Stage 2 lets the trained weights absorb the prior the SPAB table had been providing, so the released model is self-contained rather than losing that knowledge when the table is removed.

Training Data

We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.

| Source | Percentage | |---|---| | FineWeb-Edu | 43.1% | | Cosmopedia | 24.3% | | OpenMathInstruct-2 | 13.5% | | Tiny Strange Textbooks | 9.0% | | MegaScience (medicine & biology, custom curated) | 5.0% | | High-Quality English Sentences | 3.0% | | ScienceQA | 1.2% | | Orca-Math Word Problems 200k | 0.9% | | Total | 100% |

Limitations

License

Model weights and code: MIT.

Citation

``bibtex @misc{tokle2026, title = {{Tokle-3M}: Pointwise Mutual Information as a Removable Inductive Bias for Self-Attention}, author = {{Tech.us Team}}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/techdotus/Tokle-3M}} } ``

Mirrored from the Hugging Face Hub and served from the Conceptio Open Knowledge Archive. Read the original card at https://huggingface.co/techdotus/Tokle-3M.