zeusdevpro

Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K-DS4

text-generationapache-2.0gguf

huggingface.co/zeusdevpro/Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K-DS4

Updated 2026-09-22 ·Open on Hugging Face →
gguf · qwen · qwen3.8 · moe · ds4 · dwarfstar · uncensored · quantized · apple-silicon · metal · text-generation · en · zh · base_model:orcarouter/Qwen3.8-Flash-Next-Uncensored · base_model:finetune:orcarouter/Qwen3.8-Flash-Next-Uncensored · license:apache-2.0 · region:us

Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K

An uncensored quantized build of Qwen3.8-Flash-Next for local inference on Apple Silicon (Metal).

This repository provides a standalone Qwen3.8-Flash GGUF for DwarfStar (ds4). It substantially raises the precision of the original Q4 build, bringing a 128 GB system close to the quality of Unsloth's releases while staying compatible with DS4's loader format. The routed ffn_down_exps experts are upgraded to Q8_0, and the ffn_gate_exps / ffn_up_exps tensors of layers 0, 1, 2 and 45 are upgraded as well. For the remaining 44 layers the gate/up expert pairs use Unsloth's imatrix-calibrated Q4_K rather than the version found in the original file. Everything else, including the original BF16 n-gram table, is unchanged from the original. This approach yields gains comparable to Unsloth's Q5_K_XL build.

On top of the UD-Q4K mixed-quantization base, 57 routed expert tensors are upgraded to Q8_0.


1. This is not a "full Q8" quantization

The GGUF holds 1256 tensors, and only 57 of them are Q8_0 (4.5%):

| Type | Q4 base | Q8-UDQ4K | Change | |---|---:|---:|---| | F32 | 472 | 472 | — | | Q8_0 | 337 | 394 | +57 | | F16 | 298 | 298 | — | | Q4_K | 98 | 90 | −8 | | MXFP4 | 49 | 0 | −49 | | BF16 (n-gram) | 2 | 2 | — |

Of the 147 routed expert tensors: 57 are Q8_0 (38.8%) and 90 remain Q4_K (61.2%). All non-expert tensors (dense, attention, GDN, embedding, output, PLE, shared MTP parts) are left exactly as in the base.

The accurate description is: 57 routed expert tensors raised to Q8_0 on top of UD-Q4K, not a full-model Q8.

2. Which tensors were upgraded, and why

| Tensor | Layers | Count | Base → Target | |---|---|---:|---| | ffn_down_exps.weight | blk.0 ~ blk.48 | 49 | MXFP4 → Q8_0 | | ffn_gate_exps.weight | blk.0, blk.1, blk.2, blk.45 | 4 | Q4_K → Q8_0 | | ffn_up_exps.weight | same 4 layers | 4 | Q4_K → Q8_0 |

blk.48 is the model's embedded MTP block, not an ordinary transformer layer.

Why all down experts are upgraded and only 4 gate/up layers

A format constraint forces down onto a 32-block type. GGUF requires the row width (ne0) to be divisible by the block size. Measured shapes:

``text ffn_down_exps dims=(640, 2560, 512) -> ne0 = 640 -> 640 % 256 = 128 ✗ cannot use 256-block ffn_gate_exps dims=(2560, 640, 512) -> ne0 = 2560 -> 2560 % 256 = 0 ✓ 256-block available ``

640 / 256 = 2.5 is not an integer, so down cannot use 256-block types such as Q4_K / Q2_K / IQ2_XXS at all. The base uses MXFP4 there because of the format constraint, not out of preference.

Among the expert types ds4 supports, only three use block size 32:

| Format | Bytes per weight | Bits per weight | |---|---:|---:| | MXFP4 | 0.53125 | ~4.25 | | Q4_0 | 0.5625 | ~4.5 | | Q8_0 | 1.0625 | ~8.5 |

Q8_0 is the highest-precision member of the 32-block family, so upgrading down to Q8_0 takes the highest precision available under that constraint. The upgrade direction is monotonic.

gate/up have row width 2560, which permits a 256-block type, and the base already uses imatrix-calibrated Q4_K there. Q8_0 is more precise but doubles the size, and upgrading all 98 tensors would exceed a 128 GiB machine, so only 4 layers are upgraded: blk.0, blk.1 (errors propagate downstream), blk.2 (near the early n-gram / PLE injection), and blk.45 (the last transformer block, feeding the output path).

Q8_0 is a symmetric block quantizer and needs no imatrix, so this upgrade introduces no new calibration dependency.

Size increase

All 57 target tensors hold the same element count (640 × 2560 × 512 = 838,860,800):

``text down × 49: (0.8301 − 0.4150) × 49 = 20.34 GiB gate/up × 8: (0.8301 − 0.4395) × 8 = 3.12 GiB total = 23.46 GiB ``

Measured agreement: 165.11 + 23.46 = 188.57 GiB. For comparison, upgrading all 98 gate/up tensors would add 58.62 GiB and bring resident weights to about 131.5 GiB — above the 128 GiB ceiling.

Relation to other DS4 Qwen3.8 quantization builds

Sibling builds (such as ivanfioravanti/Qwen3.8-Flash-Next-DS4-Q4) use the same base recipe: gate/up = imatrix Q4_K, down = MXFP4, dense GDN/QSA/shared/output = Q8_0, embedding/control = BF16/F32, PLE kept in an external sidecar. The only difference in this build is that down is raised to Q8_0 and gate/up is raised on 4 layers.

3. Files and running

| Item | Value | |---|---| | File size | 188.57 GiB | | Resident weights | 93.19 GiB | | SHA-256 | f8f73355ccac3e654e80e0c7935f05a3bffde291440c8f5a9f5f2a325d144c62 | | Native context | 262,144 tokens (prompt + generation combined) | | Recommended --prefill-chunk | 4096 |

``sh ./ds4-server -m Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K.gguf \ --host 127.0.0.1 --port 8114 \ --ctx 262144 --prefill-chunk 4096 \ --mtp --mtp-draft 1 \ --kv-disk-dir /tmp/ds4-kv-qwen-q8 --kv-disk-space-mb 65536 \ --power 100 ``

Memory (ctx=262144, prefill-chunk=4096):

``text ds4: memory: KV 8.33 GiB (raw 6.50 + compressed 1.83) + buffers 5.79 GiB + resident model 93.19 GiB = 107.31 GiB planned ds4-server: context buffers 14456.69 MiB (ctx=262144, prefill_chunk=4096) ``

Planned 107.31 GiB; observed runtime around 121 GiB. 128 GiB unified memory is recommended, with other large applications closed. Under memory pressure, lowering --ctx is more effective than lowering --prefill-chunk.

4. File shards and reassembly

The weights are split into 2 shards to stay under Hugging Face's single-file guidance ( Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K.gguf ```

Verify:

``sh shasum -a 256 Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K.gguf # must be f8f73355ccac3e654e80e0c7935f05a3bffde291440c8f5a9f5f2a325d144c62 ``

The reassembled file must be 202,472,325,120 bytes. Any other size means a missing shard or wrong order.

5. Speed

Decode throughput observed in real use:

| Context | Decode speed | |---:|---:| | 64K | 66.46 t/s | | 128K | 52.56 t/s | | 256K | 47.05 t/s |

Loading a 23.8K-token tool prompt for the first time takes roughly 20–28 seconds.

6. Known limitations

1. --prefill-chunk 8192 OOMs at 256K context: Metal command batch failed: Insufficient Memory. Use 4096 or lower. 2. The native 262,144 context is prompt + generation combined. Once the prompt fills it, no generation is possible. For longer contexts, enable static YaRN with DS4_QWEN4_YARN_FACTOR=2 or =4. 3. The disk KV cache stores checkpoints only; it does not swap the running KV out to disk. The active context still occupies memory. 4. MTP speedup varies widely, so long outputs can decode noticeably slower than the table above.

7. Weight provenance and license

Please also comply with the license terms of both the base and upstream models.

Mirrored from the Hugging Face Hub and served from the Conceptio Open Knowledge Archive. Read the original card at https://huggingface.co/zeusdevpro/Qwen3.8-Flash-Next-Uncensored-Q8-UDQ4K-DS4.