Voodoo Quant

High Performance LLM Quantization

Voodoo Lowers Perplexity Up To 89.5%

Quantization refers to the technique used to squeeze smart AI into inexpensive hardware. Voodoo Quant is a new quantization technique which drastically increases the intelligence of quantized models. It sets a new pareto frontier in divergence to size ratio for all architectures and model sizes.

Voodoo Quant has the strongest effect at the most aggressive quantization levels, where it lowers Qwen3.5 0.8B IQ1_S PPL by 89.5%. This is exciting because now a whole new level of intelligence is available for inexpensive hardware. With the huge improvement to 1-bit that Voodoo Quant brings, there are now many more applications for aggressively compressed LLMs.

Qwen3.5 0.8B

KLD and PPL for FineWeb evaluation dataset. KLD measured against BF16 reference.

Torch KLD
Llama.cpp KLD
Llama.cpp PPL
Voodoo Model Baseline Model Baseline Size (MB) Voodoo Size (MB) Size ↑% Baseline PPL Voodoo PPL PPL ↓% Baseline KLD Voodoo KLD KLD ↓% Baseline KLD (llama.cpp) Voodoo KLD (llama.cpp) KLD ↓% (llama.cpp)
Qwen3.5-0.8B-MTP.Voodoo25_TQ1_0 - - 250.4 - - 419.37 - - 2.3383 - - 2.287051 -
Qwen3.5-0.8B-MTP.Voodoo30_IQ1_S Llama.cpp 330.0 283.1 -14.2% 1757.9 184.82 89.5% 3.748 1.4232 62.0% 3.521 1.369778 61.1%
Qwen3.5-0.8B-MTP.Voodoo35_IQ1_M Llama.cpp 338.4 321.6 -5.0% 599.59 100.17 83.3% 2.6694 0.7712 71.1% 2.586 0.661309 74.4%
Qwen3.5-0.8B-MTP.Voodoo40_IQ2_XXS Llama.cpp 338.3 365.6 8.1% 170.82 63.71 62.7% 1.3743 0.3834 72.1% 1.355 0.322473 76.2%
Qwen3.5-0.8B-MTP.Voodoo45_IQ2_M Unsloth 409.3 401.3 -2.0% 75.52 53.48 29.2% 0.6517 0.2046 68.6% 0.331673 0.181427 45.3%
Qwen3.5-0.8B-MTP.Voodoo50_IQ3_XXS Unsloth 429.1 440.1 2.6% 60.34 50.67 16.0% 0.4313 0.1421 67.1% 0.204900 0.122011 40.5%
Qwen3.5-0.8B-MTP.Voodoo55_Q3_K_S Unsloth 452.6 463.3 2.4% 58.05 48.91 15.7% 0.3818 0.0813 78.7% 0.194287 0.069977 64.0%
Qwen3.5-0.8B-MTP.Voodoo60_IQ4_XS Unsloth 524.7 515.0 -1.9% 47.73 47.55 0.4% 0.2125 0.0451 78.8% 0.042364 0.04142 2.2%
Qwen3.5-0.8B-MTP.Voodoo65_Q4_K_M Unsloth 549.7 551.0 0.2% 45.99 45.85 0.3% 0.1922 0.0485 74.8% 0.03496 0.04213 -20.5%
Qwen3.5-0.8B-MTP.Voodoo70_Q5_K_S Unsloth 585.5 571.3 -2.4% 46.38 47.38 -2.2% 0.1749 0.0225 87.1% 0.013882 0.021512 -55.0%
Qwen3.5-0.8B-MTP.Voodoo75_Q5_K_XL Unsloth 614.2 614.0 -0.0% 45.96 45.65 0.7% 0.1746 0.0155 91.1% 0.011618 0.015389 -32.5%
Qwen3.5-0.8B-MTP.Voodoo80_Q6_K Unsloth 658.1 667.5 1.4% 45.63 45.39 0.5% 0.1703 0.0084 95.1% 0.00466 0.009714 -108.5%
Official Qwen3.5-0.8B BF16 Qwen 1746.9 - - 37.2 - - 0.0000 - - 0.0000 - -

Qwen3.5 2B

KLD and PPL for FineWeb evaluation dataset. KLD measured against BF16 reference. Sizes include the MTP nextn-predictor sidecar; the Voodoo size percentage applies to the base model only.

Torch KLD
Llama.cpp KLD
Llama.cpp PPL
Voodoo Model Baseline Model Baseline Size (MB) Voodoo Size (MB) Size ↑% Baseline PPL Voodoo PPL PPL ↓% Baseline KLD Voodoo KLD KLD ↓% Baseline KLD (llama.cpp) Voodoo KLD (llama.cpp) KLD ↓% (llama.cpp)
Qwen3.5-2B-MTP.Voodoo25_TQ1_0 - - 617.3 - - 175.54 - - 2.5121 - - 2.4322 -
Qwen3.5-2B-MTP.Voodoo30_IQ1_S Llama.cpp 723.0 725.4 0.3% 171.04 35.42 79.3% 2.7719 0.9275 66.5% 2.3988 0.8071 66.4%
Qwen3.5-2B-MTP.Voodoo35_IQ1_M Llama.cpp 748.2 816.1 9.1% 73.97 25.13 66.0% 1.8778 0.4721 74.9% 1.5525 0.4469 71.2%
Qwen3.5-2B-MTP.Voodoo40_IQ2_XXS Llama.cpp 790.3 907.2 14.8% 38.26 20.2 47.2% 1.2041 0.2883 76.1% 0.8738 0.2443 72.0%
Qwen3.5-2B-MTP.Voodoo45_IQ2_M Unsloth 859.9 1001.5 16.5% 20.54 18.83 8.3% 0.4407 0.1596 63.8% 0.2469 0.1504 39.1%
Qwen3.5-2B-MTP.Voodoo50_IQ3_XXS Unsloth 931.8 1095.9 17.6% 18.81 17.35 7.8% 0.3163 0.1172 62.9% 0.1458 0.0982 32.6%
Qwen3.5-2B-MTP.Voodoo55_Q3_K_S Unsloth 1030.9 1170.7 13.6% 18.39 16.78 8.8% 0.292 0.0739 74.7% 0.1321 0.0628 52.5%
Qwen3.5-2B-MTP.Voodoo60_IQ4_XS Unsloth 1173.0 1289.5 9.9% 16.65 16.47 1.1% 0.1675 0.0351 79.1% 0.0313 0.0316 -1.0%
Qwen3.5-2B-MTP.Voodoo65_Q4_K_M Unsloth 1280.8 1373.2 7.2% 16.36 16.36 0.0% 0.156 0.029 81.4% 0.0202 0.0242 -20.0%
Qwen3.5-2B-MTP.Voodoo70_Q5_K_S Unsloth 1384.5 1428.2 3.2% 16.23 16.25 -0.2% 0.1428 0.018 87.4% 0.0088 0.0168 -92.1%
Qwen3.5-2B-MTP.Voodoo75_Q5_K_XL Unsloth 1466.7 1529.2 4.3% 16.2 16.22 -0.1% 0.1388 0.0116 91.6% 0.0076 0.0109 -44.2%
Qwen3.5-2B-MTP.Voodoo80_Q6_K Unsloth 1575.0 1674.3 6.3% 16.17 16.17 0.0% 0.1359 0.0055 95.9% 0.0042 0.0057 -35.1%

Questions & Answers

Are Voodoo models finetuned?
No, they are the original model weights.
Do Voodoo models have a different architecture?
No, their GGUF checkpoints run in unmodified Llama.cpp.
How are quality metrics so much better?
Voodoo Quant carefully selects quant levels for each tensor based on which are most important.
Will you quantize more models?
Yes. Watch our Huggingface account for more releases.
What software does Voodoo Quant support?
Voodoo Quant is generalized so it is compatible with all checkpoint formats and LLM software.

Enterprise

Voodoo Quant is available as a service for enterprise customers. We also offer Voodoo Quant Plus, which uses a minor architecture tweak compatible with all models to enable even stronger quantization. Please enquire below.

Contact