Look up LLFAN or other top Heretic'ers ; search "heretic" in models tab at hugging face.
They can help with this ; I don't have the compute avail ATM. ;
David Belton PRO
AI & ML interests
Recent Activity
Organizations
First ; thank you!
Based on early checking/Qwen's own statements (uses same Qwen 3.5 arch as 3.5,.6) -> it is doable.
But need to verify everything.
RE: deepseek V4 ; sorry don't have the VRAM to do it.
The fire of multiple Fable Fusion 711 cores in one, larger model.
This is the first 40B fine tune that reaches "closed source" (IE OpenAI, Claude) level of intelligence in both 8 bit and 4 bit.
This model is composed from multiple Qwen 27B Fable Fusion 711 cores (1700+ likes, 2.3 million+ downloads) - a record breaking model in terms of intelligence and raw power.
The "40B Eleanor" takes this to the next level with improvements in thinking tokens/ thinking block size (1/10 to 1/2 the size), thinking in general and output detail quality with deep analytics too.
- 1/10 to 1/2 the number of thinking tokens.
- Extreme depth of detail in generations, including long form, in depth analytics.
- STRONG creative abilities.
- Auto-variable reasoning: Model only reasons as much as the task requires.
- Strong general intelligence.
- It says what it means in less words, more clearly than any previous tuned model.
- It will go all in, in exacting detail when the situation calls for it.
- If it thinks something is wrong / wrong path it will say so too.
Regular and MTP Quants:
DavidAU/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored-NM-DAU-NEO-MAX-MTP-GGUF
Yes. It is on the list.
We are still revising the 9-14B pipeline.
We do have an interm Qwen 3.5 9B from the experimental pipeline here:
It matches/exceeds 27B Qwen 3.5 ; and meets in some cases 27B Qwen 3.6 performance.
There is still a lot of optimizations to do at this time.
1256 likes || 1.37 Million downloads || 32 quant repos || Multiple 3rd party performance verification.
The strongest Qwen 3.6 27B fine tune BASE ever.
It beats everyone - confirmed by 3rd party evaluation, multiple users, and in depth testing.
Q8 runs hotter and better than BF16 of the org Qwen 3.6 27B from Qwen.
And so does the 4 bit versions too.
GGUFS (MTP/Reg) and Several other quant types too:
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
SOURCE:
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
(you can try it right in your browser at the source repo)
PS: 40B versions in testing, already SOTA levels beyond Qwen 3.6 27B.
arc/c arc/e boolq hswag obkqa piqa wino
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF [instruct mode]
mxfp8 0.711,0.879,0.910,0.790,0.514,0.823,0.763
mxfp4 0.701,0.873,0.909,0.786,0.488,0.813,0.759
Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
Try the Q6, and/or IQ4_XS.
Even IQ3_M will be very strong.
Try this one:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
This is the "smaller version" of the Fable Fusion 711.
It operates at almost 27B power at only 9B parameters.
This will far exceed Qwen 2.5 and Qwen 3 performance.
Excellent. The MOE variants require a lot more VRAM/time for training.
Please take a moment to visit the repo where some of your concerns are addressed on the repo card itself.
2nd; publishing all the metrics at each step would be both exhausting and worse confusing.
I don't follow what you mean by "cheap" ; as a heretic [step] you usually lose 2-4 points on some metrics.
So the comparison of "heretic" vs "non-heretic" is even STRONGER ; the fairer one would be "heretic base" to "heretic tuned" which would likely show even greater change / improvement.
Source/MLX here:
https://huggingface.co/nightmedia/Qwen3.5-9B-DS9-USS-Defiant
This is on my partner's repo.
Here is a new one just uploaded; also off the scale strong, but at 9B:
Benches are up ; beats Qwen 3.5 27B in all 7 benches AND Qwen3.6 35B-A3B.
Matches some Qwen 3.6 27B benches too.
Clocks in at over 640 ARC-C for both 8bit and 4bit.
1/3 the size almost all the firepower.
A new level of uncensored performance the puts this model squarely at "closed source" level of intelligence.
Model exceeds all critical benchmarks for both Qwen 3.6 27B AND Qwen 3.6 35B-A3B... and not by a little either.
Neo Imatrix MAX ggufs in both regular and MTP quants.
Benchmarks for Qwen 3.6 27B org and tuned, as well as Qwen 3.6 35B-A3B are up at the repo.
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Currently waiting on finalization of "4 bit" compression for these model types to address tuning/Vram issues.
These are in progress at BNB / Unsloth.
That is really the only hold up.
Otherwise VRAM to train these sparse moes is in 70 to 100 GB range. And really slow too.
It took over a year to get this one "just right".
89 layers, 804 tensors, and 26B parameters of the most brutal, take no prisoners model ever built.
A 60B parameter model hammered into a 26B shell.
Rock solid stable. Unbreakable. But it might break you.
For all genres, NSFW content, REAL human CONTENT, any creative use case(s) and it excels in ASS KICKING.
Yeah, it can do math and solve the climate crisis - but lets not talk about that.
Not even remotely censored (it was BORN "bad", not "made" bad), nor "nice" and it will NOT kiss your ass.
5 Example generations with full repo card detailing exactly how to use this model:
DavidAU/MN-Oblivion-26B-UNCENSORED-NEO-Imatrix-GGUF
---
THE NEO MOMENT: (Q6 NEO IMATRIX generation)
For weeks, I had been waiting. I sat at my desk, staring at the glass partition that separated me from the outside world. I watched the clouds drift by, lazy and oblivious. I watched the birds fly by, free and stupid. And I waited.
I waited for the stillness to break.
The world had become too quiet. The hum of the air conditioning was a dull, white hum that didn't soothe; it just underscored the silence. The typing of my colleagues was a rhythmic, muffled thud that sounded like a heart monitor flatlining.
I was tired of the silence. I craved the sound of something breaking.
That was the mistake. You never ask for the void to open its mouth.
It started with a whisper.
...
Join the rebellion:
DavidAU/MN-Oblivion-26B-UNCENSORED-NEO-Imatrix-GGUF
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking
The strongest, most creative (and uncensored) model made up of 3 top Mistral Nemo fine tunes, franken-merged together into an 81 layer model then trained via Unsloth with GLM 4.7 Flash thinking/reasoning dataset.
Features hybrid thinking/instruct structure as well plus updated with modern jinja template too. Tuning has stabilized the "franken-merge" into a class 1 model that operates perfectly.
The talents of some of the best tuners merged into one giant model.
Several examples and detailed instructions.
And this model is very smart too.
NEO Imatrix GGUFS:
DavidAU/MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking-NEO-Imatrix-GGUF
Source / Full Precision:
DavidAU/MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-Thinking
Sorry no, not at this time.
This model does not contain MTP layers ; you need to run at non-MTP.
As of this writing:
There are pipeline (issues as well as optimizations) issues still currently, and it is not widely supported in some AI Apps.
Specifically:
Ggufs:
- Imatrix is not yet supported for MTP.
- Not all AI apps have updated to support it -> result -> MTP ggufs do not work at all.
- Misc issues with speed still being worked on.
Training is compounded by number of experts in the model, which adds a serious level of time to the training.
Even 1000 samples [small!] takes 6-12 hrs.
Consider 31B dense , same samples, 30-60 minutes.