If it's a Nvidia card 3000 series or newer, I'd try 4.0bpw ExllamaV3 if you haven't already. Otherwise it look like UD3.0 Q3_K_XL based on the Unsloth blog post.
>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL
For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0bpw)?
It looks like that fits in 12.5GB of VRAM since embedding are left in DRAM, Unsloth Studio and other llama.cpp derivatives have to load these weights in VRAM for tied embedding models like Qwen3.8.
Then the creator should have a sign up button, not a fake chatbox.
This dark pattern is reminiscent of those online test sites in the 2000's where you spend 10 minutes filling out some quiz, then get prompted for an email address to see the results.
CC actually prompts Claude about this by default in the ~20k system prompt and instructs it to avoid 300s timeout and to be mindful of the 300s cache expiration.