Repository: localaiLicense: apache-2.0

Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.
Links
Tags
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.
Links
Tags
Repository: localaiLicense: mit
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.
Links
Tags

Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.
Links
Tags
Ling-3.0-tiny in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.
Links
Tags
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.
Links
Tags
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.
Links
Tags
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.
Links
Tags
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.
Links
Tags
Qwen3.8-27B-Uncensored in the higher-quality Q8_0 GGUF format, with its integrated MTP head and shared F16 vision projector.
Links
Tags

Huihui Qwen3.8 27B in Q8_0 GGUF format, with the shared BF16 vision projector and MTP speculative decoding through llama.cpp.
Links
Tags
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.
Links
Tags
Carbon-3B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.
Links
Tags
Carbon-8B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.
Links
Tags
Ornith-1.0-9B in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.
Links
Tags
Ornith-1.5-9B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.
Links
Tags
Ornith-1.5-9B OBLITERATED in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. Its safety guardrails are removed, and the publisher recommends this quantization for better behavior fidelity.
Links
Tags
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.
Links
Tags
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q8_0 GGUF build includes the F16 vision projector and uses the embedded chat template.
Links
Tags
Thomson-1.0-Small in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.
Links
Tags
Tiel-Coder-35B-A3B in Q8_K_XL format with MTP speculative decoding and a BF16 vision projector.
Links
Tags