Model Gallery

128 models from 1 repositories

Filter by type:

Filter by tags:

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q8
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

ling-3.0-tiny-q8
Ling-3.0-tiny in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q8
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q8
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

qwen3.8-27b-uncensored-q8
Qwen3.8-27B-Uncensored in the higher-quality Q8_0 GGUF format, with its integrated MTP head and shared F16 vision projector.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated-q8
Huihui Qwen3.8 27B in Q8_0 GGUF format, with the shared BF16 vision projector and MTP speculative decoding through llama.cpp.

Repository: localaiLicense: apache-2.0

hy-mt2-1.8b-q8
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

carbon-3b-q8
Carbon-3B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.

Repository: localaiLicense: apache-2.0

carbon-8b-q8
Carbon-8B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.

Repository: localaiLicense: apache-2.0

ornith-1.0-9b-q8
Ornith-1.0-9B in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-9b-q8
Ornith-1.5-9B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-9b-obliterated-q8
Ornith-1.5-9B OBLITERATED in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. Its safety guardrails are removed, and the publisher recommends this quantization for better behavior fidelity.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

nex-n2.5-mini-q8
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q8_0 GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

thomson-1.0-small-q8
Thomson-1.0-Small in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: polyform-strict-1.0.0

tiel-coder-35b-a3b-q8-mtp
Tiel-Coder-35B-A3B in Q8_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

Page 1