Model Gallery

139 models from 1 repositories

Filter by type:

Filter by tags:

apodex-1.1-mini-q4
Apodex-1.1-mini is an Apache-2.0 Qwen3.5 mixture-of-experts model for long-horizon research, data analysis, coding, file work, and tool use. It activates about 3B of its 35.95B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the recommended Q4_K_M GGUF and F16 vision projector. An MTP-enabled build and a higher-quality Q8_0 model are available as variants.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q4-mtp
Apodex-1.1-mini with MTP speculative decoding enabled on the recommended Q4_K_M GGUF. The model carries its native MTP head, so it needs no separate draft model. The F16 vision projector supports multimodal prompts.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q8
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q4
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the UD-Q4_K_XL GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

qwen3.8-flash-next-q4
Qwen3.8-Flash-Next is Qwen's 125B-parameter, 6B-active experimental vision-language mixture-of-experts model. It targets agentic coding, reasoning, tool use, and long-context workloads with a native 262K-token context window. This default entry uses Unsloth's UD-Q4_K_XL GGUF and BF16 vision projector. Linked variants offer Q8_0 and AtomicChat's smaller IQ4_XS and Q4_K_M builds with a separate n-gram table shard.

Repository: localaiLicense: other

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-iq4
Qwen3.8 Flash Next in AtomicChat's AD-3.84bpw IQ4_XS M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-q4
Qwen3.8 Flash Next in AtomicChat's AD-4.27bpw Q4_K_M M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

ling-3.0-tiny-q4
Ling-3.0-tiny is InclusionAI's MIT-licensed hybrid reasoning MoE model with 7.9B total parameters and 1.3B active parameters per token. It targets reasoning, coding, instruction following, and agentic tasks with a native 131K-token context window. This default entry uses the Q4_K_M GGUF. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ling-3.0-tiny-q8
Ling-3.0-tiny in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

ling-3.0-flash-iq1
Ling-3.0-flash is InclusionAI's MIT-licensed hybrid reasoning MoE model with 124B total parameters and 5.5B active parameters per token. It targets coding, deep research, instruction following, and agentic workflows with a native 256K-token context window. This default entry uses the 36.5 GB AD-IQ1_M GGUF. A higher-quality 44.7 GB AD-IQ2_XS model is available as a variant.

Repository: localaiLicense: mit

ling-3.0-flash-iq2
Ling-3.0-flash in the higher-quality 44.7 GB AD-IQ2_XS GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-apex
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the APEX Balanced GGUF and BF16 vision projector. Compact APEX and MTP-enabled APEX builds are available as variants.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-apex-compact
Ornith-1.5-35B-A3B in the smaller APEX Compact GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex
Ornith-1.5-35B-A3B in the APEX Balanced GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex-compact
Ornith-1.5-35B-A3B in the APEX Compact GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q4
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

nex-n2.5-mini-q4
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q4_K_M GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

Page 1