Model Gallery

187 models from 1 repositories

Filter by type:

Filter by tags:

ornith-1.5-35b-a3b-apex
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the APEX Balanced GGUF and BF16 vision projector. Compact APEX and MTP-enabled APEX builds are available as variants.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-apex-compact
Ornith-1.5-35B-A3B in the smaller APEX Compact GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex
Ornith-1.5-35B-A3B in the APEX Balanced GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex-compact
Ornith-1.5-35B-A3B in the APEX Compact GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q4
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4
Tiel-Coder-35B-A3B is a 35B-parameter mixture-of-experts model for coding, reasoning, tool use, and vision tasks. This default entry uses the Q4_K_XL GGUF and BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4-mtp
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q5-mtp
Tiel-Coder-35B-A3B in Q5_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q6-mtp
Tiel-Coder-35B-A3B in Q6_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8-mtp
Tiel-Coder-35B-A3B in Q8_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8
Tiel-Coder-35B-A3B in the higher-quality Q8_K_XL GGUF format, with the BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

nemotron-3.5-lightning-30b-a3b-q4
NVIDIA Nemotron 3.5 Lightning is a text-only hybrid Mamba-2, attention, and mixture-of-experts model with 30B total parameters and 3B active parameters. It targets reasoning, coding, tool use, multilingual chat, and long-context agent workflows, with a context window of up to one million tokens. This entry uses the official Q4_K_M GGUF. Automatic variant selection can choose the smaller NVFP4 build or the higher-quality Q8_0 build when it fits.

Repository: localaiLicense: openmdw-1.1

nemotron-3.5-lightning-30b-a3b-nvfp4
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official NVFP4 GGUF format. This is the smallest linked build and retains the model's reasoning, coding, tool-use, multilingual, and long-context capabilities.

Repository: localaiLicense: openmdw-1.1

nemotron-3.5-lightning-30b-a3b-q8
NVIDIA Nemotron 3.5 Lightning 30B-A3B in the official high-quality Q8_0 GGUF format for hosts with enough memory.

Repository: localaiLicense: openmdw-1.1

instella-moe-16b-a3b-think
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: other

instella-moe-16b-a3b-think-q8
AMD Instella-MoE-16B-A3B-Think is a reasoning and instruction-following mixture-of-experts model with 16 billion total parameters and 3 billion active parameters. It supports long-form reasoning, chat, coding, and tool use. This entry uses the near-lossless Q8_0 GGUF quantization.

Repository: localaiLicense: other

qwen3.6-35b-a3b-uncensored-genesis-hermes-v6
Qwen3.6-35B-A3B Uncensored Genesis Hermes V6 is LuffyTheFox's multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry installs the Q8_0 GGUF together with its F16 multimodal projector for llama.cpp. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior. License: Apache-2.0.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7
Qwen3.6-35B-A3B Genesis Hermes V7 is LuffyTheFox's Apache-2.0 multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry's own payload uses the model card's recommended APEX GGUF and the shared F16 multimodal projector. Automatic variant selection may instead choose Compact APEX, an MTP-enabled APEX build, or Q8_K_P based on serving features and available memory. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-apex-compact
Qwen3.6-35B-A3B Genesis Hermes V7 in the smaller APEX Compact GGUF format, with the shared F16 multimodal projector. This build preserves the model's multimodal, reasoning, coding, and agentic capabilities for hosts with less memory than the recommended full APEX build.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7-mtp-apex
Qwen3.6-35B-A3B Genesis Hermes V7 in the full APEX GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared F16 multimodal projector.

Repository: localaiLicense: apache-2.0

Page 1