Model Gallery

181 models from 1 repositories

Filter by type:

Filter by tags:

nl2sh-1.5b-q4
nl2sh-1.5b is a 1.5B Qwen2.5-Coder fine-tune that converts plain-English requests into single POSIX or Bash commands. This Q4_K_M GGUF is 941 MB and is designed for fast CPU inference. Use the system prompt from the model card and review every generated command before execution. The model can produce destructive commands and cannot inspect the local filesystem.

Repository: localaiLicense: apache-2.0

ornith-1.5-35b-a3b-apex
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the APEX Balanced GGUF and BF16 vision projector. Compact APEX and MTP-enabled APEX builds are available as variants.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-apex-compact
Ornith-1.5-35B-A3B in the smaller APEX Compact GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex
Ornith-1.5-35B-A3B in the APEX Balanced GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex-compact
Ornith-1.5-35B-A3B in the APEX Compact GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q4
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4
Tiel-Coder-35B-A3B is a 35B-parameter mixture-of-experts model for coding, reasoning, tool use, and vision tasks. This default entry uses the Q4_K_XL GGUF and BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4-mtp
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q5-mtp
Tiel-Coder-35B-A3B in Q5_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q6-mtp
Tiel-Coder-35B-A3B in Q6_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8-mtp
Tiel-Coder-35B-A3B in Q8_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8
Tiel-Coder-35B-A3B in the higher-quality Q8_K_XL GGUF format, with the BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

pocket-35b
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-q3
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the balanced Q3_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-q2
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the smaller Q2_K GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-iq1
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the most compact IQ1_M GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct-q8
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-uncensored-genesis-hermes-v6
Qwen3.6-35B-A3B Uncensored Genesis Hermes V6 is LuffyTheFox's multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry installs the Q8_0 GGUF together with its F16 multimodal projector for llama.cpp. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior. License: Apache-2.0.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-genesis-hermes-v7
Qwen3.6-35B-A3B Genesis Hermes V7 is LuffyTheFox's Apache-2.0 multimodal, agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It combines Genesis tensor calibration with Hermes function-calling data while retaining the 35B mixture-of-experts architecture, roughly 3B active parameters per token, and the native 262K-token context window. This entry's own payload uses the model card's recommended APEX GGUF and the shared F16 multimodal projector. Automatic variant selection may instead choose Compact APEX, an MTP-enabled APEX build, or Q8_K_P based on serving features and available memory. The model card recommends Jinja chat templates and at least a 128K context for its thinking behavior.

Repository: localaiLicense: apache-2.0

Page 1