Model Gallery

305 models from 1 repositories

Filter by type:

Filter by tags:

spark-x2.5-4b
# Spark-X2.5 [](https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw) [](https://discord.gg/kTDE2Hg8aw) [](https://www.youtube.com/@SparkLLM) [](https://dev.to/sparkllm) [](https://bsky.app/profile/sparkllm.bsky.social) [](https://x.com/sparkllm) [](https://www.zhihu.com/people/zhiikz7qh7m) [](images/xhtoken-wechat.jpg) > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. ## Introduction We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages. ...

Repository: localaiLicense: apache-2.0

deepseek-v4-flash-vision-exp
# DeepSeek-V4-Flash-Vision-Exp ## Introduction We are excited to introduce **DeepSeek-V4-Flash-Vision-Exp**, our first experimental multimodal model in the DeepSeek-V4 family. It builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities. Compared to DeepSeek-V4-Flash-0731, DeepSeek-V4-Flash-Vision-Exp achieves substantial improvements on its multimodal agent capabilities, while maintaining comparable performance on text-only agent tasks. Notes: 1. For the text agent benchmarks above, DeepSeek models are evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the `max` reasoning effort level with `temperature = 1.0, top_p = 0.95`. 2. † For ApexBench and Agents' Last Exam, DeepSeek-V4-Flash-0731 ignores the multimodal elements in the input. ## Repository layout This repository contains the tokenizer, prompt encoding reference, and a minimal PyTorch inference implementation for DeepSeek-V4 Flash Vision. The reference inference covers the vision encoder and aligner, DFlash attention, MoE, Hyper-Connections, and the DSpark forward path. ...

Repository: localaiLicense: mit

qwopus3.8-27b-flash
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

glm-5.3-flash
# GLM-5.3-Flash 👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report. 📍 Use GLM-5.3-Flash API services on Z.ai API Platform. ## Introduction We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute. ## Serve GLM-5.3-Flash Locally ...

Repository: localaiLicense: mit

apodex-1.1-mini-q4
Apodex-1.1-mini is an Apache-2.0 Qwen3.5 mixture-of-experts model for long-horizon research, data analysis, coding, file work, and tool use. It activates about 3B of its 35.95B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the recommended Q4_K_M GGUF and F16 vision projector. An MTP-enabled build and a higher-quality Q8_0 model are available as variants.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q4-mtp
Apodex-1.1-mini with MTP speculative decoding enabled on the recommended Q4_K_M GGUF. The model carries its native MTP head, so it needs no separate draft model. The F16 vision projector supports multimodal prompts.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q8
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q4
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the UD-Q4_K_XL GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

nl2sh-1.5b-q4
nl2sh-1.5b is a 1.5B Qwen2.5-Coder fine-tune that converts plain-English requests into single POSIX or Bash commands. This Q4_K_M GGUF is 941 MB and is designed for fast CPU inference. Use the system prompt from the model card and review every generated command before execution. The model can produce destructive commands and cannot inspect the local filesystem.

Repository: localaiLicense: apache-2.0

s1-mini-q4
S1-mini by Superwhisper is a 0.6B English text normalizer for raw speech transcripts. It removes fillers and false starts, restores punctuation and capitalization, and formats spoken numbers, dates, currency, and email addresses as written text. This default entry uses the publisher's 462 MB Q4_K_M GGUF and greedy decoding. A higher-fidelity F16 model is available as a variant. Prefix the transcript with the styling, structure, and context control line documented on the model page.

Repository: localaiLicense: s1-mini-license

glm-5.3
# GLM-5.3 GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks: + Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam. + Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks. ## Benchmark ### Serve GLM-5.3 Locally GLM-5.3 supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide - For deployment on the `Ascend NPU` platform, inference frameworks such as vLLM-Ascend, xLLM and SGLang are supported — see here. ### Note ...

Repository: localaiLicense: other

qwen3.8-flash-next-q4
Qwen3.8-Flash-Next is Qwen's 125B-parameter, 6B-active experimental vision-language mixture-of-experts model. It targets agentic coding, reasoning, tool use, and long-context workloads with a native 262K-token context window. This default entry uses Unsloth's UD-Q4_K_XL GGUF and BF16 vision projector. Linked variants offer Q8_0 and AtomicChat's smaller IQ4_XS and Q4_K_M builds with a separate n-gram table shard.

Repository: localaiLicense: other

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-iq4
Qwen3.8 Flash Next in AtomicChat's AD-3.84bpw IQ4_XS M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-q4
Qwen3.8 Flash Next in AtomicChat's AD-4.27bpw Q4_K_M M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

dfm-mimir:vllm
DFM Mimir is an Apache-2.0, instruction-tuned HRM-Text model from Danish Foundation Models. It has about 1 billion parameters and a 4,096-token context window. The model focuses on Danish and English chat, reasoning, mathematics, and code generation, and uses only permissible post-training data. This entry serves the official BF16 safetensors checkpoint with vLLM.

Repository: localaiLicense: apache-2.0

ling-3.0-tiny-q4
Ling-3.0-tiny is InclusionAI's MIT-licensed hybrid reasoning MoE model with 7.9B total parameters and 1.3B active parameters per token. It targets reasoning, coding, instruction following, and agentic tasks with a native 131K-token context window. This default entry uses the Q4_K_M GGUF. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: mit

ling-3.0-tiny-q8
Ling-3.0-tiny in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: mit

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

Page 1