Model Gallery

596 models from 1 repositories

Filter by type:

Filter by tags:

spark-x2.5-4b
# Spark-X2.5 [](https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw) [](https://discord.gg/kTDE2Hg8aw) [](https://www.youtube.com/@SparkLLM) [](https://dev.to/sparkllm) [](https://bsky.app/profile/sparkllm.bsky.social) [](https://x.com/sparkllm) [](https://www.zhihu.com/people/zhiikz7qh7m) [](images/xhtoken-wechat.jpg) > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. ## Introduction We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages. ...

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

supra2-100m-instruct
Supra2-100M-Instruct is a compact English chat model trained from scratch by SupraLabs on the Qwen3 architecture. It has 100 million parameters, a 2,048-token context window, and is intended for lightweight experiments and constrained edge deployments. This entry uses the publisher's official F16 GGUF build.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

dfm-mimir:vllm
DFM Mimir is an Apache-2.0, instruction-tuned HRM-Text model from Danish Foundation Models. It has about 1 billion parameters and a 4,096-token context window. The model focuses on Danish and English chat, reasoning, mathematics, and code generation, and uses only permissible post-training data. This entry serves the official BF16 safetensors checkpoint with vLLM.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q4
IBM Granite 4.2 8B is a multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q4
IBM Granite 4.2 30B is the family's flagship multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q8
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q4
Dirk is a Qwen3.8 27B vision-language model with a concise chat template for agentic coding, reasoning, tool use, and general knowledge tasks. It preserves the model's MTP head for speculative decoding and supports a 262K-token context window. This default entry uses the Q4_K_XL GGUF and F16 vision projector. A choice of Q5_K_XL, Q6_K_XL, and Q8_K_XL builds is available through variants.

Repository: localaiLicense: apache-2.0

nex-n2.5-mini-q4
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q4_K_M GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

nex-n2.5-mini-q5
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q5_K_M GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

nex-n2.5-mini-q6
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q6_K GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

nex-n2.5-mini-q8
Nex-N2.5-mini is a 35B-parameter Qwen3.5 MoE model for coding, tool use, and computer and browser tasks with image input. This Q8_0 GGUF build includes the F16 vision projector and uses the embedded chat template.

Repository: localaiLicense: apache-2.0

spark-x2.5-1.7b-q4
Spark-X2.5-1.7B is XHToken's 1.7B text model for conversation, reasoning, coding, and multilingual tasks. This build uses Q4_K_M GGUF weights, the embedded Jinja chat template, and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-1.7b-q8
Spark-X2.5-1.7B is XHToken's 1.7B text model for conversation, reasoning, coding, and multilingual tasks. This build uses Q8_0 GGUF weights, the embedded Jinja chat template, and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-4b-q4
Spark-X2.5-4B is XHToken's 4B text model for conversation, reasoning, coding, and multilingual tasks. This entry uses Q4_K_M GGUF weights; Q6_K and Q8_0 builds are available as variants. All builds use the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-4b-q6
Spark-X2.5-4B in Q6_K GGUF format, with the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

Page 1