Skip to content
#

qwen3-8-flash-next

Here are 20 public repositories matching this topic...

The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

  • Updated Oct 3, 2026
  • Python

Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode. New: opt-in uncensored mode (runtime abliteration, no new weights).

  • Updated Oct 2, 2026
  • Python

The best model under 200B at near 10 token/s exact batch throughput on a single laptop CPU, with an automatic 8 GB RAM path. Native C, no GPU or Python. | 200B 以下最强模型,单颗笔记本 CPU 精确批处理接近 10 token/s,最低 8 GB 内存自动运行。原生 C 语言,无需 GPU 或 Python。

  • Updated Sep 9, 2026
  • C

Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.

  • Updated Sep 30, 2026
  • Python
Qwen-Image-2.1-Uncensored

Qwen-Image-2.1-Uncensored is a local image model: qwen ai, qwen studio ai, qwen coder notes, chat qwen ai. GGUF, ComfyUI, no safety checker. Windows 10/11 zip, extract and run. Official free download. Download:➧

  • Updated Sep 27, 2026
  • C++
Qwen3.8-Omni-Flash-Uncensored-FREE

Qwen3.8-Omni-Flash - qwen 3.8 omni flash, qwen 3.8, qwen 3.8 flash and qwen ai desktop. Text, image, audio, video, 1M context, function calling. Windows 10/11 zip, extract and run. Official free download. Download:🡇

  • Updated Sep 18, 2026
  • C++

Add this topic to your repo

To associate your repository with the qwen3-8-flash-next topic, visit your repo's landing page and select "manage topics."

Learn more