Dev Radar
Support
LiveUpdated 2026-09-19 18:06 UTC

glm-5.3-flash-exl3-k2

glm-5.3-flash-exl3-k2 — We’re on a journey to advance and democratize artificial intelligence through open…

glm-5.3-flash-exl3-k2 is We’re on a journey to advance and democratize artificial intelligence through open source and open science.. It is ranked #1351 on the Dev Radar, in AI dev tools, first seen 20 days ago and shared in 3 posts (27.6k views).

Visit huggingface.co

vcruz305/GLM-5.3-Flash-EXL3-K2 · Hugging Face Hugging Face Log In Sign Up ","pad_token":"<|endoftext|>"},"chat_template_jinja":"[gMASK]<sop>\n{%- set effective_reasoning_effort = reasoning_effort if reasoning_effort is defined and reasoning_effort in ['low', 'high'] else 'max' -%}\n{%- if effective_reasoning_effort is not none -%}<|system|>Reasoning Effort: {{ effective_reasoning_effort | capitalize }}{%- endif -%}\n{%- set clear_thinking = clear_thinking if clear_thinking is defined else false -%}\n{%- if tools -%}\n{%- macro tool_to_json(tool) -%}\n {%- set ns_tool = namespace(first=true) -%}\n {{ '{' -}}\n {%- for k, v in tool.items()…

What people said about glm-5.3-flash-exl3-k2 on X

vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 → 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. 🚀 The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the…

@ViC305, 16 days ago · 128 likes · see the post

GLM-5.3-Flash EXL3 K2 packs a 320B / 18B-active model into 91.017 GiB and runs it on ONE DGX Spark. Native MTP at 64K averaged 17.29 tok/s across 4 workloads and reached 20.60 tok/s on structured output. 𝗤𝗨𝗔𝗟𝗜𝗧𝗬-𝗙𝗜𝗥𝗦𝗧 𝗞𝟮 This is not “put the entire model at 2-bit.” I applied EXL3 K2 MCG trellis…

@ViC305, 20 days ago · 61 likes · see the post

A very special THANK YOU to @plotarmordev and @MiaAI_lab for their hard work on EXL3 and pushing me in the right direction for this release! Even though I built most of the kernels from scratch, there were some extra improvements including the superGEMM work that was brought into my plugin (unbeknownst to me) by my…

@ViC305, 15 days ago · 33 likes · see the post

Alternatives to glm-5.3-flash-exl3-k2

glm-5.3-flash-exl3-k2 in numbers

FAQ

What is glm-5.3-flash-exl3-k2?

We’re on a journey to advance and democratize artificial intelligence through open source and open science. It was first shared on X 20 days ago and is ranked #1351 on the Dev Radar.

Is glm-5.3-flash-exl3-k2 free?

It is open source.

Who shared glm-5.3-flash-exl3-k2?

1 account on X, including @ViC305, in 3 posts totalling 27.6k views.

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:06 UTC. Full method.