Back To Blog

GLM-4.7 Flash Now Available on VOLT Intelligence

VOLT Team
 / Jan 24, 2026
GLM-4.7 Flash Now Available on VOLT Intelligence

GLM-4.7-Flash is now live on VOLT Intelligence. Z.ai's 30B MoE model brings GLM-4.7's coding DNA to a lighter architecture built for cost-efficient inference.

GLM-4.7-Flash delivers 91% of the flagship's coding ability at roughly 10% of the compute. It activates just 3B parameters per forward pass, which means predictable latency and dramatically lower serving costs. For high-volume inference or resource-constrained deployments, it's a tradeoff that makes a ton of sense.

Benchmarks

GLM-4.7-Flash dominates its weight class. On SWE-bench Verified, it hits 59.2%, nearly tripling Qwen3-30B-A3B-Thinking's 22.0%. Tool orchestration on τ²-Bench reaches 79.5%, outperforming models with far more parameters. Math reasoning stays sharp at 91.6% on AIME 2025.

Benchmark GLM-4.7-Flash Qwen3-30B-A3B-Thinking GPT-OSS-20B
SWE-bench Verified 59.2% 22.0% 34.0%
τ²-Bench (Tool Use) 79.5% 49.0% 47.7%
BrowseComp 42.8% 22.9% 28.3%
AIME 2025 (Math) 91.6% 85.0% 91.7%

Source: Z.ai

The tradeoff shows on harder reasoning. HLE drops to 14.4% versus GLM-4.7's 42.8%. And SWE-bench trails the flagship's 73.8%. For maximum accuracy on complex codebases, the full model is still the better choice.

Where GLM-4.7-Flash Crushes

This model is built for throughput. Coding assistants handling hundreds of concurrent users. Automated pipelines processing large volumes of code. Development environments where latency matters more than peak reasoning depth.

Z.ai kept Interleaved Thinking and improved frontend generation from the flagship. The 200K context window handles large codebases without truncation.

Running GLM-4.7-Flash on VOLT Intelligence

Even a 30B model needs GPU resources to self-host properly. Through VOLT Intelligence, GLM-4.7-Flash is accessible via a single API alongside the full GLM-4.7 and the complete model library.

Route high-volume tasks to Flash, complex reasoning to the flagship. Same endpoint, different model parameter.

Ready to run inference at 1/10th the cost?

GLM-4.7-Flash delivers 91% of flagship coding performance.