llamaperf

Qwen3.5

Alibaba · 2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.
Tone: positive
quant:
WinterMix58 (mlx)

Post describes a new MLX quantization method (WinterMix) for Qwen3.5-122B-A10B, with two builds: WinterMix58 (82 GiB) and WinterMix48 (68 GiB). Benchmarks show perplexity improvements over existing MLX quants. The author is enthusiastic about the results and the method's advantages for agentic workflows on Apple Silicon.

Tone: positive
visionsummarization

Model based on Qwen3.5-4B. Trained on 8xH100 for 3 days. Supports Safetensors, GGUF, MLX weights. Requires as little as 4GB VRAM. Multiple quantizations available (GPTQ, W8A8, FP8, Q4, Q6). Tested with vLLM, SGLang, llama.cpp.