Cerebras deploys Qwen 3.8 27B model with 1500 tokens/sec inference
Cerebras now serves Qwen 3.8 27B—27 billion parameters—with public endpoints achieving approximately 1500 tokens per second inference speed. The model offers 64k context in its free tier and 128k in paid tier, though all versions remain unpruned original models. Cerebras uses selective weight-only quantization for storage while preserving full precision during operations. This speed boost could help low-cost AI access, but end users face no guaranteed cost reductions since pruning techniques are unavailable on public endpoints. The free tier’s 64k context limit also constrains basic use cases. Source is brief and the detail sits with them—this signal doesn’t resolve whether speed translates to real affordability for users.
Source: Hacker News
MANY MINDED