4-Bit LLM Beats Full-Precision: The Compression Tradeoff Just Got Inverted
A new technique called Quantization-Aware Healing (QAH) compresses GPT-OSS 120B to 60B at 4 bits while beating the original on reasoning and math benchmarks — a breakthrough for efficient model deployment.