<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Quantization on Neputer Blog</title><link>https://blog.neputer.com/tags/quantization/</link><description>Recent content in Quantization on Neputer Blog</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 26 Aug 2026 09:04:08 +0545</lastBuildDate><atom:link href="https://blog.neputer.com/tags/quantization/index.xml" rel="self" type="application/rss+xml"/><item><title>4-Bit LLM Beats Full-Precision: The Compression Tradeoff Just Got Inverted</title><link>https://blog.neputer.com/news/2026-08-26-4-bit-llm-beats-full-precision-the-compression-tradeoff/</link><pubDate>Wed, 26 Aug 2026 09:04:08 +0545</pubDate><guid>https://blog.neputer.com/news/2026-08-26-4-bit-llm-beats-full-precision-the-compression-tradeoff/</guid><description>Multiverse Computing&amp;#39;s Quantization-Aware Healing lets a 4-bit, 60B model outperform its bfloat16 120B source on 7 of 9 benchmarks, upending deployment assumptions.</description></item></channel></rss>