AI快讯 / AI 开源项目

turboquant-model

Optimize LLM inference with near-optimal 4-bit weight quantization and on-the-fly dequantization for lower memory use and faster matmul

原文来源AI 开源 Releases
查看官方原文

Optimize LLM inference with near-optimal 4-bit weight quantization and on-the-fly dequantization for lower memory use and faster matmul