<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>推理加速 on 苏然的博客</title>
		<link>https://bml.asia/tags/%E6%8E%A8%E7%90%86%E5%8A%A0%E9%80%9F/</link>
		<description>Recent content in 推理加速 on 苏然的博客</description>
		<generator>Hugo</generator>
		<language>zh-CN</language>
		
		
		
		
			<lastBuildDate>Thu, 17 Sep 2026 04:19:55 +0800</lastBuildDate>
		
			<atom:link href="https://bml.asia/tags/%E6%8E%A8%E7%90%86%E5%8A%A0%E9%80%9F/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>曙光 ParaCache 与 KV Cache 之战</title>
				<link>https://bml.asia/posts/sugon-paracache-kvcache-analysis/</link>
				<pubDate>Fri, 28 Aug 2026 15:27:10 +0800</pubDate>
				<guid>https://bml.asia/posts/sugon-paracache-kvcache-analysis/</guid>
				<description>&lt;p&gt;　　中科曙光在数博会上推出了新一代词元加速方案 ParaCache，采用全国产化技术路线，并已在首个全国产十万卡 AI 超集群曙光 8000 的推理场景完成验证。要理解这个方案的价值，得先明白它打的是哪里——靶点是 &lt;strong&gt;KV Cache（键值缓存）&lt;/strong&gt;，而要解决的头号问题，官方表述得很直白：&amp;ldquo;中间结果调度和复用不畅导致的重复计算&amp;rdquo;。配合赛迪顾问报告里的判断，KV Cache 正成为存储系统的新型负载：长上下文、多轮交互、高并发推理让它的规模持续膨胀，既占显存，又拖慢响应和集群并发。&lt;/p&gt;</description>
			</item>
			<item>
				<title>DFlash2 用投机解码给 Qwen 提速</title>
				<link>https://bml.asia/posts/dflash2-qwen3-27b-speculative-decoding/</link>
				<pubDate>Wed, 26 Aug 2026 00:04:05 +0800</pubDate>
				<guid>https://bml.asia/posts/dflash2-qwen3-27b-speculative-decoding/</guid>
				<description>&lt;p&gt;　　Qwen3.8-27B 开源短短十天，Hugging Face 下载量就破了 260 万，配套的加速生态也几乎同步成熟。其中最具讨论度的一件事，是 Inco AI 开源的 &lt;strong&gt;&lt;code&gt;Qwen3.8-27B-DFlash2&lt;/code&gt;&lt;/strong&gt; 投机解码（speculative decoding）草稿模型：它只有 &lt;strong&gt;1.92B 参数、约 3.85GB&lt;/strong&gt;，配合主模型在 SGLang 上跑，官方模型卡给出的单并发输出吞吐最高达到自回归基准的 &lt;strong&gt;3.43 倍&lt;/strong&gt;。网上随之出现“提速明显”“4–5 倍”的说法。作为长期使用大模型做服务的工程师，我的第一反应不是看峰值倍数，而是追问：这笔加速到底从哪来、在什么工况下成立、又会在哪里缩水。把投机解码的原理、DFlash2 的架构选择与实测数据叠在一起看，答案其实相当清晰，也远比标题里的倍数更有信息量。&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
