<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>投机解码 on 苏然的博客</title>
		<link>https://bml.asia/tags/%E6%8A%95%E6%9C%BA%E8%A7%A3%E7%A0%81/</link>
		<description>Recent content in 投机解码 on 苏然的博客</description>
		<generator>Hugo</generator>
		<language>zh-CN</language>
		
		
		
		
			<lastBuildDate>Thu, 17 Sep 2026 04:19:55 +0800</lastBuildDate>
		
			<atom:link href="https://bml.asia/tags/%E6%8A%95%E6%9C%BA%E8%A7%A3%E7%A0%81/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>DFlash2 用投机解码给 Qwen 提速</title>
				<link>https://bml.asia/posts/dflash2-qwen3-27b-speculative-decoding/</link>
				<pubDate>Wed, 26 Aug 2026 00:04:05 +0800</pubDate>
				<guid>https://bml.asia/posts/dflash2-qwen3-27b-speculative-decoding/</guid>
				<description>&lt;p&gt;　　Qwen3.8-27B 开源短短十天，Hugging Face 下载量就破了 260 万，配套的加速生态也几乎同步成熟。其中最具讨论度的一件事，是 Inco AI 开源的 &lt;strong&gt;&lt;code&gt;Qwen3.8-27B-DFlash2&lt;/code&gt;&lt;/strong&gt; 投机解码（speculative decoding）草稿模型：它只有 &lt;strong&gt;1.92B 参数、约 3.85GB&lt;/strong&gt;，配合主模型在 SGLang 上跑，官方模型卡给出的单并发输出吞吐最高达到自回归基准的 &lt;strong&gt;3.43 倍&lt;/strong&gt;。网上随之出现“提速明显”“4–5 倍”的说法。作为长期使用大模型做服务的工程师，我的第一反应不是看峰值倍数，而是追问：这笔加速到底从哪来、在什么工况下成立、又会在哪里缩水。把投机解码的原理、DFlash2 的架构选择与实测数据叠在一起看，答案其实相当清晰，也远比标题里的倍数更有信息量。&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
