<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>AI 基础设施 on 苏然的博客</title>
		<link>https://bml.asia/tags/ai-%E5%9F%BA%E7%A1%80%E8%AE%BE%E6%96%BD/</link>
		<description>Recent content in AI 基础设施 on 苏然的博客</description>
		<generator>Hugo</generator>
		<language>zh-CN</language>
		
		
		
		
			<lastBuildDate>Thu, 17 Sep 2026 04:19:55 +0800</lastBuildDate>
		
			<atom:link href="https://bml.asia/tags/ai-%E5%9F%BA%E7%A1%80%E8%AE%BE%E6%96%BD/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>曙光 ParaCache 与 KV Cache 之战</title>
				<link>https://bml.asia/posts/sugon-paracache-kvcache-analysis/</link>
				<pubDate>Fri, 28 Aug 2026 15:27:10 +0800</pubDate>
				<guid>https://bml.asia/posts/sugon-paracache-kvcache-analysis/</guid>
				<description>&lt;p&gt;　　中科曙光在数博会上推出了新一代词元加速方案 ParaCache，采用全国产化技术路线，并已在首个全国产十万卡 AI 超集群曙光 8000 的推理场景完成验证。要理解这个方案的价值，得先明白它打的是哪里——靶点是 &lt;strong&gt;KV Cache（键值缓存）&lt;/strong&gt;，而要解决的头号问题，官方表述得很直白：&amp;ldquo;中间结果调度和复用不畅导致的重复计算&amp;rdquo;。配合赛迪顾问报告里的判断，KV Cache 正成为存储系统的新型负载：长上下文、多轮交互、高并发推理让它的规模持续膨胀，既占显存，又拖慢响应和集群并发。&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
