<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>MoE on 苏然的博客</title>
		<link>https://bml.asia/tags/moe/</link>
		<description>Recent content in MoE on 苏然的博客</description>
		<generator>Hugo</generator>
		<language>zh-CN</language>
		
		
		
		
			<lastBuildDate>Thu, 17 Sep 2026 04:19:55 +0800</lastBuildDate>
		
			<atom:link href="https://bml.asia/tags/moe/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>DeepSeek V4.1 Flash 正式版：非对称架构 CED 的技术账</title>
				<link>https://bml.asia/posts/deepseek-v41-ced-asymmetric/</link>
				<pubDate>Thu, 10 Sep 2026 21:32:05 +0800</pubDate>
				<guid>https://bml.asia/posts/deepseek-v41-ced-asymmetric/</guid>
				<description>&lt;p&gt;&lt;span style=&#34;text-indent:2em;display:block;&#34;&gt;9 月 10 日中午，内测了两天的 V4.1 Flash 正式转正并全量上线，调用时把模型名换成 deepseek-flash 即可。这个 552B 参数的 MoE 模型被官方称作新一代结构系列里的最小尺寸版本，却没有接着堆参数，而是拿出了一套全新的 Causal-Encoder-Decoder 架构，从预训练开始完全重练：45 万亿 token 的多模态语料，喂出了首个原生支持视觉理解的正式版模型。跑分直接把自家旗舰 V4 Pro 比了下去，性能、费用、速度、总用时几个维度上都宣称全面超越，9 月 14 日之后连 V4 Pro 的请求都会被重路由到 V4.1 Flash，按新单价计费。Flash 把 Pro 干掉这种事，放在半年前听着还像段子，现在成了官方公告里的原话。&lt;/span&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>DeepSeek V4.1 Flash 内测观察：把速度卷到 507 tokens/s 的技术账</title>
				<link>https://bml.asia/posts/deepseek-v41-flash-speed/</link>
				<pubDate>Tue, 08 Sep 2026 19:27:37 +0800</pubDate>
				<guid>https://bml.asia/posts/deepseek-v41-flash-speed/</guid>
				<description>&lt;p&gt;&lt;span style=&#34;text-indent:2em;display:block;&#34;&gt;9 月 8 日下午，DeepSeek 在官方交流群里悄悄放出了 V4.1 Flash 的中间版本内测，没有发布会，也没有官网更新，只是一条群通知：新模型采用了新的结构，原生支持多模态，能力更强、速度更快、成本更低。开发者只要保持 base_url 不变，把模型名换成 deepseek-v4.1-flash-expires-on-0910 就能调用，每个账号限流 20 并发，计费与 V4 Flash 完全持平，到了 9 月 10 日这个模型名会自动过期下线。消息传开后，社交平台上很快出现了实测数据：有人跑出最高 507 tokens/s 的输出速度，也有人让模型生成一段鹈鹕骑自行车的 SVG 动画，稳定在 300 tokens/s 以上。作为参照，人类正常的阅读速度大约每秒 5 到 10 个字，这个速度意味着模型吐字的效率已经远远甩开了人的阅读上限。&lt;/span&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>腾讯混元 Hy4 preview 开源了</title>
				<link>https://bml.asia/posts/tencent-hunyuan-hy4-preview/</link>
				<pubDate>Fri, 28 Aug 2026 14:31:19 +0800</pubDate>
				<guid>https://bml.asia/posts/tencent-hunyuan-hy4-preview/</guid>
				<description>&lt;p&gt;　　腾讯混元今天放出了 Hy4 preview，参数配置一看就是冲着开源第一梯队去的：总参数 770B，激活参数却只有 49B，稀疏度压到约百分之六点四——这意味着它用接近 49B 稠密模型的推理成本，撑起了 770B 的知识容量，是典型的激进 MoE 设计。上下文长度直接拉到 1M，一次性装下整个代码仓库或几百页长文档不成问题，这对代码理解、长文分析这类真实工作负载来说，比单纯堆参数更有实际意义。目前模型已开源，并在腾讯云 TokenHub、OpenRouter 上线，WorkBuddy、CodeBuddy、元宝、ima 等腾讯产品也同步首发，等于把模型能力和自家工具链绑在了一起。&lt;/p&gt;</description>
			</item>
			<item>
				<title>小红书开源 Dots3 的 AI 变局</title>
				<link>https://bml.asia/posts/xiaohongshu-dots3-kaoyuan-shuang-ping/</link>
				<pubDate>Thu, 20 Aug 2026 16:59:04 +0800</pubDate>
				<guid>https://bml.asia/posts/xiaohongshu-dots3-kaoyuan-shuang-ping/</guid>
				<description>&lt;p&gt;　　同一款模型，在海内外收获的待遇可以差到像两个世界。小红书的 Dots3-Note Preview 就是这样——它在海外热热闹闹地开源，国内却几乎听不见声响。这种温差本身就比任何参数都值得琢磨：到底是这款模型不行，还是我们根本用错了观察它的角度？&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
