阿里巴巴千问团队 @Alibaba_Qwen…
This is a dev post classified by Jev as Other (a launch), kept by the Dev Radar because it carries real work, not commentary.
千问音频全家桶一夜上新!五大模型齐发组完整栈 阿里巴巴千问团队 @Alibaba_Qwen 正式发布了Qwen-Audio-3.1系列语音大模型。这不是小升级,而是一次“全家桶式”更新:原来的ASR(语音识别)、TTS(语音合成)和Realtime(实时对话)三大核心模型全面进化,同时一口气新推出了两个下一代模型——专门做音频理解的ASR-Next和专门做音频创作的TTS-Next。 现在这五款模型一起上线,直接组成了一套从“听懂→生成→实时互动→创作”的完整音频能力栈。官方还同步宣布全线大幅降价:TTS大约便宜70%,Realtime便宜85%,ASR最高直接降95%,门槛一下子低了很多。 具体能力上,新版ASR不仅多语种、方言识别更准,还能自动把语气词、重复话去掉,转写成更通顺有逻辑的文字。ASR-Next则更进一步,能同时识别多个说话人、标时间戳,还能听懂情绪、环境音和机器声,做声音描述、事件定位和音频问答。 TTS支持多语言方言合成,同一音色能自然跨语言切换,还能用简单指令控制情绪、语速和风格。TTS-Next更猛,一次就能同时生成人声、音效和背景音乐,适合做有声书、播客、游戏和广告。 Realtime则实现了真正的边说边听、随时打断,像打电话一样自然,甚至能感知到你心情不好就放慢语速、更共情地回复。
Posted by ME News (50.1k followers) 1 h ago · 0 likes · 357 views · view the original post on X. Kept by the Dev Radar as Other. Tools mentioned: Qwen-Audio-3.1-TTS, Qwen3.8-Max-0902.
More dev work like this
- China Mobile just open-sourced a bridge between Physical AI models and humanoid robots. — @CyberRobooo
- China just dropped a Claude Opus 5 level model that runs locally. — @thesupermannx
- MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. — @_LuoFuli
- I’m building something awesome. 👀 — @LeSiOO
- The team behind @ThawDigital is building onchain capital markets infrastructure across… — @utila_io
- Mainnet is a few weeks away. — @0xMiden
- Packed particle layout using 16-bit int/fp struct fields. Borrowed my old ideas/code… — @SebAaltonen
- 🚨重磅!Qwen 正式发布 Qwen Intelligence,个人智能真正走进手机! — @NFT_Chen
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 22.6k posts from 5k X accounts over the last 21 days, 2.6k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-23 14:48 UTC. Full method.