🤖 What if an AI agent could keep seeing, listening, talking, and working — all at the…
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
🤖 What if an AI agent could keep seeing, listening, talking, and working — all at the same time? Developer @speechjsp built Gander, a multimodal duplex interaction agent that combines continuous audio-visual perception, real-time conversation, and asynchronous agent execution. Gander’s Cerebellum is fine-tuned from MiniCPM-o 4.5, with additional training and system-level improvements for native duplex interaction. Instead of treating voice interaction and agent execution as separate steps, Gander brings them together in one continuous interaction loop.
Posted by OpenBMB (10.8k followers) 3 days ago · 64 likes · 3k views · view the original post on X. Kept by the Dev Radar as AI dev tools.
More dev work like this
- Nimble is able to process images! — @madiator
- Your agent already solved this. Finding the session is the problem. — @DanKornas
- Trains your own large language model from scratch using plain PyTorch. — @tom_doerr
- gpt 5.6 terra = gpt 6 astra — @notjazii
- your rtx 3060 was running bonsai 2 at 26 tok/s this morning and now does 40 tok/s, i… — @sudoingX
- Turning a novel into a game is more than turning chapters into dialogue. — @DanKornas
- Meet DeepSeek-V4.1-Flash: a blazing fast image-text-to-text model. It takes both images… — @HuggingModels
- How to use Jev, and where it gives you the most advantage: — @0x_rody
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 12.2k posts from 4.7k X accounts over the last 21 days, 1.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 18:06 UTC. Full method.