网页截图比文本更适合喂给大模型检索,这个结论我第一眼是不信的。看完伯克利这个团队的 README,觉得他们有道理。
This is a dev post classified by Jev as AI dev tools (a tool drop), kept by the Dev Radar because it carries real work, not commentary.
网页截图比文本更适合喂给大模型检索,这个结论我第一眼是不信的。看完伯克利这个团队的 README,觉得他们有道理。 给模型喂网页,常规做法是先解析成文字再切块,表格、图表、版式在这一步就丢了,后面模型答不上来是因为它根本没看见。 PixelRAG 反过来,把网页、PDF 直接渲染成一张张截图,在图上做检索,答案在表格里就把那块截图原样递给模型读。 已斩获 10000+ GitHub Star! GitHub:http://github.com/StarTrail-org/PixelRAG 他们把整个维基百科 828 万个页面建好了索引挂在线上,不用配置不用密钥,可以直接调,拿一张图当查询条件也行。 还顺手出了一个 Claude Code 插件,叫 pixelbrowse,Claude 看网页时截一张图自己读,不去抓源码,图表和布局跟人看到的一样。 想给自己的文档建索引也行,Mac 的 M 系列芯片上一份 PDF 三分钟左右建完,不用显卡。
Posted by GitHubDaily (84.6k followers) 1 days ago · 8 likes · 3.1k views · view the original post on X. Kept by the Dev Radar as AI dev tools. Tools mentioned: pixelrag.
More dev work like this
- The Dot platform has processed 6.18B tokens this week, representing 34.33% of our… — @usedotai
- What is Jev, the tool buzzing across Silicon Valley lately? — @hayden090807
- PSA - Claude code: Turn off the prompt suggestions, save ~10% of your limits/spend — @akshdeeps_001
- 智谱 GLM Coding Plan 开启中秋加国庆双节畅享活动。 — @0xLogicrw
- 😱 什么?!仅 0.6B 本地开源决策模型在 Typed Decisions 上压过 Laya ! — @NFT_Chen
- Runs the 2.78-trillion-parameter Kimi K3 model on a single CPU using 8 GB of RAM without… — @tom_doerr
- CLAUDE.md 里写了一堆规则,Claude Code 照样有一半当没看见,这事很多人都碰到过。 — @GitHub_Daily
- 🤯卧槽!开源圈又杀出个比 Laya 更狠的 System-1 决策核模型 AgentJev-0.6B ! — @NFT_Chen
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 20.7k posts from 4.9k X accounts over the last 21 days, 2.4k tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-22 08:21 UTC. Full method.