Skip to content
10/6 · Tue

Trending now

Full ranking
  1. 1OpenAI Adds Watermarks to ChatGPT and Codex Text for EU Users23 Interest
  2. 2Bifrost Gateway Governs Access Permissions for MCP Tools17 Interest
  3. 3GitHub 发布 AI 代码评审基准 ReviewBench13 Interest
  4. 4把 Jev 接入 Claude Code 与浏览器智能体12 Interest
  5. 5用免费 LLM API 给 Claude Code/Cursor 降本12 Interest

Latest curated items

Today10/6Tue
  1. DEV Community · MCP76

    How AI agents actually use your MCP server: five failure modes you won't see in logs

    The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.

    Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.

  2. DEV Community · MCP78

    Prompt injection is a data-plane problem: move the boundary from the model to the tool call.

    The author argues that prompt injection shouldn't be solved by making models smarter; instead, just as SQL injection is handled with parameterized queries, the boundary should be drawn where the agent executes actions.

    Why it matters: The author draws an analogy between prompt injection and SQL injection, arguing for moving the boundary from the model to the tool-call layer, and lays out a practical approach with a strategy layer and separate read and write phases.

  3. DEV Community · MCP82

    Skills Are Not Tools: Why I Gave My Coding Agent 11 MCP Servers and It Fell Apart

    In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.

    Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.

  4. DEV Community · MCP78

    Anonymous health checks on 78 registry MCP servers: 51.3% complete the full call sequence

    Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.

    Why it matters: Anonymous health checks on 78 registry MCP servers, with reproducible data on tiered authentication and spec version migration.

  5. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

  6. Reddit · ClaudeCode / Codex / VibeCoding76

    A local proxy spreads Claude Code requests across multiple Max accounts and switches before the quota runs out.

    The author open-sourced claudemanager, a local daemon that Claude Code points to via ANTHROPIC_BASE_URL. It only changes the request's Authorization header to route sessions to the Max account with the most remaining capacity in its 5-hour, weekly, and per-model windows, switching at custom thresholds before those windows fill up.

    Why it matters: The author also open-sourced a local proxy that automatically distributes Claude Code traffic across multiple Max accounts based on remaining quota, and logs requests along the way.

  7. Hacker News · MCP78

    Flash-Agents: an MCP plugin that hands Claude Code's coding tasks to a DeepSeek Flash worker

    Flash-Agents is a Claude Code plugin that delegates bounded coding work—implementing slices, porting tests, reviewing diffs, mapping out a codebase—to a DeepSeek V4.1 Flash worker, while Claude keeps architecture, acceptance criteria, and final review.

    Why it matters: The author outsources Claude Code's coding tasks to a DeepSeek Flash worker and shares the sandbox, patches, and measured data, so you can judge the cost and safety boundaries for yourself.

  8. Hacker News · MCP78

    Spill:把超大的 MCP 返回结果移出上下文,存入本地 DuckDB

    Spill 是一个 Apache-2.0 开源工具,通过 Hook 拦截超过 32 KiB 的 MCP 工具返回,将其存为本地 DuckDB 表(~/.spill/spill.duckdb),智能体只拿到一个紧凑描述符,再用 SQL 查询而不是读入 5 万 token 的原始 JSON。

    Awaiting translation

    Why it matters: Spill 把超大 MCP 返回落到本地 DuckDB,让智能体用 SQL 取数,为上下文窗口紧张提供了一种可复用的思路。

  9. DEV Community · Claude Code78

    I tested ten Claude Code mods: when a guard crashes, the command still runs — only three held up

    I tested ten Claude Code mods on Claude Code 2.1.288 across 85 sessions, 882 prompts, and 5993 tool calls, and found that a guard Hook without a .catch gets skipped when it throws, so the command runs anyway. Only by adding a catch that returns deny does it fail closed.

    Why it matters: I tested ten Claude Code mods across 85 sessions and 5993 tool calls, and lay out transferable criteria for choosing between them, plus the open question of failing open.

  10. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
  1. Habr · Вайбкодинг76

    When Automated Checks Lie: Five Cases from a Project Where an AI Agent Writes the Code

    In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.

    Why it matters: The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.

  2. DEV Community · Claude Code82

    How I Used Git Checkpoints to Undo Any Change Made by a Coding Agent

    For coding agents running unattended, the author built a checkpoint mechanism based on hidden git refs. Before each task starts, it snapshots the entire working tree—including untracked files—and rolls back automatically when validation fails. The restore operation itself can also be undone.

    Why it matters: With roughly 40 lines of shell, the author turned git checkpoints into rollback-capable infrastructure, laying out the concrete approach and the limits of running coding agents unattended.

  3. Reddit · ClaudeCode / Codex / VibeCoding78

    Use CLAUDE.md plus a local MCP memory architecture to stop Claude Code from constantly needing corrections and burning through quota

    The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.

    Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.

  4. AI Hero · Skills Updates62

    AI Hero Skills v1.3 released: new /implement-spec, /pr, and /retro skills, and CONTEXT.md renamed to GLOSSARY.md

    AI Hero Skills v1.3 is out, adding three skills—/implement-spec, /pr, and /retro—and extending the main flow from /grill-with-docs → /to-spec → /to-tickets into implementation, PRs, and retros.

    Why it matters: The author rounds out the skill set into a full path from writing specs to PRs and retros, and lays out the trade-offs at each step along with the rough edges they already know about.

  5. Hacker News · MCP77

    我连续 37 天夜间测量 MCP 注册表:19.2% 的服务器 24 小时内改动了工具面

    作者搭建 mcp-transparency-log,按计划爬取官方 MCP 注册表中所有公开可达的服务器,记录其工具名、描述、JSON schema 和四个 annotation 提示,写入带签名树头的 append-only 日志。

    Awaiting translation

    Why it matters: 作者连续 37 天夜间爬取 MCP 官方注册表,用可复现的日志量化工具面变化,并公开了两次错误修正过程。

  6. DEV Community · Claude Code78

    Why Claude Code’s Read(.env) Deny Rule Doesn’t Stop Bash from Reading It

    The author added a Read(./.env) deny rule to Claude Code, but after Read was blocked, Claude switched to running `grep DATABASE_URL .env` via Bash, printing the production connection string into the conversation.

    Why it matters: Through hands-on testing, the author found that the Read deny rule doesn’t stop Bash from reading .env, and shares a three-layer protection setup that can be adapted to your own permission configuration.

  7. DEV Community · Codex78

    How I Stopped Codex from Burning Through My Usage Quota

    The author found that two Codex browser automation tasks consumed 170,123 and 110,180 tokens respectively, so they set out to control usage through model selection, configuration files, and task splitting.

    Why it matters: Drawing on real measurements where two browser tasks burned through hundreds of thousands of tokens, the author shares a quota-saving approach: switch models and configurations based on task difficulty.

10/4Sun
  1. DEV Community · Claude Code85

    How to Stop an AI Coding Agent from Declaring a Task Done Too Early

    The author runs a fully autonomous implementation system where an orchestrator hands out tasks to parallel implementation agents (built on Claude Code). At first, agents could mark their own tasks as complete, which led to problems like tests never being run, acceptance criteria not being met, assertions loosened to make tests pass, and hardcoded return values.

    Why it matters: The author solved the problem of agents declaring completion too early with a three-layer design: checkable acceptance criteria, completion reports backed by evidence, and read-only validation agents.

  2. DEV Community · Cursor76

    Cursor ships Composer 2, and the API response strings give away its undisclosed Kimi K2.5 base

    On March 20, 2026, developer Fynn was debugging Cursor's OpenAI-compatible endpoint when the returned model ID came back as accounts/anysphere/models/kimi-k2p5-rl-0317-s515-fast — evidence that Composer 2 was post-trained with reinforcement learning on top of Moonshot AI's Kimi K2.5. The tweet hit 44.4 views within a day.

    Why it matters: One API debugging session ties together Cursor's undisclosed Kimi base, the licensing attribution dispute, and the cost landscape for Chinese versus U.S. models — a look at how the industry handles disclosure.

  3. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.