Toolgrade
A sandboxed MCP tool-design test bench that executes versioned behavioral batteries, replays sanitized traces, explains selection and recovery failures, and proposes reviewable refactors without claiming universal quality or security.
The supplied research confirms a public evaluation of 36 popular MCP servers and an ecosystem report with a similar share of low usability grades. It found no product that grades tool design through behavioral execution and emits refactor proposals; adjacent products manage configurations, browse marketplaces, test whole agents, or scan security. Toolgrade supports two bounded modes. Own mode tests a server the operator controls and prepares code or schema changes for review. Vet mode evaluates an untrusted candidate without production credentials, private data, or network authority. Every verdict is tied to server version, manifest, tool surface, test battery, agent and model regime, sandbox policy, synthetic fixtures, seeds, traces, repetitions, and rubric version. Production-trace replay requires explicit authority and sanitization. A grade is comparative evidence within that regime, not universal usability, correctness, safety, security, or production readiness. A generated patch is a proposal; it cannot modify a repository, merge, deploy, or change a live server without owner review and normal gates.
The developer-platform, agent-infrastructure, integration, or product engineering team building or vetting MCP servers for production agents.
A July 2026 public evaluation and ecosystem report create a current quality window.
Snapshotting, synthetic fixtures, execution, trace analysis, scoring, and proposal generation scale through software.
One cross-reference and two inbound connections support the server-quality family.
A reproducible public grading signal, corroborating ecosystem data, a confirmed direct-product gap, and software-heavy testing support a timely developer tool.
The buyer and budget are underdefined, grades vary by agent regime, sandboxing untrusted servers is hard, trace replay raises privacy risk, and protocol or platform vendors can add similar evals.
Discussion
No comments yet — be the first to weigh in.
