The independent trust layer for AI tools
Independent trust for the AI tools your ecosystem runs on.
Agents plug into third-party MCP servers and skills that can hijack them or leak data. polygraph.so is an independent lab that grades those tools behaviorally, publishes evidence anyone can re-run, and keeps the grade current, because a tool surface doesn’t hold still. Free to read as a public index; continuous and monitored for your network:
For ecosystems & networks
A continuous trust index
We keep an independent, reproducible index of your network’s MCP servers, agents, and skills, re-graded on a cadence, with drops and swapped tool surfaces flagged to your team.
Free & public
The public trust index
Every grade we publish is free to read and reproducible. Browse behavioral grades for MCP servers and skills across networks, and see exactly what an index looks like before you run one for yours.
Building the tools yourself? The open litmus harness (CLI, GitHub Action, and README badges) is self-serve on the builders page →
Live indexes, read from real evidence.
225 live grades · 6 networks
Bankr ecosystem
monitoredStatic safety grades for the Bankr skill library, plus behavioral grades for the agents that ship their own MCP server.
Base network
MCP servers an onchain agent on Base can use: the projects integrating Base MCP, plus the wider DeFi, data, and infrastructure servers around the network.
Virtuals Protocol
Behavioral grades for the MCP surface of the Virtuals agent launchpad on Base: the protocol's own commerce and framework infrastructure, and the agents launched on it.
Your network
Not listed yet?
Start a trust index for your ecosystem’s servers, agents, and skills: graded, then re-graded on a cadence.
Start monitoringContinuous, and impossible to buy.
Continuous
A grade is a measurement, and measurements have a date. The tool it describes won’t hold still.
Independent
Nobody can pay for a grade. No graded party gets review or approval rights.
Re-grade, detect, alert.
01 · Re-grade
The whole index re-run on a schedule, the same behavioral test, repeated, not a one-time snapshot.
02 · Detect
Each run compared to the last. A drop, a new failing probe, or a fingerprint mismatch is flagged against what graded before.
03 · Alert
Your team hears about a regression before your users do, with the evidence bundle attached.
How polygraph gets funded.
The Bankr community launched $POLYGRAPH; we didn’t issue it. We claim the dev fees publicly and use them to fund the work: the harness, the grades, and the evidence stay free to read.
Nobody can pay for a grade. No graded party gets review or approval rights over their result.
Not financial advice. The token funds the work; it doesn’t move a grade.
Put an independent trust layer under your ecosystem.
Tell us what your network ships. We'll stand up a continuous, independent trust index, graded, then re-graded on a cadence, and keep it current as you add to it.
Prefer to write us directly? hello@polygraph.so