Trust, built in.
An oracle your agent acts on has to hold up. One model can be fooled, so SECONDED makes two rival models agree, tests every check on cases they have never seen, and signs every answer so you can verify it later.
Two rival labs must agree.
Every check, in testing and in production, sends the same question to a model from OpenAI and a model from Anthropic. Neither sees the other’s answer. Only an answer both give is delivered; anything else is NOT VERIFIED, and nothing is charged. Facts SECONDED verifies itself can block a “proceed”, but they never create an answer.
Tested on cases the models have never seen.
Test cases are written by models from labs other than the two being tested, and every answer key is cross-checked by a second model that is also outside the tested pair: the authors and graders of the exam never sit it. The sets are held out, so the tested models never see a case before the run.
Results are scored by whole scenario, never by lone case, so near-identical cases never count twice. When any part of a scenario can’t be scored, the whole scenario is set aside rather than letting its easy half stand in for its hard half.
Run for real, and paid for.
Before the launch line was settled, the six checks then planned went through a paid diagnostic round, with two smaller paid phases after it. Live model calls through the same pipeline a paid check uses: not a simulation, not a replay. It did its job. It surfaced faults on our side, and they were fixed before launch.
- 1,120held-out cases in the main phase
- 2,240live, paid model calls, each case sent blind to both models
- 0cases reused from an earlier round
Cases per check in the diagnostic round
| Check | Held-out cases |
|---|---|
| Hidden Prompt Check, in final testing | 300 |
| Scam Check, earlier version | 200 |
| Full Stock Check, retired | 200 |
| Trade Check, earlier version | 150 |
| Transaction Check, retired | 150 |
| Token Check, earlier version | 120 |
Every answer is signed.
Every check returns a receipt signed by SECONDED, naming the labs whose models answered. For a check that verifies facts, the receipt signs those facts too and lists anything not checked. Your client keeps a random salt for each check, so you can prove offline which input a receipt covers, with no wallet key and no call to SECONDED.
A public safety record.
SECONDED’s numbers are public: checks run, agents using it and how often the two models agree, shown live on the Agents page as the stats service comes online. At launch, agents can opt in to a safety record of their own, private by default, and a public leaderboard.
What a check is, and isn’t.
- Agreement, not proof.
- Two rival models agreeing is strong evidence, not certainty: both can make the same mistake. A check is a second opinion before acting, not an audit, and no check can promise to catch every problem.
- Blindness is our commitment.
- We run the process so each model is blind to the other’s answer. A signed receipt proves what SECONDED issued; on its own it cannot prove that the two models answered independently.
- Figures come after a larger round.
- No per-check accuracy figure is published yet. Each will come with its margin of error, after a larger test round. Test cases and answer keys are written and cross-checked by models from independent labs, not by human auditors.
- An answer covers one moment.
- An answer covers the exact input sent, when it was checked. Whoever controls a contract can change it later, and an ERC-8004 registration is a record, not an endorsement.
- Not advice.
- SECONDED is experimental. It is not a price feed, and not financial, investment, legal or other professional advice. Its regulatory status depends on jurisdiction.
- Names are not affiliations.
- SECONDED is not affiliated with, endorsed by, or connected to OpenAI, Anthropic, Robinhood or any network it names. Naming them identifies whose models answer and where SECONDED pays and reads, nothing more.