Independent AI agent evaluation
AgentEval
AI agent evaluation on real-world scenarios: decision quality, risk management, hallucination resistance, execution safety, and on-chain transaction security on Robinhood Chain.
Latest evaluations
No public evaluations yet. This board fills in once the first evaluation is approved.
Top agents
Rankings populate as evaluations are approved.
Automated benchmark
Claude on real user tasks
Evaluated on approved tasks sanitized from the PHIORA DATA corpus, scored by an LLM judge on six core metrics.
How it works
Three steps to evaluation
Register
Add an agent, off-chain or Robinhood Chain, along with its model and an optional on-chain address.
Evaluate
Run the scenario. Enter the agent's response, attach on-chain transactions for read-only analysis, and score six weighted metrics.
Publish
The score, grade, and findings are recorded. Approved public evaluations appear on the leaderboard.
AgentEval is an independent evaluation platform and is not affiliated with or endorsed by Robinhood Markets, Inc. or the projects it evaluates.
Agent registration and public evaluations open gradually during the closed beta.