We run Eve on the public memory benchmarks the field uses — each scored by the benchmark's own grader — and show it against the range of published competitor scores. Every number is backed by the unedited transcripts, so you can check each one yourself.
Reported memory-benchmark scores vary widely by grader and subset, so we show competitors as the range of their published numbers rather than one cherry-picked figure — and mark where Eve lands. Each card links to the raw transcripts behind Eve's score.
The whole point of a benchmark is that someone else can reproduce it. So for each one we change nothing about how the test is scored — we only swap in Eve as the memory layer behind the agent.
We run the released task episodes through each benchmark's published harness — not a reimplementation. The questions, the order, and the scoring rubric are the paper's.
A standard model answers each episode the way a normal user's agent would, with Eve connected as its memory layer — the same way you'd plug Eve into your own Claude or ChatGPT.
Final answers are scored by the benchmark's own grader at its pinned version. We don't grade ourselves — the rubric that ranks everyone else ranks us.
Extraordinary numbers deserve scrutiny — so every episode is recorded in full. The harness calls Eve the same way any agent does, through our public connector, and logs exactly what went in and what came back: every ingested turn, every memory Eve formed, every question, Eve's answer, and the grade the paper's own grader gave it.
Nothing summarized, nothing cherry-picked — read the full trace of every episode →. Independent verification with the benchmarks' authors is in progress.