Why not a leaderboard, a playground, or an eval platform?
Leaderboards give you an aggregate score you can't verify. Playgrounds give you a chat you can't cite. Eval platforms make you build the pipeline yourself. Nobody publishes a specific, dated, fully reproducible test of one model on one realistic task — exact prompt, exact parameters, raw output, at a permanent URL.
Swipe to see the full comparison →
| Category | Can you verify it? | Permanent & citable? | Task-specific? |
|---|---|---|---|
| Leaderboards | No — aggregate scores over academic benchmarks | The ranking itself changes weekly | No — one number, not your task |
| Playgrounds | N/A — the chat is ephemeral | No — closes when you close the tab | Sometimes, but not shareable or re-checkable |
| Eval platforms | Yes, once you build the pipeline yourself | Depends entirely on your own setup | Yes, if you build it |
| LLM Ground | Yes — the raw, unedited output is published | Yes — every run is a permalink, kept forever | Yes — every probe is one specific, realistic task |
What a proving-ground result looks like
A probe is one realistic task; a run is that probe against one model on one date. Every box below shows a piece of what a published run page contains — with mock content, clearly marked, since Phase 1 hasn't run yet.
Why should you believe a number on this site?
Because you don't have to take it on faith. Every claim here traces back to a raw output, a published rubric, or a versioned probe definition — the same discipline that the fabricated, unmethodologied rankings previously on this domain didn't have.
Why this gets more valuable over time
Every run is permanent, so the dataset compounds monthly instead of resetting. A competitor starting next year can't manufacture a baseline they didn't record — the history is the moat.
- Phase 1In progress
Curated probe library
20–30 hand-built probes across structured extraction, code, instruction-following, and consistency, run against 8–12 current models.
- Phase 2Planned
Regression tracking
The full probe suite re-runs monthly. This is the moat: a longitudinal record of model behaviour that can't be backfilled by a competitor starting later.
- Phase 3Planned
Bring your own probe
Authors add their own probes, run them with their own key, and optionally publish to the public library after review.
- Phase 4Planned
Paid tier
Private probe suites and scheduled regression runs for teams who depend on a specific model behaving consistently.
The public probe library, every historical run, and the full methodology are free — always.
Frequently asked
These are close to the exact questions people already type into ChatGPT, Claude, or Perplexity — and today those assistants have no good source to cite, because leaderboards don't answer task-specific questions and playgrounds don't publish anything. Here's the direct answer to each one.
