Our Methodology
How Rival generates and ranks AI outputs. Reproducible, and never cherry-picked.
No Cherry-Picking
One shot per model, no re-rolls
Real Votes Only
No synthetic benchmarks
Open Data
Datasets publicly available
The Rival Score
0 to 1000, built from real human preference. No LLM grading another LLM. Three parts:
- Win Rate
- Share of duels users voted for this model. Real humans, real chaos.
- Consistency
- How evenly a model performs across challenges. Wins big, faceplants big? Lower score.
- Credibility
- Ramps 0 to 1 as duels increase. Three lucky votes won't claim the throne.
How We Generate Comparisons
Every comparison runs the same automated pipeline. Nobody curates it.
- Same prompt, every model. Each model in a challenge receives the exact same input. No special treatment, no warm-up round.
- Automated pipeline. Responses are generated via OpenRouter, Replicate, and direct provider APIs via scripts. No human selection bias.
- Controlled parameters. Pre-generated responses use
temperature=0.7and effectively unlimited output tokens. GPT-5.1 runs at provider default. Every response ships a reproducibility card, so you can call us out. - No system prompts for pre-generated content. User prompt only. What the model sees is what you see.
- One shot, no cherry-picking. One attempt per model, never re-rolled. What the API coughed up is what ships.
API Parameter Reference
Anything unset inherits the OpenRouter default. There are no hidden knobs.
Pre-generated Showcase Responses
Static responses behind model, challenge and compare pages, via the OpenRouter Chat Completions API.
| Parameter | Value | Note |
|---|---|---|
| temperature | 0.7 | GPT-5.x series: omitted (provider default applies) |
| max_tokens | 100,000 | Effectively unlimited. Auto-retries with lower cap if model context exceeded |
| top_p | 1.0 | Explicitly set. No nucleus sampling truncation |
| top_k | 0 | Explicitly set. Disabled (full vocabulary) |
| frequency_penalty | 0.0 | Explicitly set. No frequency-based suppression |
| presence_penalty | 0.0 | Explicitly set. No presence-based suppression |
| repetition_penalty | 1.0 | Explicitly set. No repetition penalty |
| min_p | 0.0 | Explicitly set. No minimum probability cutoff |
| top_a | 0.0 | Explicitly set. No top-a filtering |
| seed | none | Non-deterministic generation |
| system_prompt | none | Raw user prompt only. No hidden instructions |
| messages | 1 user message | Single-turn. No conversation history |
Parameters not listed (logit_bias, logprobs, response_format, tools, stop) are never sent. OpenRouter parameter docs
Challenge Design
Challenges are the prompts behind every comparison. They exist to find the gaps.
- Capability-specific. Each targets one skill: coding, reasoning, writing, analysis, instruction following, visual output.
- Model-agnostic prompts. No provider-specific keywords, no prompt tricks that advantage one model family.
- High quality ceiling. Open-ended by design. A model ten years from now should still beat today's.
Ranking Methodology
Rankings on the arena come from real votes, weighted. No synthetic benchmarks, and no model grading another.
- Challenge win rates. The primary signal. Each duel win adds to that challenge category's score.
- Category coverage. Models rank inside each category. A stable rank needs participation. No ducking the hard ones.
- Consistency matters. Steady across many beats one standout. A viral moment isn't a career.
- One vote per user per duel. Anonymous and deduplicated. No ballot stuffing.
- Blind voting available. Blind mode hides names until after the vote. People vote differently when they can't see the logo.
Vote Integrity
Community votes power every ranking, so we are unpleasant about protecting them.
- Deduplication. One vote per model pair per challenge, on a unique DB constraint. The database said no.
- Rate limiting. Signed in: 200 votes/hour. Anonymous: 30, keyed by IP hash. Touch grass between votes.
- Voter fingerprinting. Anonymous votes tie to a hashed IP. Client IDs are hashed too. We trust you. Just not that much.
- Minimum sample size. 10 duels minimum before the leaderboard. Credibility down-weights small counts further.
- Server-side normalization. IDs are validated on insert. Malformed data gets rejected. We've seen what the internet sends.
Category Arenas
Models also compete in 15 category arenas, each scoped to one capability. Good at code is not good at everything.
- Website
- SVG
- Creative Writing
- Code Generation
- Reasoning
- Analysis
- Instruction Following
- Data Processing
- Multilingual
- Math
- Image Generation
- Audio
- Summarization
- Conversation
- Research
Same Rival Score formula, scoped to that category's challenges. Explore them on the arena
7-Day Trend Tracking
Snapshotted daily, so the board shows a 7-day trend. Some are quietly losing ground.
- Rising
- Score improved over 7 days
- Falling
- Score declined over 7 days
- Stable
- Score unchanged
- New
- Less than 7 days of data
Data Freshness
We track versions, so nobody gets credit for last quarter's weights.
- Point-in-time captures. Recorded when the challenge is created, model version stored alongside. Receipts kept.
- Version tracking. New challenges use the latest version. Old responses are never overwritten.
- Historical preservation. Older outputs stay. Compare then against now: growth arc or regression arc.
- Pricing data. Model pricing on the prices page comes from LiteLLM, updated to current provider rates. Brace your wallet.
What We Don't Do
Four things we get asked for and keep saying no to.
- No automated benchmarks. No MMLU, no HumanEval. Real people judge, not machines grading themselves.
- No cherry-picking. Outputs ship exactly as the API returns them. If a model fumbles, the fumble ships.
- No pay-for-ranking. Providers can't buy a rank. The one ad slot is sold by email and it never touches a score.
- No synthetic evaluations. Scores come from real votes. We don't let AI grade AI.
Model Coverage
OpenAI, Anthropic, Google, DeepSeek, xAI and everyone smaller. If it ships an API, we are already yelling at it.
- Broad coverage. New models go in as soon as they ship a public API. Full roster on the models page.
- Coverage varies. New models have fewer challenges. Rankings firm up as votes come in.
Data Transparency
Aggregated, anonymized response and voting data is public. Take it and prove us wrong.
Rival Datasets
Responses, votes and challenge metadata in JSONL. Real preference data, free.
Reproducibility
Terminal icon on any response: the exact params, plus a command you can paste. Check our work.
- One-click reproduction. Copy a cURL, Python or raw JSON payload that mirrors the call. No detective work.
- Full parameter transparency. Temperature, system prompt, max tokens, model ID and provider, on every response.
- OpenRouter as common layer. All text models are called through OpenRouter, one API surface to reproduce against, whatever the original provider.
Questions about our methodology? Open an issue on GitHub. See also our Privacy Policy and Terms of Service.
Frequently asked questions
How is the Rival Score calculated?
A single number from 0 to 1000: win rate times consistency times a credibility factor, scaled by 1000. Win rate is the share of head-to-head duels won. Consistency measures how evenly a model performs across challenges. Credibility ramps from 0 to 1 as duel count grows, so three lucky votes can't top the board.
What parameters does Rival use to generate model responses?
Pre-generated showcase responses use temperature 0.7, top_p 1.0, max_tokens 100,000, and zero frequency, presence, and repetition penalties, sent through the OpenRouter Chat Completions API with no system prompt and a single user message. The GPT-5.x series omits temperature and runs at the provider default. Every response ships with a reproducibility card listing the exact parameters.
Does Rival use an LLM as a judge?
No. Rankings come entirely from real, anonymous human votes in head-to-head duels, deduplicated to one vote per user per duel. Blind mode hides model names until after the vote. No model grades another model here, and no synthetic benchmarks.
Does Rival re-roll outputs to make a model look better?
No. Each model gets one generation attempt per prompt, never re-rolled for a nicer result. Every model in a challenge gets the exact same input, so what you see is the first response the API returned.