CURATED · OCT 2026
Frontier model & task routing hub

Pick a task.
Get the right model.

48 everyday tasks. 24 curated models. Find a starting point, open a chatbot, and inspect the evidence. Compare 12 closed-source and 12 open-weight models — including permissive and custom licenses. A documented shortlist, not a live leaderboard.

0
MODELS PROFILED
0 / 0
CLOSED / OPEN-WEIGHT
0
TASK ROUTES
0
CAPABILITY AXES SCORED

What are you trying to do?

STEP 01 / SELECT TASK

Every task is a weighted blend of 11 capability axes — reasoning, coding, prose, multimodality, context length, structured output, agentic tooling, multilingualism, cost, deployability, and live data. Tap one.

Model explorer

STEP 02 / DIG DEEPER

The full roster — strengths, failure modes, how to access each model, and exactly what its structured-output formats look like. Use the model links or open a profile to inspect sources. Open weights are not necessarily fully open-source; check each license.

24 models you can chat with

OPEN & TRY

One model-specific OpenRouter chat link per profile — a shared browser interface, not 24 separate apps. Sign-in and credits may be required. Third-party hosting, model availability and supported features vary; links are not a live availability check. Confirm the selected model before sending sensitive data.

How routing works

METHOD

SOTA-SELECT is an editorial scoring engine, not a live leaderboard. Numerical scores and task weights are transparent estimates, not measured benchmark results. Official documentation describes capabilities; firsthand reports illustrate specific risks. Neither independently validates a fit score. Source coverage is incomplete and shown in each profile.

① Profile

Each model receives editorial 0–100 estimates on 11 capability axes. They are directional priors, not percentages of accuracy. Compare the source scope and test your own prompts; unsupported exact scores should not drive procurement.

② Weight

Each task is a weight vector over the same axes — "agentic coding" leans on CODE+AGENT; "deep research" leans on LIVE+CTX; "self-hosting" almost entirely on DEPLOY.

③ Rank

Fit score = Σ(weight × axis-score). The top five are shown with the axes that drove each pick, so every recommendation is explainable, not a black box.

④ Verify

Review official model cards and linked human reports. Unreviewed documentation links and community searches are labeled, not presented as evidence. Source coverage and limitations are recorded in RESEARCH.md — check before committing to production.