48 everyday tasks. 24 curated models. Find a starting point, open a chatbot, and inspect the evidence. Compare 12 closed-source and 12 open-weight models — including permissive and custom licenses. A documented shortlist, not a live leaderboard.
Every task is a weighted blend of 11 capability axes — reasoning, coding, prose, multimodality, context length, structured output, agentic tooling, multilingualism, cost, deployability, and live data. Tap one.
The full roster — strengths, failure modes, how to access each model, and exactly what its structured-output formats look like. Use the model links or open a profile to inspect sources. Open weights are not necessarily fully open-source; check each license.
One model-specific OpenRouter chat link per profile — a shared browser interface, not 24 separate apps. Sign-in and credits may be required. Third-party hosting, model availability and supported features vary; links are not a live availability check. Confirm the selected model before sending sensitive data.
SOTA-SELECT is an editorial scoring engine, not a live leaderboard. Numerical scores and task weights are transparent estimates, not measured benchmark results. Official documentation describes capabilities; firsthand reports illustrate specific risks. Neither independently validates a fit score. Source coverage is incomplete and shown in each profile.
Each model receives editorial 0–100 estimates on 11 capability axes. They are directional priors, not percentages of accuracy. Compare the source scope and test your own prompts; unsupported exact scores should not drive procurement.
Each task is a weight vector over the same axes — "agentic coding" leans on CODE+AGENT; "deep research" leans on LIVE+CTX; "self-hosting" almost entirely on DEPLOY.
Fit score = Σ(weight × axis-score). The top five are shown with the axes that drove each pick, so every recommendation is explainable, not a black box.
Review official model cards and linked human reports. Unreviewed documentation links and community searches are labeled, not presented as evidence. Source coverage and limitations are recorded in RESEARCH.md — check before committing to production.