# SOTA·SELECT — source register and research limitations

Catalog revision: **2026-10-09**. This is a curated starting shortlist, **not a verified
ranking of the latest October 2026 releases**. Established version names are retained
so readers can inspect concrete documentation. Availability may change.

## What changed

The former dossier asserted exact current-generation specs, prices, benchmark positions
and purported community quotations without traceable source URLs. Those assertions
and quotes have been removed from the app and replaced with conservative descriptions,
primary-source links and explicitly scoped human reports. The old document is not a
verified source. Git history preserves it for audit purposes.

## How recommendations work

Each model has eleven editorial scores from 0 to 100: reasoning, coding, writing,
multimodal input, context, structured output, agents, languages, cost, deployment and
live retrieval. Each task has weights summing to 1. Fit is the weighted sum; the top
five models are shown with their two largest weighted contributions.

**These numbers are editorial priors, not benchmark measurements.** A score of 90 does
not mean 90% accuracy. The linked documentation and reports inform qualitative checks
and caveats, but do not independently substantiate exact scores or rank ordering.
The app is not performing live research or measuring current price/performance.

Provider features are distinct from model features. Live web access requires retrieval
tools; structured output varies by serving stack; multimodal chat support varies by
interface. A high fit score is not a guarantee that every required modality or feature
is available. Check these requirements before using a recommendation. For health, legal
or financial tasks, use qualified review and appropriate data controls.

## Reviewed human reports and comments

The following source bodies were read using the GitHub API on **2026-10-09**.
Summaries in the app are paraphrases, not invented quotations. These are firsthand
technical experiences, not broad consumer satisfaction surveys. Negative reports are
subject to selection bias and are not evidence of general inferiority.

| Source | Published / author | What it supports | Limits |
|---|---|---|---|
| [OpenAI SDK #2486](https://github.com/openai/openai-python/issues/2486) | 2025-07-22 · ejm-betterup | A report of validation failure parsing incomplete structured responses | SDK-level; not a review of each listed OpenAI version; no numerical score validation |
| [Claude Code #100130](https://github.com/anthropics/claude-code/issues/100130) | 2026-10-07 · youngsoo-spec | A Max 20x subscriber reports weekly limits interrupting work | Subscription-level, not an exact-model comparison; not an OpenRouter limit |
| [llama.cpp #11970](https://github.com/ggml-org/llama.cpp/issues/11970) | 2025-02-20 · vnicolici | Quantized DeepSeek-R1 serving and cache behavior can affect perceived latency | Older Windows setup and low-bit quantizations; not current general performance |
| [Comment on #11970](https://github.com/ggml-org/llama.cpp/issues/11970#issuecomment-2671215654) | 2025-02-20 · ggerganov | Tentative tokenization explanation | Same thread, not independent corroboration or a resolved diagnosis |

No verified human review is attached for the other profiles. The UI explicitly says so.
Reddit searches are provided for discovery only, not cited as reviewed evidence. All
profiles link to Artificial Analysis as an evaluation discovery resource; no numerical
benchmark result from that site is claimed here.

## Official reference register

These are reference links curated on 2026-10-09, **not independently re-fetched vendor
documentation in this update**. The execution environment allowed GitHub API research
but not arbitrary vendor/model-host requests. Verify current content, supported inputs,
context, pricing, licensing, retirement status and availability at the provider.

### Closed-source references (12)

| Model | License label | Primary reference |
|---|---|---|
| GPT-4.1 | Proprietary | [Official reference](https://platform.openai.com/docs/models/gpt-4.1) |
| GPT-4.1 mini | Proprietary | [Official reference](https://platform.openai.com/docs/models/gpt-4.1-mini) |
| OpenAI o3 | Proprietary | [Official reference](https://platform.openai.com/docs/models/o3) |
| OpenAI o4-mini | Proprietary | [Official reference](https://platform.openai.com/docs/models/o4-mini) |
| Claude Opus 4 | Proprietary | [Official reference](https://www.anthropic.com/news/claude-4) |
| Claude Sonnet 4 | Proprietary | [Official reference](https://www.anthropic.com/news/claude-4) |
| Claude 3.5 Haiku | Proprietary | [Official reference](https://www.anthropic.com/news/3-5-models-and-computer-use) |
| Gemini 2.5 Pro | Proprietary | [Official reference](https://ai.google.dev/gemini-api/docs/models) |
| Gemini 2.5 Flash | Proprietary | [Official reference](https://ai.google.dev/gemini-api/docs/models) |
| Grok 3 | Proprietary | [Official reference](https://docs.x.ai/docs/models) |
| Sonar Pro | Proprietary | [Official reference](https://docs.perplexity.ai/getting-started/models/models/sonar-pro) |
| Amazon Nova Pro | Proprietary | [Official reference](https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html) |

### Open-weight references (12)

| Model | License label | Primary reference |
|---|---|---|
| DeepSeek-R1 | MIT | [Official reference](https://github.com/deepseek-ai/DeepSeek-R1) |
| DeepSeek-V3 | DeepSeek model license | [Official reference](https://github.com/deepseek-ai/DeepSeek-V3) |
| Qwen3-235B-A22B | Apache 2.0 | [Official reference](https://github.com/QwenLM/Qwen3) |
| QwQ-32B | Apache 2.0 | [Official reference](https://huggingface.co/Qwen/QwQ-32B) |
| Llama 4 Maverick | Llama 4 community license | [Official reference](https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct) |
| Llama 3.3 70B Instruct | Llama 3.3 community license | [Official reference](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct) |
| Mistral Small 3.1 24B | Apache 2.0 | [Official reference](https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503) |
| Gemma 3 27B IT | Gemma terms | [Official reference](https://huggingface.co/google/gemma-3-27b-it) |
| gpt-oss-120b | Apache 2.0 | [Official reference](https://github.com/openai/gpt-oss) |
| Kimi K2 Instruct | Modified MIT | [Official reference](https://github.com/MoonshotAI/Kimi-K2) |
| GLM-4.5 | MIT | [Official reference](https://github.com/zai-org/GLM-4.5) |
| OLMo 2 32B Instruct | Apache 2.0 | [Official reference](https://huggingface.co/allenai/OLMo-2-0325-32B-Instruct) |

## Chat access (24 model entries)

Each model has an OpenRouter browser-chat URL with a model identifier in its `models`
query parameter. This is one shared chatbot platform exposing multiple LLMs, not 24
independent apps. An account and credits may be required, and third-party hosting has
its own privacy terms. The directory is not a live availability monitor. Confirm the
selected model and active providers; older versions may become unavailable. No claim
is made that each vendor’s consumer app exposes these exact checkpoints.

## Updating the catalog

Model records, numerical priors and URLs live in `MODELS` in `index.html`. Task vectors
live in `TASKS`; additional everyday tasks inherit a documented parent vector via
`EXTRA_TASKS`. Reviewed human sources live in `HUMAN_SOURCES`, including model/family
scope, author, publication date, summary, limitations and a permalink. Add independent
reviews only after reading them; do not count multiple comments on one issue as
independent corroboration. Run the catalog tests after editing data.
