Модели

207 канонических LLM-моделей от всех провайдеров

Показаны модели 1–24 из 207

Mistral Large 4

Франция

Mistral AI's multimodal model for reasoning, coding, and agentic workloads, with 1.05T total and 49B active parameters. Public preview with a 1M-token context; downloadable weights and the exact license are still marked as coming soon.

Контекст
1.0M
Добавлена
окт. 2026 г.

Kolibri 1

Aleph Alpha's German-English reasoning model with 78B total and 3.46B active parameters, tool calling, and Apache 2.0 weights. Supports up to 1,048,576 tokens; the developer recommends at most 262,144 tokens for serving efficiency and complex tasks.

Контекст
1.0M
Добавлена
окт. 2026 г.

Gemini 4 Argon

Соединенные Штаты

Google's top-tier model anchoring the Gemini 4 generation, larger than the previous Pro line and built for complex workloads, coding, and cybersecurity. Can locate, validate, and patch critical software vulnerabilities autonomously and raises the output limit to 1M tokens. At launch, access is limited to selected cybersecurity organizations through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow; no public availability date has been announced.

Контекст
1.0M
Добавлена
окт. 2026 г.

Pareto 26.10 Preview

Unbiased's preview composite model for research, coding, and agentic workflows. Accepts text and images, supports tool calling, and provides a 1,048,576-token context window. Preview behavior may change without notice.

Контекст
1.0M
Добавлена
окт. 2026 г.

Ling 3.1 Flash

Китай

InclusionAI's hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters, built for coding, tool use, and long-document analysis. The currently served context window is 262,144 tokens.

Контекст
262K
Добавлена
сент. 2026 г.

GPT-6.1 Sol

Соединенные Штаты

OpenAI's upgrade to GPT-6 Sol, delivering near-Astra performance for complex work at a lower cost. Suited for agentic coding, computer use, and document-heavy professional work, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, up to 128K output tokens, and cheaper cached input than GPT-6 Sol.

Контекст
1.1M
Добавлена
сент. 2026 г.

Claude Sonnet 5.5

Соединенные Штаты

Anthropic's Sonnet-class model offering the best combination of speed and intelligence, succeeding Claude Sonnet 5 as a direct upgrade. Especially strong at building features, fixing bugs, and well-scoped everyday coding and knowledge work, with adaptive thinking, a 1M-token context window, and up to 128K output tokens.

Контекст
1.0M
Добавлена
сент. 2026 г.

Perceptron Mk1.5

Соединенные Штаты

Perceptron's embodied reasoning model for physical agents. Accepts text, image, video, and audio input and answers with text plus optional structured annotations such as points, boxes, polygons, and tracks, with a 36K-token context window.

Контекст
37K
Добавлена
сент. 2026 г.

GLM-5.3-Prime

Китай

The high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5-2x the output throughput through inference acceleration. Text input and output with reasoning, tool use, and a 1M-token context window.

Контекст
1.0M
Добавлена
сент. 2026 г.

Ember-1

Соединенные Штаты

A specialized reasoning model from Fireworks Research, built on Kimi K3 and designed to make every token go further by producing shorter reasoning traces. Supports image input, tool use, and a 1M-token context window.

Контекст
1.0M
Добавлена
сент. 2026 г.

Qwen 3.8 Max Prime

Китай

A higher-throughput variant of Qwen 3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. Accepts text, image, and video input with reasoning, tool use, and a 1M-token context window.

Контекст
1.0M
Добавлена
сент. 2026 г.

Aion 3.5 Mini

The smaller, lower-cost sibling of Aion 3.5, a multi-model roleplaying and storytelling system from AionLabs built on the GLM family of models. Supports reasoning, tool use, and a 262K-token context window.

Контекст
262K
Добавлена
сент. 2026 г.

Aion 3.5

A multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. Uses a collaborative generation process in which multiple specialized models contribute to each response, with reasoning, tool use, and a 262K-token context window.

Контекст
262K
Добавлена
сент. 2026 г.

Solar Mini 4

Республика Корея

Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K-token context window. Built for agentic use cases where response speed and cost matter, with reasoning and tool use.

Контекст
524K
Добавлена
сент. 2026 г.

MiMo-V2.6-Pro

Китай

Xiaomi's flagship foundation model, built at a scale of over 1T parameters for the most demanding workloads. Natively multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window; weights are published on Hugging Face.

Контекст
1.1M
Добавлена
сент. 2026 г.

Claude Opus 5.5

Соединенные Штаты

Anthropic's Opus-class model for long-running agentic coding and knowledge work, succeeding Claude Opus 5 at a lower price. Particularly strong at multi-step changes in large codebases and sustained autonomous tasks, with always-on adaptive thinking steered by an effort parameter, a 1M-token context window, and up to 128K output tokens. Anthropic's recommended starting point for most workloads.

Контекст
1.0M
Добавлена
сент. 2026 г.

MiMo-V2.6-Flash

Китай

Xiaomi's open-source mixture-of-experts foundation model with 309B total parameters and 15B activated per token, using a hybrid attention mechanism for efficient long-context inference. Multimodal across text, image, video, and audio input, with reasoning, tool use, and a 1M-token context window.

Контекст
1.1M
Добавлена
сент. 2026 г.

GPT-6 Sol

Соединенные Штаты

The cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. Suited for demanding professional work, agentic coding, and document-heavy tasks, with multimodal input, function calling, reasoning effort controls, a 1.05M-token context window, and up to 128K output tokens.

Контекст
1.1M
Добавлена
сент. 2026 г.

GPT-6 Luna

Соединенные Штаты

The fast, most efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol for focused, high-volume tasks. Suited for latency-sensitive workloads such as chat, classification, extraction, and lightweight agentic steps, with multimodal input, function calling, reasoning controls, a 1.05M-token context window, and up to 128K output tokens.

Контекст
1.1M
Добавлена
сент. 2026 г.

Grok 4.7

Соединенные Штаты

xAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. Particularly strong at long-running software engineering tasks and verifying its own work, with image and file input, reasoning, tool use, a 500K-token context window, and an OpenAI-compatible API.

Контекст
500K
Добавлена
сент. 2026 г.

Qwen 3.8 Omni Flash

Китай

Alibaba's omni-modal reasoning model and the first Qwen model built around agentic capabilities with native audio-video understanding. Suited for audio-video analysis and summarization and multimodal agents, with tool use and a 1M-token context window.

Контекст
1.0M
Добавлена
сент. 2026 г.

Pareto

A multimodal composite model from Unbiased built for research, coding, and agentic workflows, aiming at frontier-level performance across a broad range of general-purpose tasks. Supports image input, tool use, and a 262K-token context window.

Контекст
262K
Добавлена
сент. 2026 г.

Ternary Bonsai 2 27B

Соединенные Штаты

PrismML's open-weight 27B-parameter reasoning model derived from Qwen 3.8 27B and shrunk with ternary compression. Supports coding, mathematics, tool calling, and image understanding with a 262K-token context window.

Контекст
262K
Добавлена
сент. 2026 г.

GLM-5.3-FlashX

Китай

The high-speed variant of Z.ai's GLM-5.3-Flash, a natively multimodal model delivering inference speeds of up to 200 tokens per second. Built on the same hybrid sparse and linear attention architecture, with image and video input, reasoning, tool use, and a 1M-token context window.

Контекст
1.0M
Добавлена
сент. 2026 г.