weightsapi.INFERENCEConsole
Navigation
Model catalog

Model catalog

Choose the right balance of capability, context and cost.

No KYC

No identity-document upload or KYC step in the current sign-in flow. Access and usage remain linked to your account.

No added content filters

WeightsAPI adds no content filter at the gateway. A model’s own training can still lead it to refuse a request.

15 open-weight models. Publisher releases.

Published model sources verified 2026-09-04. Check service status for current availability. API model list ↗

ModelInput / 1MOutput / 1MContextType
LLlama 3.1 8B$0.03$0.05128K*Text
DDeepSeek V4 Flash$0.09$0.18128K*Text
GGemma 3 27B$0.10$0.15128K*Text + images
MMistral Small 3$0.10$0.2032K*Text
QQwen3-Coder-Next$0.12$0.80128K*Text
HHermes 4 70B$0.13$0.40128K*Text
MMistral Small 4$0.15$0.60128K*Text + image
QQwen3 32B$0.15$0.20128K*Text
DDolphin Mistral 24B Venice Edition$0.20$0.90128K*Text + image
QQwen3 235B A22B$0.20$0.60128K*Text
DDeepSeek V3$0.25$0.85128K*Text
LLlama 3.3 70B$0.35$0.40128K*Text
DDeepSeek R1$0.50$2.15128K*Text
MMixtral 8×22B$0.60$0.6064K*Text
LLlama 3.1 405B$1.20$1.20128K*Text

Choose by workload

Llama 3.1 8B

Fast classification, extraction and lightweight chat.

$0.03 / $0.05Llama 3.1 Community

DeepSeek V4 Flash

A text model for long-context reasoning, coding and tool workflows, with a configured 1,048,576-token window.

$0.09 / $0.18MIT

Gemma 3 27B

Image understanding and multilingual conversation.

$0.10 / $0.15Gemma Terms of Use

Mistral Small 3

Efficient text tasks, function calling and European-language chat.

$0.10 / $0.20Apache 2.0

Qwen3-Coder-Next

A text-only coding model for codebase exploration and tool workflows, with a published 262,144-token native context.

$0.12 / $0.80Apache 2.0

Hermes 4 70B

Creative conversation, roleplay and reasoning with configurable instructions, published by Nous Research.

$0.13 / $0.40Meta Llama Community

Mistral Small 4

An official Mistral model combining text and image input, tool calling and configurable reasoning, with a published 262,144-token window.

$0.15 / $0.60Apache 2.0

Qwen3 32B

A balanced choice for code, multilingual chat and reasoning.

$0.15 / $0.20Apache 2.0

Dolphin Mistral 24B Venice Edition

Conversations and characters with application-controlled instructions, from Dolphin and Venice.

$0.20 / $0.90Apache 2.0

Qwen3 235B A22B

Strong general-purpose instruction following and tool use.

$0.20 / $0.60Apache 2.0

DeepSeek V3

Code generation, complex instructions and structured tasks.

$0.25 / $0.85MIT

Llama 3.3 70B

General-purpose assistants, summarization and tool use.

$0.35 / $0.40Llama 3.3 Community

DeepSeek R1

Difficult reasoning, mathematics and code analysis.

$0.50 / $2.15MIT

Mixtral 8×22B

Multilingual generation and function-calling workflows.

$0.60 / $0.60Apache 2.0

Llama 3.1 405B

Demanding general-purpose generation and complex instructions.

$1.20 / $1.20Llama 3.1 Community