No filter added at the gateway
No moderation pass or topic blocklist is added by us. The model’s own training can still refuse.
WeightsAPI adds no content filter of its own between your app and the model. Call Dolphin Mistral 24B Venice Edition ($0.20 input / $0.90 output per million tokens) or Hermes 4 70B ($0.13 / $0.40), two fine-tunes built to reduce refusals, through an OpenAI-compatible endpoint that SillyTavern takes as a Custom (OpenAI-compatible) source. Sign up with a username and password, then fund one USD balance in XMR, BTC, USDT, USDC, ETH, SOL, LTC or TRX with no platform deposit fee.
Why us: Hermes 4 70B at $0.13 / $0.40 is not on the public model lists of Venice, OpenRouter, NanoGPT or Chutes (checked 25 September 2026), and Dolphin costs the same here as on Venice’s own API, payable in Monero without an email.
Create your account now; check each model on /status before topping up. Dedicated GPUs: activation date confirmed before you pay.
No moderation pass or topic blocklist is added by us. The model’s own training can still refuse.
A 70B creative and roleplay fine-tune from Nous Research, per million input / output tokens.
WeightsAPI’s gateway checks your API key, its model allowlist and your balance, then passes your messages to the model you chose. It adds no content filter of its own: no moderation pass, no topic blocklist. That is the exact claim on our FAQ and trust page.
Three things still shape each answer: the model’s training (Dolphin and Hermes are tuned to reduce refusals, not to guarantee answers), your instructions (system prompt, character card, sampling settings), and the serving configuration, since the chat template must match the deployment.
No added filter does not make anything legal. You remain responsible for lawful use and third-party rights, and each model’s license applies: Apache 2.0 for Dolphin, and the Llama 3.1 terms and use conditions for Hermes 4 70B, a Llama 3.1 70B derivative.
Two catalog models are fine-tunes built for this work; three original releases cover cheaper drafts or a different voice. All run on the same key and balance, so you can switch per character or per scene.
| Model | API model ID | Type | Input / output | Context |
|---|---|---|---|---|
| Hermes 4 70B | NousResearch/Hermes-4-70B | Fine-tune; text | $0.13 / $0.40 | 128K |
| Dolphin Mistral 24B Venice Edition | dphn/Dolphin-Mistral-24B-Venice-Edition | Fine-tune; text + image | $0.20 / $0.90 | 128K |
| Llama 3.3 70B | meta-llama/Llama-3.3-70B-Instruct | Original; text | $0.35 / $0.40 | 128K |
| Qwen3 235B A22B | Qwen/Qwen3-235B-A22B-Instruct-2507 | Original; text | $0.20 / $0.60 | 128K |
| Mistral Small 3 | mistralai/Mistral-Small-24B-Instruct-2501 | Original; text | $0.10 / $0.20 | 32K |
Our pick for most roleplay and long-form fiction is Hermes 4 70B: the larger model, output at $0.40 per million tokens against $0.90 for Dolphin, and positioned by Nous Research for creative writing, roleplay and conversation mixed with reasoning. Choose Dolphin for image input or its character-first tuning, built with Venice. Compare them in Dolphin vs Hermes 4 70B.
The original releases are not refusal-reduced fine-tunes. Use Mistral Small 3 for cheap short scenes within 32K, or Llama 3.3 70B and Qwen3 235B A22B for longer ones. All 15 models are on the models page.
In SillyTavern’s API Connections, choose Chat Completion, then Custom (OpenAI-compatible). Enter https://weightsapi.com/v1 as the endpoint without appending /chat/completions, paste your key without the Bearer prefix, and select NousResearch/Hermes-4-70B or dphn/Dolphin-Mistral-24B-Venice-Edition. Send a Test Message to confirm. Details: SillyTavern integration guide.
Two settings protect your balance. Each request reserves credit for its input plus the maximum output you allow, then settles on actual usage, so a sensible response length keeps credit free in long chats. And give SillyTavern its own key with a lifetime spending cap and a Dolphin and Hermes allowlist, revocable from the console (key controls). The same endpoint suits Open WebUI and LibreChat.
Your character card and story context reach the model without a moderation layer from us. Fiction with villains, dark themes or hard drama is not blocked by a blanket filter; our line is lawful use and the model’s license. Create a username account.
A 70B creative fine-tune at $0.13 / $0.40. Venice, OpenRouter, NanoGPT and Chutes don’t list the 70B (25 September 2026). OpenRouter carries Hermes 4 only as the 405B at $1.00 / $3.00, plus Hermes 3 70B at $0.70 / $0.70; NanoGPT lists Hermes 4 405B at $0.30 / $1.20 and Hermes 3 70B at $0.408 / $0.408. See the Hermes 4 70B page.
$0.20 / $0.90, the same as Venice’s API for Venice Uncensored 1.2 and OpenRouter’s listing, which Venice serves. NanoGPT charges $0.40 / $1.80. No premium for Monero or for skipping the email. See the Dolphin page.
XMR, BTC, LTC, SOL, ETH, TRX, USDT on TRON or Ethereum, or USDC on Ethereum; the full USD amount becomes credit within about a minute once the network confirmations are reached, and you pay only the network fee. For API credit, Venice takes USD or USDC, OpenRouter USDC with a 5% crypto fee, Chutes TAO or crypto through Stripe. Every asset and network.
No email, phone, ID or card, so your stories are not tied to an inbox, and the gateway does not store prompt or response text, though the inference provider receives each request to answer it. No KYC is not anonymity: usage and payment records stay linked to your account. What we ask and record.
For a character app with steady traffic, the balance prepays a dedicated GPU: $416 per month for an L40S (48 GB), $1,352 for an H100 SXM (80 GB), no per-token charges on that machine, fit confirmed by quote. Venice, OpenRouter and NanoGPT sell no monthly GPU on pages we reviewed; Chutes bills private GPUs from $1.80 per hour plus a $5.40 deployment fee. See dedicated GPU plans.
| Item | What applies |
|---|---|
| Gateway filter | None added. Model training, your instructions and serving configuration can still cause refusals. |
| Price and unit | USD per 1 million tokens: Hermes 4 70B $0.13 / $0.40, Dolphin $0.20 / $0.90; 13 more models from $0.03 input. Pay as you go. |
| Minimum | USD 20 per top-up, every time; smaller remaining balances stay usable. No free trial. |
| Crypto and networks | BTC (1 confirmation), ETH (12), USDT on TRON (19), USDC on Ethereum (12), XMR (10), SOL (32), LTC (6); TRX and USDT on Ethereum also listed. |
| Fees | No platform deposit fee; you pay network fees. No cache, batch or volume discount. |
| Delays | Quote fixed for 10 minutes; credited within about a minute once the network confirmations are reached. |
| KYC scope | No ID or KYC step in the current flow. Payment records are kept; late or wrong-amount transfers go to review. |
| Availability | Available — deposit from $100, credited in about one minute |
| Dedicated GPUs | $416 to $9,736 per month; one month from activation, no automatic renewal. |
| Key conditions | No withdrawals to a wallet. Lawful use and model licenses apply. |
Same models, different terms. Prices are USD per million input / output tokens from public pricing pages or model lists.
| Provider | Dolphin Venice Edition | Hermes 4 70B | Crypto for API credit | Signup |
|---|---|---|---|---|
| WeightsAPI | $0.20 / $0.90 | $0.13 / $0.40 | XMR, BTC, LTC, SOL, ETH, TRX, USDT, USDC; no platform deposit fee; USD 20 per top-up | Username and password |
| Venice.ai | $0.20 / $0.90 (Venice Uncensored 1.2) | Not listed | USDC (credits also sold in USD) | Account for most features; USDC wallet access without one |
| OpenRouter | $0.20 / $0.90 (Venice: Uncensored, served by Venice) | Not listed (Hermes 4 405B: $1.00 / $3.00) | USDC; 5% fee on crypto purchases | Email or other contact information |
| NanoGPT | $0.40 / $1.80 (Venice Uncensored) | Not listed (Hermes 4 405B: $0.30 / $1.20) | BTC, Lightning, LTC, XMR, DOGE, DASH, ZEC and more; from $0.10 | No account needed |
| Chutes | Not listed | Not listed | TAO; other crypto through Stripe | Username plus Bittensor hotkey (CLI) |
Content terms differ too: Chutes’ terms prohibit generating “harmful, illegal, or offensive content”, a broad standard for a horror writer. We add no gateway filter and ask for lawful use.
Where others fit better: NanoGPT to start with a few dollars and no account; Venice for its consumer app, privacy modes and separate Role Play Uncensored model ($0.50 / $2.00); OpenRouter for a 460-model catalog and cheaper original models, such as Llama 3.3 70B at $0.10 / $0.32 against our $0.35 / $0.40. More in our OpenRouter alternative comparison.
What will it really cost? Tokens at published rates plus your network fee; no subscription. Roleplay is input-heavy because each turn resends the chat history. At the pricing calculator’s default 80% input / 20% output split, USD 100 buys about 540 million Hermes 4 70B tokens or about 290 million Dolphin tokens (our arithmetic).
Why a USD 20 minimum? Each top-up is a separate on-chain payment with its own network fee and verification; a USD 20 floor keeps that overhead small, and with no platform deposit fee all of it becomes credit. It is per top-up, not a balance to maintain. To spend $5 on a test, NanoGPT fits better.
Can I use it today? Create your account, deposit from $100, credited in about one minute. A published price alone does not mean a model is serving.
How do I pay, and how long does it take? Checkout fixes the crypto amount for 10 minutes, and the transfer must use the listed network; credit lands within about a minute once the network confirmations are reached (BTC 1, LTC 6, XMR 10). Wrong-network transfers have no recovery policy. See billing and crypto deposits.
Will I be asked for ID later? Not in the current account flow. Payment records stay tied to your account, and late or wrong-amount transfers go to review.
Can I get my money back? Credit cannot be withdrawn to a wallet, so top up for planned usage. The only refund is cancelling a GPU order before activation, back to your balance. Charges are grouped by UTC hour, model and key.
Choose WeightsAPI if you write interactive fiction, run SillyTavern characters or build a companion or game-dialogue app, want Hermes 4 70B or Dolphin with no gateway filter, and prefer paying in Monero, Bitcoin or stablecoins to handing over an email and a card.
Look elsewhere if you want to spend under USD 20, need guaranteed capacity or a latency commitment today, or need a withdrawable balance.
Next step: create a username account, store the recovery code offline, make a SillyTavern key with a cap and a Dolphin and Hermes allowlist, and top up once /status shows your model.
Username and password, no email, no card. Create your account, deposit from $100, credited in about one minute.
Short answers to what roleplay and creative users ask before signing up.
WeightsAPI is the precise version: no content filter added at the gateway, Dolphin Mistral 24B Venice Edition and Hermes 4 70B (fine-tunes designed to reduce refusals), and a balance funded in XMR, BTC, LTC, SOL, ETH, TRX, USDT or USDC. Models can still refuse, and lawful use applies.
No filter is added at the gateway, and the gateway does not store prompt or response text; the inference provider still receives each request. The model’s training, your instructions and the serving configuration decide the answer.
Yes. Both are tuned to reduce refusals, and Dolphin’s authors call it uncensored, but neither guarantees an answer every time. Rewording the system prompt or switching between the two on the same key usually helps.
$0.20 per million input tokens and $0.90 per million output tokens, with a 128K service context and image input. As of 25 September 2026, Venice and OpenRouter list the same price; NanoGPT lists $0.40 / $1.80.
WeightsAPI lists NousResearch/Hermes-4-70B at $0.13 input and $0.40 output per million tokens, 128K context. As of 25 September 2026, Venice, OpenRouter, NanoGPT and Chutes did not list it; OpenRouter and NanoGPT carry only the larger Hermes 4 405B.
Chat Completion, Custom (OpenAI-compatible) source, endpoint https://weightsapi.com/v1, your key without the Bearer prefix, and the Hermes or Dolphin model ID. Give that key its own spending cap and model allowlist.
Yes: XMR with 10 confirmations, alongside BTC, LTC, SOL, ETH, TRX, USDT and USDC. Each top-up is at least USD 20, with no platform deposit fee; credit cannot be withdrawn to a wallet.
You can request it. A dedicated GPU order asks which model you want to deploy and is paid from the same balance: $416 per month for an L40S (48 GB) up to $9,736 for 8 × H100, with no per-token charges on that machine for one month from activation. The operator confirms model compatibility, precision, context and concurrency before activation, and the activation date is confirmed before you pay.
Endpoint, key and model fields.
Prices, context, license, alternatives.
Prices, image input, behavior notes.
What signup asks for and what we record.
Assets, networks and confirmations.
Model-by-model prices and payment terms.