Account access and model behavior
WeightsAPI’s sign-in flow has no identity-document upload or KYC step. Access and usage remain linked to your account. WeightsAPI adds no gateway content filter. The catalogue includes original models and published fine-tuned variants, including Dolphin and Hermes; their training and configuration can still cause refusals. Check service status for model availability.
Where request content goes
The gateway does not persist prompts, generated responses or tool arguments in its application database. Serving a paid request nevertheless requires sending the prompt or messages, and any tools, to an authoritative tokenizer; the inference provider receives the request needed to generate a response.
Those processing steps matter even without application content logs. Tokenizer and inference-provider retention, infrastructure logs, operational access and deletion behavior have not been independently verified. Treat the gateway’s storage behavior as one boundary, not an end-to-end zero-retention guarantee. The data-handling documentation describes the current scope.
Records the gateway keeps
Authentication and accounting require records beyond message content:
- Accounts and keys: platform account identifier and creation time; key identifier, name, visible prefix, SHA-256 hash, spending cap, allowed models, revocation flag and creation time. The complete key secret is returned only when created.
- Request accounting: account/key/model references, reservation amount, frozen input/output rates, state, creation time, token totals when settled and a provider response identifier when captured.
- Usage and ledger: hourly request and token totals by model and key; USD amounts, accounting references, entry type and timestamps.
- Payments: funding order reference, asset, network, fixed receiving address, USD target and recorded status; transaction, exchange-rate and confirmation details only when supplied and recorded. Withdrawal requests also retain destination, amount and review state.
- Other service records: endpoint metadata and confirmed-deposit records used to establish account access eligibility; any historical trial counters are not an access mechanism.
Retention durations, deletion procedures and access responsibilities for these records remain to be finalized.
Dedicated GPU orders retain the selected hardware and model, monthly price, payment and cancellation status, and service dates and endpoint details when provisioning is confirmed.
Controls with a defined scope
Stored API keys use hashes for authentication; names and prefixes remain readable for account management. The console supports separate keys, model allowlists, spending caps and revocation. Use separate keys for environments and applications, and keep secrets on your server. A revoked key cannot admit new requests, while requests already admitted can still settle.
AI access is attached to the signed-in account and requires one confirmed top-up of at least USD 100. There is no public free-trial quota. Once qualified, an account may spend remaining credit below USD 100 when sufficient for the request; every later top-up still has a USD 100 minimum. Any historical trial-counter records, if present, need a separate retention review and do not grant access.
Accounting is explicit about uncertainty
Available credit is the ledger balance minus active reservations. Before dispatch, the gateway reserves the maximum request cost using authoritative input-token counts and the output cap. Settlement uses provider-reported usage and the rates saved with that reservation. Missing or inconsistent usage keeps credit pending reconciliation instead of charging an estimated token count.
A funding order records the asset, network, receiving address, USD target and crypto quote. Credit requires a matched, verified transaction; creating an order or refreshing its status cannot confirm payment. Keep the order reference and public transaction identifier until the account shows confirmed credit. See billing.
Regions and reliability need evidence
Dallas, Ashburn and Portland are intended GPU regions. No deployed fleet, routing guarantee or verified data residency is claimed. A configured endpoint is not evidence of healthy service.
The status page reports availability and monitoring state. An empty incident history does not establish uptime. No availability percentage, delivery date or latency commitment is guaranteed.
Decide what is suitable before sending data
Evaluate outputs for your use case and check each model’s license and restrictions. An open-weight model or a lack of an added model-filtering layer does not remove your responsibility for lawful use, third-party rights or safe application behavior. Use synthetic data while the provider chain remains unverified.
Read the data-handling information and service conditions before committing a workload. Operator identity, jurisdiction, official contacts, retention periods and data-rights procedures are not published. No certification, independent audit or contractual service level is claimed.