LLM Inference APIPricing
Inference API: Simple, Transparent Pricing
Prepaid credits power every request on our pay-as-you-go inference API, with transparent per-token rates and no hidden subscription fees. Your balance covers usage exactly, ensuring predictable costs from the first token to the last.
No subscription. Prepaid credit never expires.
Cost calculator
Worked examples
Summarizing a long technical document
Processing a 50,000-token prompt to generate a 1,000-token summary costs $0.0125 for input and $0.001 for output, totaling $0.0135. This workload fits well within our 64k context window and demonstrates low-cost batch processing.
Power-user chat session
A session using a 10,000-token context window with 2,000 output tokens costs $0.0025 for input and $0.002 for output. The total is $0.0045, showing that even with substantial context, the per-token model remains economical for interactive use.
High-volume content generation
Generating 50,000 tokens of output from a small 1,000-token prompt costs $0.00025 for input and $0.05 for output. The total is $0.05025, illustrating that output tokens drive the majority of the cost in content-heavy workflows.
Token Pricing
Our inference API charges strictly by token volume, with input tokens at $0.25 per million and output tokens at $1.00 per million. This structure reflects the computational difference between reading context and generating text. Input tokens cover the 64,000-token context window, allowing you to pass large prompts or conversation histories. Output tokens are billed for the model's generation. Because we do not bundle features like image or audio generation, you only pay for the text you actually process. This pay-as-you-go AI API model ensures that high-volume users are not overpaying for unused capabilities. The pricing is transparent, with no per-request overhead or minimum charges. You can calculate your exact cost by multiplying your token counts by these rates. This approach keeps costs predictable for both short interactions and long-running LLM hosting scenarios.
Prepaid Credits
Every request is deducted from your prepaid balance, which never expires. You top up with a minimum of $10 using crypto (USDT or USDC). This model ensures that your usage can never exceed what you have loaded, providing strict budget control. Unlike subscription services, you pay only for what you consume, making it an efficient LLM proxy for variable workloads. Credits are applied instantly, allowing immediate access to the API. There are no monthly fees or recurring charges, so you can top up only when needed. This flexibility supports developers who need burst capacity or want to test the uncensored LLM without committing to a long-term contract. Your balance remains available indefinitely, so you can pause usage without losing your funds.
Bonus Credits
We reward larger top-ups with bonus credits to extend your usage further. Adding $50 grants a 5% bonus, while $100 adds a 10% bonus to your account. These bonuses are applied to your prepaid balance and used for standard token consumption. This feature reduces the effective cost per token for power users who maintain a higher balance. The bonus credits do not expire and are subject to the same usage rules as your primary funds. This structure supports consistent LLM gateway usage without the friction of frequent small transactions. It also makes the inference API more cost-effective for teams or individuals with predictable, ongoing needs. Bonus credits are added automatically upon successful top-up.
What your credit buys
Same price for every feature — function calling and JSON mode cost nothing extra.
| Feature | Support |
|---|---|
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Max context | 64,000 tokens (prompt + completion together) |
| Structured output | JSON object mode via response_format json_object |
| Parallel requests | 8 requests at the same time per key |
| Requests per minute | 300/min per key |
| Payment | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Volume bonus | +5% from $50, +10% from $100 |
| Subscription | paid credit never expires, no subscription |
Questions and answers
Do prepaid credits expire?
No, your prepaid credits never expire. You can top up your account at any time, and the balance remains available indefinitely until you use it for API requests.
What happens if I exceed my request limit?
If you exceed the 300 requests per minute limit, you will receive a standard rate limit error. Your prepaid balance is not affected by rate limits; you simply need to wait for the window to reset before sending more requests.
Can I use the API for commercial purposes?
Yes, our <strong>uncensored LLM</strong> is available for lawful adult use, including commercial applications. As long as the content does not involve sexual content with minors, you can use the API for any lawful purpose, including business, research, or entertainment.
How do I track my token usage?
You can monitor your token usage directly in your account dashboard. Each API response includes token counts for both input and output, allowing you to verify your costs. You can also check your remaining balance at any time to ensure you have sufficient credits for your next request.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key