Inference API: Quickstart Guide
Get started with our uncensored inference API in under five minutes. Change your base URL and API key, then run your existing LLM code against our open-weight model with transparent per-token pricing.
Base URL and Authentication
Our API is fully OpenAI compatible. To switch, update your client configuration to point to our base URL and provide your API key. The base URL is https://api.llminferenceapi.com/v1. Authentication is handled via the standard Authorization: Bearer header. You receive your key immediately after signing up with just an email and password—no phone number or credit card required for the trial.
- Base URL:
https://api.llminferenceapi.com/v1 - Header:
Authorization: Bearer YOUR_API_KEY - Model ID:
uncensored
This setup works with any official OpenAI SDK or compatible client library. You keep your existing code structure; only the endpoint and key change.
First Request
Make your first call to the chat completions endpoint. The model ID is uncensored. This model is an open-weight large language model tuned to answer without content refusals for lawful adult use, not GPT or Claude.
curl https://api.llminferenceapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The request body includes the standard messages array. You can specify temperature, max tokens, and other standard parameters. The API returns a standard chat completion response with token usage statistics. If you encounter a 401 error, your key is invalid. A 402 error means your prepaid credits are exhausted.
Python SDK Integration
Using the Python SDK is straightforward. Initialize the client with our base URL and your API key. The code structure mirrors the standard OpenAI pattern. This allows you to swap in our llm proxy service without rewriting your application logic.
from openai import OpenAI
client = OpenAI(base_url="https://api.llminferenceapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Ensure you pass the model='uncensored' argument. The response object contains the generated text and usage metrics. This approach is ideal for developers who want a drop-in openai compatible api solution. No additional dependencies are required beyond the standard SDK.
Node.js SDK Usage
For Node.js developers, the integration is equally simple. Set the baseURL and apiKey in your client configuration. The SDK handles the HTTP requests to our ai chat api endpoints automatically.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llminferenceapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Call chat.completions.create() with the model ID uncensored. You can stream responses or handle them as complete objects. This flexibility supports both batch processing and real-time user interactions. The API supports tool calling, allowing you to define functions that the model can invoke.
Streaming Responses
For real-time user experiences, enable streaming. Set the stream parameter to true in your request. The API returns a Server-Sent Events (SSE) stream of partial responses.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Each chunk contains a partial delta of the response. You can render these tokens to the user as they arrive, reducing perceived latency. This is critical for chat interfaces where immediate feedback is expected. The streaming response includes the same token usage information as non-streaming requests.
Limits, Errors, and Context
Our API enforces a 64,000 token context window for both input and output combined. This supports long documents and complex conversations. Rate limiting is set to 300 requests per minute per key. If you exceed this, you receive a 429 error. Request bodies are limited to 8 MB.
Common errors include 401 (invalid key), 402 (insufficient credits), and 429 (rate limit). Credits never expire, so you can top up at your convenience. Prepaid credits start at $10, with bonuses for larger deposits. This pay as you go ai api model ensures you only pay for what you use.
Under the hood: specs
A quick checklist for developers: format, limits, features, billing.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.llminferenceapi.com/v1 |
| Authentication | Authorization: Bearer YOUR_KEY |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| JSON mode | JSON object mode via response_format json_object |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Context window | 64,000 tokens (prompt + completion together) |
| Request size | 8 MB request body |
| Rate limit | 300 requests per minute per key |
| Parallel requests | up to 8 in parallel per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Credit expiry | no monthly fee; paid credit does not expire |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Payment | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Account | sign in with Google or with e-mail + password |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Key management | one key per account, regenerate any time (the old one stops working) |
HTTP errors
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is this API suitable for production applications?
Yes, it is designed for production use with standard rate limits and reliable endpoints. However, it does not offer SLA guarantees or enterprise certifications like SOC2. It is best for applications that prioritize uncensored output and flexible pricing over formal enterprise guarantees.
What does 'uncensored' mean for this model?
The model does not refuse lawful adult, fictional, security-research, or controversial topics. It will generate content that other models might block due to safety filters. The only hard limit is that sexual content involving minors is always blocked.
Do I need a credit card to start?
No. You can sign up with just an email and password. Every new account receives $0.50 in trial credit valid for 7 days. You can add credits later using crypto (USDT or USDC) when you are ready to scale.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key