Get API key

LLM Inference APIDocs

Inference API: Quickstart Guide

Get started with our uncensored inference API in under five minutes. Change your base URL and API key, then run your existing LLM code against our open-weight model with transparent per-token pricing.

Base URL and Authentication

Our API is fully OpenAI compatible. To switch, update your client configuration to point to our base URL and provide your API key. The base URL is https://api.llminferenceapi.com/v1. Authentication is handled via the standard Authorization: Bearer header. You receive your key immediately after signing up with just an email and password—no phone number or credit card required for the trial.

  • Base URL: https://api.llminferenceapi.com/v1
  • Header: Authorization: Bearer YOUR_API_KEY
  • Model ID: uncensored

This setup works with any official OpenAI SDK or compatible client library. You keep your existing code structure; only the endpoint and key change.

First Request

Make your first call to the chat completions endpoint. The model ID is uncensored. This model is an open-weight large language model tuned to answer without content refusals for lawful adult use, not GPT or Claude.

curl https://api.llminferenceapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The request body includes the standard messages array. You can specify temperature, max tokens, and other standard parameters. The API returns a standard chat completion response with token usage statistics. If you encounter a 401 error, your key is invalid. A 402 error means your prepaid credits are exhausted.

Python SDK Integration

Using the Python SDK is straightforward. Initialize the client with our base URL and your API key. The code structure mirrors the standard OpenAI pattern. This allows you to swap in our llm proxy service without rewriting your application logic.

from openai import OpenAI

client = OpenAI(base_url="https://api.llminferenceapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Ensure you pass the model='uncensored' argument. The response object contains the generated text and usage metrics. This approach is ideal for developers who want a drop-in openai compatible api solution. No additional dependencies are required beyond the standard SDK.

Node.js SDK Usage

For Node.js developers, the integration is equally simple. Set the baseURL and apiKey in your client configuration. The SDK handles the HTTP requests to our ai chat api endpoints automatically.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llminferenceapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Call chat.completions.create() with the model ID uncensored. You can stream responses or handle them as complete objects. This flexibility supports both batch processing and real-time user interactions. The API supports tool calling, allowing you to define functions that the model can invoke.

Streaming Responses

For real-time user experiences, enable streaming. Set the stream parameter to true in your request. The API returns a Server-Sent Events (SSE) stream of partial responses.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Each chunk contains a partial delta of the response. You can render these tokens to the user as they arrive, reducing perceived latency. This is critical for chat interfaces where immediate feedback is expected. The streaming response includes the same token usage information as non-streaming requests.

Limits, Errors, and Context

Our API enforces a 64,000 token context window for both input and output combined. This supports long documents and complex conversations. Rate limiting is set to 300 requests per minute per key. If you exceed this, you receive a 429 error. Request bodies are limited to 8 MB.

Common errors include 401 (invalid key), 402 (insufficient credits), and 429 (rate limit). Credits never expire, so you can top up at your convenience. Prepaid credits start at $10, with bonuses for larger deposits. This pay as you go ai api model ensures you only pay for what you use.

Under the hood: specs

A quick checklist for developers: format, limits, features, billing.

ParameterDetails
CompatibilityOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Base URLhttps://api.llminferenceapi.com/v1
AuthenticationAuthorization: Bearer YOUR_KEY
EndpointsPOST /v1/chat/completions · GET /v1/models
Modeluncensored
Completion lengthup to 16,000 tokens per request (default 2,048)
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
StreamingYes — server-sent events; the last chunk carries token usage
JSON modeJSON object mode via response_format json_object
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Context window64,000 tokens (prompt + completion together)
Request size8 MB request body
Rate limit300 requests per minute per key
Parallel requestsup to 8 in parallel per key
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Credit expiryno monthly fee; paid credit does not expire
Bonus credit+5% on $50+, +10% on $100+
Free trial$0.50 of credit valid 7 days, no card needed
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
PaymentUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Billingprepaid credit, charged by real token usage; errors and refusals are free
Accountsign in with Google or with e-mail + password
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Key managementone key per account, regenerate any time (the old one stops working)

HTTP errors

Every error is JSON with a type you can switch on. You are never charged for an error.

HTTPTypeWhat to do
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busytemporary overload, retry shortly

Questions and answers

Is this API suitable for production applications?

Yes, it is designed for production use with standard rate limits and reliable endpoints. However, it does not offer SLA guarantees or enterprise certifications like SOC2. It is best for applications that prioritize uncensored output and flexible pricing over formal enterprise guarantees.

What does 'uncensored' mean for this model?

The model does not refuse lawful adult, fictional, security-research, or controversial topics. It will generate content that other models might block due to safety filters. The only hard limit is that sexual content involving minors is always blocked.

Do I need a credit card to start?

No. You can sign up with just an email and password. Every new account receives $0.50 in trial credit valid for 7 days. You can add credits later using crypto (USDT or USDC) when you are ready to scale.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key