Get API key

LLM Inference APIGuide

Troubleshooting Common AI Chat API Mistakes

Debugging an AI chat API integration often fails due to misconfigured headers, misunderstood token limits, or improper streaming handling. This guide addresses the most common implementation errors developers encounter when integrating OpenAI-compatible endpoints, ensuring your code runs reliably in production.

Updated:

Key points

  1. Always calculate token usage based on the specific model's tokenizer, not just character counts, to avoid context window overflows.
  2. Handle streaming errors by checking the HTTP status code before parsing the JSON stream, as network drops can leave streams in an inconsistent state.
  3. Ensure your request headers strictly match the API specification, particularly the Content-Type and Authorization fields, to prevent silent 400 or 401 errors.
  4. Implement rate limit backoff strategies immediately, as exceeding 300 requests per minute will result in 429 errors that halt your application.

Understanding Context Windows

One of the most frequent reasons for API failures is exceeding the context window. The context window defines the total number of tokens allowed in a single request, including both the input prompt and the generated completion. When this limit is reached, the API will reject the request with an error, often indicating that the sequence is too long.

Developers often mistake character count for token count. A single word can represent multiple tokens depending on the tokenizer. For example, a 64,000 token context window, such as the one provided by our inference api, allows for substantial conversation history or large document processing, but it is not infinite.

  • Monitor token usage: Use the official tokenizer for your model to accurately count tokens before sending a request.
  • Truncate intelligently: If you exceed the limit, remove the oldest messages from the conversation history rather than the most recent ones.
  • Account for overhead: Reserve some tokens for the model's response. If your prompt uses 63,000 tokens, you only have 1,000 tokens left for the completion.

Failure to manage this limit results in dropped connections or incomplete responses. Always verify your token counts against the model's documentation before deploying to production.

Handling Streaming Errors

Streaming responses via Server-Sent Events (SSE) are essential for a good user experience, but they introduce complexity in error handling. Unlike standard JSON responses, a stream can break mid-way through. If a network error occurs, your client might receive partial data, leaving the stream in an undefined state.

When implementing a stream consumer, you must handle the stream's lifecycle carefully. Check the HTTP status code before attempting to parse the stream. If the connection drops, you should log the error and decide whether to retry or display a message to the user.

Additionally, ensure your client correctly handles the end-of-stream marker. Some libraries expect a specific closing event, while others rely on the connection closing. Misunderstanding this can lead to hanging processes or memory leaks.

Always implement a timeout for your stream requests. If the API does not send a response within a reasonable time, abort the request to free up resources. This is crucial for maintaining stability in high-concurrency environments.

Token Counting and Limits

Token counting is not just about staying within the context window; it is also about cost management. Each token has a specific price, and miscalculating usage can lead to unexpected bills. While our pricing is transparent, with per-token rates for input and output, you still need to track usage accurately.

Most developers use a library to count tokens, but it is critical to use the correct tokenizer for the model you are using. Different models use different tokenizers, and using the wrong one can lead to significant discrepancies in counted tokens. For instance, a tokenizer trained on English text may handle punctuation differently than one trained on code.

Keep an eye on your usage limits. Our API allows 300 requests per minute per key. If you exceed this, you will receive a 429 Too Many Requests error. Implementing a simple counter in your application can help you stay within these limits and avoid service interruptions.

Finally, remember that token counts can vary slightly between different implementations of the same tokenizer. Always test your token counting logic with a few known inputs to ensure consistency.

Header Configuration Pitfalls

Headers are the configuration layer of your API requests. Misconfiguring them is a common source of 400 Bad Request or 401 Unauthorized errors. The two most critical headers are Content-Type and Authorization.

The Content-Type header must be set to application/json. If it is missing or incorrect, the API may not parse your request body correctly. The Authorization header must include your API key in the format Bearer YOUR_API_KEY. A common mistake is forgetting the Bearer prefix, which results in an authentication error.

  • Check for typos: Ensure your API key is copied correctly, including any trailing spaces or newlines.
  • Verify headers: Use a tool like curl or Postman to inspect the headers being sent.
  • Handle case sensitivity: Some APIs are case-sensitive for header names, though most modern APIs are not.

Always verify your headers before sending a request. A small mistake in a header can cause the entire request to fail, leading to confusion and wasted debugging time.

Rate Limit Management

Rate limits are in place to ensure fair usage and prevent abuse. Our API allows 300 requests per minute per key. If you exceed this limit, you will receive a 429 Too Many Requests error. This error includes a Retry-After header, which indicates how long you should wait before making another request.

To manage rate limits effectively, implement a backoff strategy. Instead of retrying immediately, wait for a period that increases exponentially with each retry. This prevents your application from overwhelming the API during peak times.

Monitor your usage metrics. Most APIs provide a dashboard or API endpoint to track your request volume. Use this data to optimize your application's request pattern. If you are making too many small requests, consider batching them together.

Remember that rate limits are per key, not per account. If you have multiple keys, each key has its own limit. Plan your key distribution accordingly to avoid hitting limits unexpectedly.

Error Code Interpretation

Understanding error codes is crucial for debugging. The most common errors you will encounter are 400 Bad Request, 401 Unauthorized, 429 Too Many Requests, and 500 Internal Server Error.

  • 400 Bad Request: This usually indicates a problem with the request body, such as missing fields or invalid JSON. Check the error message for details on which field is incorrect.
  • 401 Unauthorized: This indicates an issue with your API key. Verify that the key is correct and has not been revoked.
  • 429 Too Many Requests: This indicates you have exceeded the rate limit. Implement a backoff strategy to handle this gracefully.
  • 500 Internal Server Error: This indicates a problem on the server side. Retry the request after a short delay.

Always log the error response body. It often contains valuable information about what went wrong, such as the specific field that caused the error. This can save you hours of debugging time.

Optimizing Request Bodies

The request body is the core of your API interaction. Optimizing it can improve performance and reduce costs. A common mistake is sending too much data in a single request. If your prompt is too large, you may exceed the context window or incur higher costs.

Structure your JSON carefully. Ensure that all required fields are present and that optional fields are only included when needed. For example, if you do not need streaming, do not include the stream parameter. This reduces the payload size and simplifies the response.

Use tools like curl or Postman to test your request bodies. This allows you to verify that the JSON is valid and that the API is interpreting it correctly. It also helps you identify any unnecessary data being sent.

Finally, consider caching responses for identical requests. If you are sending the same prompt multiple times, you can store the response locally and avoid making the API call again. This can significantly reduce latency and costs for repetitive tasks.

Debugging Tool Calling

Tool calling allows the model to execute functions based on user input. Debugging tool calling can be challenging because it involves multiple steps: sending the request, receiving the tool call, executing the function, and sending the result back to the model.

Ensure that your function definitions are accurate. The schema must match the actual function signature. If the schema is incorrect, the model may generate invalid arguments, leading to errors when you try to execute the function.

Log the tool call arguments and the function output. This allows you to verify that the model is generating the correct arguments and that your function is executing as expected. If there is an error, the log will help you identify the issue.

Handle errors gracefully. If the function execution fails, send an error message back to the model so it can adjust its response. This provides a better user experience and allows the model to recover from errors.

Questions and answers

How do I calculate token usage for my API requests?

Use the official tokenizer library provided for your specific model. Character count is not a reliable indicator of token count, as different characters can represent different numbers of tokens. Most SDKs provide a utility function to count tokens accurately.

What happens if I exceed the rate limit?

You will receive a 429 Too Many Requests error. The response will include a <code>Retry-After</code> header indicating how long you should wait before retrying. Implementing an exponential backoff strategy is recommended to handle this gracefully.

Can I use any OpenAI-compatible SDK with this API?

Yes, any SDK that supports the OpenAI API format can be used by simply changing the <code>base_url</code> and <code>API_KEY</code> environment variables. This includes Python, Node.js, and other popular languages.

How do I handle streaming errors in my application?

Check the HTTP status code before parsing the stream. If the connection drops, log the error and decide whether to retry or display a message to the user. Implement a timeout to prevent hanging processes.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key