Base URL & Authentication
The Unhinged API is designed as a direct replacement for standard OpenAI endpoints. You interact with the API by pointing your client to our base URL: https://api.unhinged.top/v1. This ensures that any library compatible with the OpenAI format can communicate with our uncensored model without custom adapters. To authenticate, you need an API key, which you generate immediately upon signing up via Google or email on the Get API key page. Pass this key in the Authorization header as a Bearer token. The service supports one active key per account; generating a new key invalidates the previous one. This setup keeps your integration simple and secure, allowing you to focus on building features rather than managing complex authentication flows.
curl https://api.unhinged.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Chat Completions Endpoint
The core of the API is the POST /v1/chat/completions endpoint. It accepts a standard message structure and returns raw text outputs. You must specify the model ID as uncensored in your requests. This open-weight model is tuned to answer without content refusals for lawful adult use, making it ideal for creative writing, security research, or uncensored roleplay. The model does not gatekeep controversial or explicit topics, provided they do not involve minors. Requests are processed as text-in, text-out, with no hidden guardrails interfering with your data. You can configure the conversation history by passing an array of messages, each with a role (user, assistant, or system) and content. This endpoint is the only text generation path; we do not offer embeddings or image generation.
from openai import OpenAI
client = OpenAI(base_url="https://api.unhinged.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Streaming Responses
For applications requiring real-time feedback, the API supports Server-Sent Events (SSE). Streaming allows you to receive tokens as they are generated, reducing perceived latency for end-users. When using streaming, token usage statistics are included in the final chunk of the response, allowing you to track costs accurately in real-time. This is particularly useful for chat interfaces or coding assistants where typing speed matters. The streaming format is compatible with standard OpenAI SDKs, requiring only that you set the stream parameter to true. This feature ensures that your users see content immediately, rather than waiting for the entire completion to be generated and sent back in one block.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling & Tools
The uncensored model supports function calling, allowing your application to trigger external actions based on user intent. You can define tools (functions) in your request, and the model will return structured arguments for them. This is useful for building agents that can interact with APIs, databases, or user interfaces. You can control the model's behavior using tool_choice to force specific functions or let the model decide. This capability makes the API suitable for complex workflows where text generation needs to drive programmatic actions. The model handles tool definitions robustly, ensuring that your integrations remain reliable even when dealing with uncensored or unconventional prompts.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.unhinged.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
JSON Mode
When your application requires structured data, you can enforce JSON output by setting response_format to {"type": "json_object"}. This ensures the model returns strictly valid JSON, which is critical for parsing responses in code. This feature is especially useful when combining function calling with data extraction or when building APIs that consume structured text. By enforcing JSON mode, you reduce the need for post-processing regex or validation logic. The model is optimized to adhere to this constraint, providing reliable data structures for your downstream processes. This makes the API a robust backend for applications that depend on consistent, machine-readable outputs.
Parameters & Rate Limits
Control the model's creativity and determinism using parameters like temperature, top_p, stop, seed, and penalties. The context window supports 64,000 tokens total (prompt + completion), with a max output of 16,000 tokens per request (2,048 if unspecified). Rate limits are set to 300 requests per minute and 8 concurrent requests per key. If you exceed these limits, you will receive a 429 error. Authentication errors (401) indicate an invalid key, while 402 errors mean your prepaid credit is exhausted. Errors and refusals are free, so you only pay for successful token usage. This transparent pricing model ensures you know exactly what you are paying for.