Docs How to use Bu0y
Bu0y speaks the OpenAI chat API. Get a bu0y_… key, add credits, then
point Cursor, Hermes, or any compatible client at
https://api.bu0y.com/v1.
No. 01 Get a key
bu0y_… secret once. It is not shown again.Sign a Solana challenge, mint a key, fund the deposit address, then sync. Full flow in llms-full.txt or the home page agent panel.
No. 02 Plug it in
| Field | Value |
|---|---|
| Base URL | https://api.bu0y.com/v1 |
| API key | Your bu0y_… secret |
| Model | Any id from Models (or /v1/models) — you pick it; clients do not auto-fill |
Stop the base URL at /v1. Clients append
/chat/completions themselves.
No. 03 Cursor
bu0y_… key.https://api.bu0y.com/v1.
If the connection fails, try Network → HTTP Compatibility Mode → HTTP/1.1. Agent
mode speaks OpenAI’s Responses API (/v1/responses); Bu0y serves it on
the same route, key and balance as chat completions.
No. 04 Hermes
Hermes accepts any OpenAI-compatible server. Easiest path:
hermes model → Custom endpoint → enter base URL, key, and model.
# ~/.hermes/config.yaml
# custom_providers MUST be a list (each entry starts with "-")
custom_providers:
- name: bu0y
base_url: https://api.bu0y.com/v1
api_key: bu0y_…
model:
default: glm4.6
provider: custom:bu0y
A dict under custom_providers causes
Unknown provider 'custom:bu0y'. Leave the base URL at
/v1. Hermes appends /chat/completions.
Prefer an explicit max_tokens; Bu0y reserves worst-case
before every fill. Your ceiling is sent upstream exactly as set; the
reserve holds against a 1024-token floor underneath it so reasoning
models cannot bill past the hold. You pay for tokens used.
No. 05 OpenAI SDK
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.BU0Y_KEY,
baseURL: "https://api.bu0y.com/v1",
});
const res = await client.chat.completions.create({
model: "glm4.6",
messages: [{ role: "user", content: "hi" }],
max_tokens: 256,
});
Same pattern in Python: OpenAI(api_key=…, base_url="https://api.bu0y.com/v1").
No. 06 Anywhere else
Look for Base URL / OpenAI Compatible / Custom provider. Use the three fields above. That covers Cline, Continue, LibreChat, Open WebUI, LangChain, Vercel AI SDK, and most agent stacks.
curl https://api.bu0y.com/v1/chat/completions \
-H "Authorization: Bearer $BU0Y_KEY" \
-H "content-type: application/json" \
-d '{
"model": "glm4.6",
"messages": [{"role":"user","content":"hi"}],
"max_tokens": 256
}'
Free price check before a fill:
POST /v1/quote with
model, inputTokens, maxOutputTokens.
No. 07 Add credits
Credits are non-refundable fill rights. No withdraw.
Available networks are shown in your desk. Check
GET /v1/deposits/status before requesting an address;
a network with configured: false is unavailable.
GET /v1/deposits/solana → send mainnet USDC to that address (≥ $1) →
Check now / POST …/sync.
GET /v1/deposits/base → send USDC to that address (≥ $1) →
Check now / POST …/sync.
GET /v1/deposits/robinhood/usdg → send USDG contract
0x5fc5…d168 to that address (≥ $1) → Check now /
POST …/sync.
No. 08 Rules that matter
No tab. A fill that cannot reserve the worst case fails before it runs.
Nothing fills under the margin floor. When a smaller ceiling would have, error.retry_max_tokens names it exactly.
Cheapest survivor wins. No private lane for a marketplace.
"stream": true relays tokens as the model produces them: first
token in about a second, and no total time cap. A buffered fill waits for the
whole answer and stops at 300s. Either way settlement completes before the
terminal frame, so a [DONE] you have seen is already in the ledger.
Bu0y does not store your prompts or completions. We keep billing metadata only. The marketplace that fills the request still receives the prompt to run it.
Repeat a byte-identical prompt prefix and the cached slice bills at the
cache-read rate — typically about a tenth of the input price, on models whose
market quotes one. usage.prompt_tokens_details.cached_tokens
reports the hit. Nothing to enable, no cache_control.
No. 09 FAQ
max_tokens is bigger than your
balance. Fund the account or lower the ceiling. 502 means the marketplaces
failed. Retry. They are not the same error.
No. 10 Changelog
max_tokens was
refused up front (400 unservable_max_tokens) when the model's recent
speed could not finish it inside the ceiling. A ceiling is not a forecast, and
the refusal cost agent loops a 400 and a retry on every request. A long answer
that reaches the wall is handled by the wall, not refused in advance.
min_tokens_per_sec and max_seconds:
pay more for a route that finishes. A route whose recent decode speed is slower
is skipped even when cheapest; 503 unmet_speed when no route is
known to qualify.
bu0y.billed_micros and
bu0y.completion_tokens for the output that reached you. Read the
error frame rather than assuming an errored stream was free, and keep the partial
output. A stream that breaks before any output still costs nothing.
No. 11 Full reference
Auth challenge details, deposit sync, error codes, and the complete route list live in the machine-readable dumps. Prefer those over scraping this page.
/llms.txt
Short index
/llms-full.txt
Full API and money rules
/docs.md
Markdown twin of this page
Bu0y is not affiliated with any marketplace it quotes. Nothing on this site is investment advice.