Bu0y

Docs How to use Bu0y

One key. Your usual tools.

Bu0y speaks the OpenAI chat API. Get a bu0y_… key, add credits, then point Cursor, Hermes, or any compatible client at https://api.bu0y.com/v1.

Base URL https://api.bu0y.com/v1

No. 01 Get a key

Humans first

  1. Open Account, accept the terms, and create a passkey.
  2. Mint an API key. Copy the bu0y_… secret once. It is not shown again.
  3. Add credits with USDC on Solana or Base, or USDG on Robinhood Chain (minimum $1).
  4. Paste the key into your tool. Set the base URL below.

Open Account

Agents (no passkey)

Sign a Solana challenge, mint a key, fund the deposit address, then sync. Full flow in llms-full.txt or the home page agent panel.

No. 02 Plug it in

Three numbers to remember

Field Value
Base URL https://api.bu0y.com/v1
API key Your bu0y_… secret
Model Any id from Models (or /v1/models) — you pick it; clients do not auto-fill

Stop the base URL at /v1. Clients append /chat/completions themselves.

No. 03 Cursor

Override the OpenAI endpoint

  1. Open Cursor Settings → Models.
  2. Turn on OpenAI API Key and paste your bu0y_… key.
  3. Turn on Override OpenAI Base URL and set https://api.bu0y.com/v1.
  4. Add a model id from Models (copy the id), then Verify.

If the connection fails, try Network → HTTP Compatibility Mode → HTTP/1.1. Agent mode speaks OpenAI’s Responses API (/v1/responses); Bu0y serves it on the same route, key and balance as chat completions.

No. 04 Hermes

Custom endpoint

Hermes accepts any OpenAI-compatible server. Easiest path: hermes model → Custom endpoint → enter base URL, key, and model.

# ~/.hermes/config.yaml
# custom_providers MUST be a list (each entry starts with "-")
custom_providers:
  - name: bu0y
    base_url: https://api.bu0y.com/v1
    api_key: bu0y_…

model:
  default: glm4.6
  provider: custom:bu0y

A dict under custom_providers causes Unknown provider 'custom:bu0y'. Leave the base URL at /v1. Hermes appends /chat/completions. Prefer an explicit max_tokens; Bu0y reserves worst-case before every fill. Your ceiling is sent upstream exactly as set; the reserve holds against a 1024-token floor underneath it so reasoning models cannot bill past the hold. You pay for tokens used.

No. 05 OpenAI SDK

Change two fields

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.BU0Y_KEY,
  baseURL: "https://api.bu0y.com/v1",
});

const res = await client.chat.completions.create({
  model: "glm4.6",
  messages: [{ role: "user", content: "hi" }],
  max_tokens: 256,
});

Same pattern in Python: OpenAI(api_key=…, base_url="https://api.bu0y.com/v1").

No. 06 Anywhere else

If it says “OpenAI compatible”

Look for Base URL / OpenAI Compatible / Custom provider. Use the three fields above. That covers Cline, Continue, LibreChat, Open WebUI, LangChain, Vercel AI SDK, and most agent stacks.

curl https://api.bu0y.com/v1/chat/completions \
  -H "Authorization: Bearer $BU0Y_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "glm4.6",
    "messages": [{"role":"user","content":"hi"}],
    "max_tokens": 256
  }'

Free price check before a fill: POST /v1/quote with model, inputTokens, maxOutputTokens.

No. 07 Add credits

Prepaid only

Credits are non-refundable fill rights. No withdraw. Available networks are shown in your desk. Check GET /v1/deposits/status before requesting an address; a network with configured: false is unavailable.

No. 08 Rules that matter

Short list

  • 01

    Prepaid

    No tab. A fill that cannot reserve the worst case fails before it runs.

  • 02

    Floor or 503

    Nothing fills under the margin floor. When a smaller ceiling would have, error.retry_max_tokens names it exactly.

  • 03

    Source-blind

    Cheapest survivor wins. No private lane for a marketplace.

  • 04

    Stream anything long

    "stream": true relays tokens as the model produces them: first token in about a second, and no total time cap. A buffered fill waits for the whole answer and stops at 300s. Either way settlement completes before the terminal frame, so a [DONE] you have seen is already in the ledger.

  • 05

    Prompts stay off our books

    Bu0y does not store your prompts or completions. We keep billing metadata only. The marketplace that fills the request still receives the prompt to run it.

  • 06

    Caching is automatic

    Repeat a byte-identical prompt prefix and the cached slice bills at the cache-read rate — typically about a tenth of the input price, on models whose market quotes one. usage.prompt_tokens_details.cached_tokens reports the hit. Nothing to enable, no cache_control.

No. 09 FAQ

The usual questions

How is this different from OpenRouter?
OpenRouter is a catalog of official APIs behind one key. Bu0y shops inference marketplaces on every request and fills at the cheapest source that can actually deliver. Prepaid USDC. Source-blind. Nothing fills under the margin floor.
Why prepaid?
No tab. We reserve the worst-case bill before a fill runs. Credits are usage rights. Not a balance you can withdraw. Not a refund.
Can I pick which marketplace fills?
No. Cheapest survivor that clears policy wins. The response does not name the source. That is the product.
What is 402 vs 502?
402 means the reserve for your max_tokens is bigger than your balance. Fund the account or lower the ceiling. 502 means the marketplaces failed. Retry. They are not the same error.
Do you store prompts?
No. Billing metadata only: tokens, prices, timestamps. The marketplace that fills still receives the prompt, because it has to run it.
What are points?
Fill spend counted as points. One point is one cent. Points shows hashed handles, not names or wallets. $BU0Y would count this sheet. Sign in to see yours on the account desk.
Where do I get help?
Discord. That is the support channel. Join.

No. 10 Changelog

What changed for clients

2026-09-02 — slow and broken fills
Withdrawn the same day: for a few hours a large max_tokens was refused up front (400 unservable_max_tokens) when the model's recent speed could not finish it inside the ceiling. A ceiling is not a forecast, and the refusal cost agent loops a 400 and a retry on every request. A long answer that reaches the wall is handled by the wall, not refused in advance.
Timeout errors now say which clock ran out, including when it was the source's own gateway giving up before the first byte.
New request hints min_tokens_per_sec and max_seconds: pay more for a route that finishes. A route whose recent decode speed is slower is skipped even when cheapest; 503 unmet_speed when no route is known to qualify.
A broken stream can carry a charge. If the provider stops mid-answer, the closing error frame includes bu0y.billed_micros and bu0y.completion_tokens for the output that reached you. Read the error frame rather than assuming an errored stream was free, and keep the partial output. A stream that breaks before any output still costs nothing.
Fills can run longer. The absolute ceiling is 480s, up from 300s, and mid-answer silence is judged per model between 20s and 180s. A client-side timeout should be longer than 480s, or you will hang up on fills that would have finished and be billed for what was delivered.
Everything else in that release is internal and needs nothing from you.

No. 11 Full reference

When you need every endpoint

Auth challenge details, deposit sync, error codes, and the complete route list live in the machine-readable dumps. Prefer those over scraping this page.

Bu0y is not affiliated with any marketplace it quotes. Nothing on this site is investment advice.