BYOLLM Cloud docs

Everything is on this page. Use your browser’s find (Ctrl+F) — it searches all of it at once, which is the whole reason this is one page.

About BYOLLM

Use your own AI on websites — the models you already have, on your own machine, for sites that never learn which one answered.

What is BYOLLM?

BYOLLM lets you use your own AI on websites. You install one small program on your computer. Then, websites that support BYOLLM can use the AI you already have — a free model running on your machine, or an AI service you already pay for — instead of the website paying for AI and passing the cost to you.

Why would I want this?

For you: your favorite model goes everywhere you do. New models work the moment you get them — not when a site gets around to adding them. Your prompts are encrypted end-to-end to your own device. Sites never learn which model you use, and your subscriptions are never shared. And sites that do not pay for AI can charge you less, or nothing.

For sites and developers: no AI bills — your users bring their own compute. No floating money: you do not pay inference bills up front and hope to collect later, and you never ask people to prepay just to try you. Free trials cost you nothing to offer. You can ship the features you kept private for fear of the API bill. One small integration, and your users choose the models.

Your device

The byollm program runs on your computer. It knows which AI services you have set up: free open-source models on your machine, metered services you pay per use, or your own subscriptions like Claude Pro/Max. When a website you have enabled sends work, your device runs it with the service you chose. Your prompts are encrypted end-to-end to your own device. byollm.cloud passes them along and cannot read them.

Sites

A website that wants to use BYOLLM says what it needs — “writing help,” “chat,” and so on. When you connect the site, you pick which of your services answers each one. Every result arrives labeled with whether it ran on your own device. You can turn a site off at any time, and it stops getting your work.

Teams (optional)

A team lets you share what runs on your devices with people you name — the free open-source models on your machine, or a metered service with a spending limit you set. Your subscription accounts (like Claude Pro/Max) are never shared with anyone. That is a rule, not a setting.

byollm.cloud (or your own relay)

Many sites, many devices, many people. byollm.cloud keeps track of who has allowed what and sends each job to the right device. It never sees your prompts. If you would rather run this part yourself, the relay is open source — you can run your own instead of using byollm.cloud.

Connect your device

Install the daemon, pair it with BYOLLM Cloud, approve it by fingerprint, and route your first job. Nothing runs on a device you have not approved.

1. Install Node, then the daemon

The daemon is a Node command-line tool. Install Node 20 or newer for your platform, then install byollm globally.

macOS (Homebrew):

brew install node
npm install -g byollm@latest

Linux (Debian/Ubuntu shown — use your distribution’s Node LTS package or nvm):

sudo apt-get install -y nodejs npm   # or: use nvm for Node 20+
npm install -g byollm@latest

Windows (PowerShell):

winget install OpenJS.NodeJS.LTS
npm install -g byollm@latest

2. Run setup

byollm setup

One command does the whole thing: it finds the model CLIs and local servers you already have, writes the config, pairs this device with your account, and offers to keep it running in the background. It also checks that a CLI it found can actually answer — a signed-out one reports its version quite happily — and offers to run the sign-in for you when it cannot.

byollm connect is still there and does one part of that: pairing, and nothing else. Reach for it when you are pairing an already-configured device with a second hub, or re-pairing one whose credential was revoked. On a device that is already paired it says so and does not repeat the ceremony.

Setup pairs before you have a model server, on purpose: comparing the fingerprint is the step that matters, and it is the one a person has to be present for. If nothing is serving models yet the daemon says so and pairs anyway — nothing routes to the device until a backend is healthy, and then work starts arriving on its own with no second pairing.

3. Approve it, by eye

The daemon prints a short pairing code and a key fingerprint. Sign in at the devices page, enter the code, and compare the fingerprint on screen to the one in your terminal before approving. That comparison is the whole security ceremony: you are pinning this exact device’s key to your account, and nothing routes to a device you have not seen the fingerprint of.

Pairing once covers every site you connect. Which sites your device serves is decided by your consents on the connected sites page — new consents reach your daemon automatically; revoking one stops routing without touching the pairing.

4. Have something to run models with

The daemon does not bundle a model server — it talks to one you already run. Any OpenAI-compatible server works through one backend type: Ollama, MLX (Apple Silicon), llama.cpp, or vLLM. Your Claude Pro/Max subscription is a separate backend type. The backends and models guide covers setting up each one, per platform.

The quickest way to have something, if you have no preference yet — this is what the default config expects, so nothing else needs editing:

brew install ollama    # macOS; ollama.com has Linux and Windows
ollama pull llama3.2
Pairing refuses until a backend answers. If nothing is running you get “No backend is reachable, so there is nothing to offer this app yet” — deliberate, because a device that pairs while advertising nothing is a device that silently never receives work. Run byollm services to see what it tried and where.

5. Check it, then use it

byollm services   # what is configured, healthy, and advertised
byollm sites      # which sites this device serves, with pinned fingerprints

Connect a site on the connected sites page — you will read exactly what that site may send before you agree — and its jobs start routing to your device. Connecting to sites is free at every tier, and your devices only ever run your jobs unless you explicitly join a team that shares compute.

How do I keep the daemon running?

byollm start asks your machine to start the daemon at login and restart it if it stops. Which supervisor does that depends on the machine, and the command names the one you got: launchd on macOS, a systemd user unit on Linux, Task Scheduler on Windows.

Windows: supported, with one thing to know. The scheduled task is registered for you, not for the machine, so it needs no administrator rights — it runs one program, as you, when you log in. Some managed machines still block task creation by policy. When that happens byollm sets itself to start from your Startup folder instead and says so, and the difference worth knowing is that the Startup folder does not restart it if it stops. Task Scheduler would have.

Either way, byollm run in a terminal serves jobs for as long as the window is open. That is the answer while a supervisor is being sorted out, and it is a complete one — nothing about how the daemon was started changes what it will do.

byollm start       # start at login, and restart if it stops
byollm status      # which supervisor, and whether it is running now
byollm stop        # take it back off

Where things live

Configuration is one file: ~/.byollm/config.json (%USERPROFILE%\.byollm\config.json on Windows). Everything the daemon will ever do is in that file — which backends exist, and which model serves each job kind. A job can never name a model, a URL, a path or a flag; there is no field on the wire for any of them. What runs on your device is your configuration’s decision, never a request’s.

Backends and models

How a job kind maps to a backend and a model you chose. The owner decides everything; a job decides nothing.

Which job kind should I use — generate or chat?

llm.generate is one prompt in, one completion out. Its payload is { prompt, system? } — use it for anything stateless: summarize this, draft that, classify, extract, rewrite. If your call would be a single string and an answer, it is generate.

llm.chat is a conversation: its payload is { messages: [{role, content}...], system? } with roles system/user/assistant — use it when prior turns matter: assistants, multi-step refinement, anything where the model must see what was already said. The site sends the history it wants the model to see on every call; nothing is stored between jobs.

Both are text-only and carry no options — no temperature, no model name, no flags. That is deliberate (the payload is data handed to a model, never configuration), and it is also why the kinds are a closed list: an unknown kind is refused, never guessed, and adding one is a protocol change with its own threat review — not a payload field.

How does a job reach the model I chose?

A job carries a kind (llm.generate or llm.chat) and data — never a model name, a URL, or a flag. Your config declares services, and each names the kinds it answers. That is the protocol’s law, not a missing feature: the payload is data handed to a model, never configuration, so nothing a site sends can change what your device runs.

// ~/.byollm/config.json
{
  "services": {
    // Keys ("mlx", "claude") are names YOU choose. The "type" field names
    // the adapter. One HTTP adapter covers Ollama, MLX, llama.cpp and
    // vLLM — they all speak OpenAI-compatible /v1/chat/completions.
    "mlx": {
      "type": "openai-http",
      "baseUrl": "http://127.0.0.1:8080/v1",
      "model": "my-voice-model",
      "kinds": ["llm.generate"]
    },
    "claude": {
      "type": "claude-cli",
      "model": "claude-opus-5",
      "kinds": ["llm.chat"]
    }
  },
  "concurrency": 2
}
Two services may answer the same kind. When they do, neither is advertised until you name the winner in defaults — for example "defaults": { "llm.generate": "mlx" }. Nothing is chosen for you, because the wrong guess is the metered one. There is still no per-request model selection, on purpose.

Ollama — macOS, Linux, Windows

Install from ollama.com, then serve any library model. Base URL: http://127.0.0.1:11434/v1.

ollama pull llama3.2
# custom LoRA: build a named model from a base + adapter
# Modelfile:
#   FROM llama3.2
#   ADAPTER ./my-lora
ollama create my-voice-model -f Modelfile

MLX — Apple Silicon Macs

MLX is Apple’s framework for running models on M-series Macs, and its server speaks the same OpenAI-compatible API. Install with pip and serve the model:

pip install mlx-lm
mlx_lm.server --model mlx-community/Meta-Llama-3.1-8B-Instruct-4bit \
  --port 8080

Point a backend at http://127.0.0.1:8080/v1 and set the route’s model to the name the server reports. Flags vary by mlx-lm version — mlx_lm.server --help is authoritative.

How do I use a LoRA adapter, and how do I know it loaded?

If you have --adapter-path ./my-lora in your notes — from an older guide, ours included — current mlx-lm ignores that flag: the server starts, answers every request, and serves the base model. Nothing errors, so the only way to find out is to notice the answers are not your model’s.

An adapter is applied per request instead, by naming it in the request body:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'content-type: application/json' -d '{
    "model": "mlx-community/Meta-Llama-3.1-8B-Instruct-4bit",
    "messages": [{"role": "user", "content": "..."}],
    "adapters": "./my-lora",
    "temperature": 0
  }'

Verify with a difference, never with a 200. A reply that arrives and reads plausibly proves the server is up and proves nothing about your adapter. Send the same prompt twice at temperature: 0 — once with adapters and once without — and compare the text. Identical output means the adapter is not loaded. Deterministic decoding is what makes that comparison evidence rather than an impression.

byollm cannot send that field yet. The HTTP backend posts a fixed body — model, messages, stream — so a route pointed at this server answers as the base model, correctly and silently. Serving a fine-tune through byollm needs per-request parameters the backend does not have; that is specced in byollm_016 and is not shipped. Until it is, this section is how you confirm an adapter works against mlx-lm directly.

llama.cpp — macOS, Linux, Windows

llama-server -m ./model.gguf --lora ./my-lora.gguf --port 8081

Base URL http://127.0.0.1:8081/v1. GGUF adapters apply with --lora; check your build’s flags.

vLLM — Linux (CUDA)

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora --lora-modules my-voice-model=./my-lora --port 8000

Base URL http://127.0.0.1:8000/v1, route model my-voice-model.

Your Claude Pro/Max subscription

The claude-cli backend runs jobs through the Claude command-line tool signed into your own account. It is subscription-class: the daemon locks it to your own jobs at every tier, forever — a teammate’s job that would land on it is refused by the machinery, not by convention. Details in the subscriptions guide.

Verify, always

byollm services

Shows what is configured, what is healthy, and what is therefore advertised. A backend that is down is never advertised, so you never receive work you cannot run. If a route’s server is not answering on its baseUrl, fix that first — the daemon will not paper over it.

Use your Claude, Codex or API accounts

Route your own work through subscriptions you already pay for — on every site you connect. Locked to you by the protocol, at every tier, forever.

Can I use my Claude or Codex subscription here?

If you pay for Claude Pro or Max, for ChatGPT Plus or Pro, or hold API keys with any provider, your device can serve your own jobs with them. Write on a site that supports BYOLLM and the drafts come from the account you already pay for — no second bill, no copy of your prompts held by the site, nothing metered by us.

Claude Pro/Max

Sign into the Claude CLI on your device, then add the backend and route a kind to it:

// ~/.byollm/config.json
{
  "services": {
    "claude": { "type": "claude-cli", "model": "claude-opus-5", "kinds": ["llm.chat"] }
  }
}

Codex

Sign into the Codex CLI on your device, then add the backend and route a kind to it. Same shape as Claude, because it is the same class of thing: a subscription you already pay for, running on your machine.

// ~/.byollm/config.json
{
  "services": {
    "codex": { "type": "codex-cli", "model": "gpt-5-codex", "kinds": ["llm.generate"] }
  }
}

byollm setup finds it on its own if it is installed. It also checks the CLI can actually answer rather than merely that the binary is there — a signed-out CLI reports a version quite happily — and offers to run the sign-in for you when it cannot.

On a machine with no browser — a server you reached over SSH, a container, a hosted device — sign Codex in first:

codex login --device-auth

Plain codex login waits on a localhost OAuth callback, which a machine you are not sitting at cannot show you; it hangs rather than failing, so there is nothing to read. Doing it first is the whole fix: byollm setup only offers a sign-in to a CLI that cannot answer, so one that already answers is passed straight over.

API accounts (metered)

Provider API keys work through HTTP backends the same way. Metered backends carry a cost ceiling in config so a runaway job cannot spend without a limit you set.

Can I share my subscription with my team?

Subscription-class backends serve only their owner’s jobs. This is enforced in the daemon per job — a teammate’s or a stranger’s job that names a route backed by your subscription is refused, whatever your tier, whatever your config. Vendor terms forbid third-party use of personal subscriptions, and the protocol makes keeping that promise mechanical rather than a thing you have to remember. Sharing with a team covers free-class backends only: open-source and custom models you host.

Why is my subscription locked to me?

BYOLLM’s job is to let compute you already own follow you around the internet — not to resell anyone’s subscription. The self-lock is what makes the first thing possible without ever becoming the second. It is also why the pricing works the way it does: connecting your own devices to sites is free at every tier, because your compute serving you costs us almost nothing and is the entire point.

Share devices with your team

One capable box — a Mac Studio running a big open model or your team’s own fine-tune — serving you and five people you name, from any site they use. That is what Team buys.

What does Team actually do?

Your compute follows you everywhere for free. Serving other people is what Team is: their jobs, your devices. $24 a month, flat — you and five people you name, and a seat costs them nothing: it grants access to your devices’ models, not a plan of their own.

How do I share my device with my team?

Subscribe to Team on the subscriptions page, then invite members by email from your team’s roster page. An invitation is not a membership: nothing happens until the person accepts under their own login. A Team holds six — you and five others — and pending invites count toward it: six members and six open invitations would be a roster of twelve that passed the check at every step, so an invitation takes its seat when it is sent.

On the device you are sharing, give the service you want to offer "offer": "team". It has to be a free-class service — an open-source model, or your team’s custom fine-tune (see services and models). Subscription services are never shared, so a device whose only service is your Claude account has nothing to offer a roster.

The roster is what admits, and nothing else has to be done. A device offering "offer": "team" runs a teammate’s job only when the job arrives carrying a grant the control plane signed for it — and the control plane signs one only for somebody on your roster. Your device checks that signature itself, against the key it pinned when you paired it, so admission does not rest on the relay having filtered correctly. Inviting a teammate and their accepting is the whole step. Taking them off the roster ends it at their next job: there is no cache to expire and nothing to revoke on the device.

What can my teammates see?

A member signs in, accepts the invitation, and then consents per site with the shared-compute disclosure in front of them. The sentence they read is blunt on purpose: jobs may run on a device the team shares, and the owner of that device can read what runs on it — every device keeps a local log of everything it has run, by design. Nothing routes for a member until they have read that and agreed.

Trust in a team is real and mutual: members trust the device owner with their prompts; the owner’s device holds strangers-to-nobody content. Team makes that arrangement explicit and revocable instead of informal — either side can end it on the connections page, and revocation is immediate.

Will my teammates run up my bill?

Jobs from anyone who is not the device’s owner run under community budgets: capped jobs per hour and per day, a wall-clock ceiling per job, and an output-size ceiling — all owner-configurable. Prompts from other people are also retained differently: full text for a bounded number of days, then reduced to a hash. Your device serving your team never means your device belonging to it.

What happens when Team ends?

If the subscription lapses, sharing stops — members’ jobs run only on their own devices again, nothing is deleted, and resubscribing restores the roster as it was. Members are told on their connections page in one sentence; nothing about their account otherwise changes.

Add BYOLLM Cloud to your site

Register and verify your domain, mount three files, and every user can bring their own model to your product — their compute, your features, nobody’s data in the middle.

1. Register and verify your domain

On the sites page, register your apex domain (one product spanning ten subdomains counts once — the domain is the unit). The dashboard shows a DNS TXT record; add it at your DNS host and verify. Verification proves you control the domain and pins your site’s identity key — the key every user’s device will verify your requests against. Registering costs nothing on any plan, and there is no limit on how many you register — what a plan governs is how many appsyou connect to your own devices, not how many sites you build.

2. Generate your site identity — once

npx @byollm/server@latest keygen   # prints BYOLLM_SITE_KEYS=...

Once — not per deploy, never at startup. Devices pin this identity when their owners approve pairing with your site through the hub; regenerating it strands every one of them. Store it as a sensitive environment variable, shared across all the apps on your domain (on Vercel: team Settings → Environment Variables, linked to each project).

3. Mount the handlers — the three-file integration

The npm README’s walkthrough is authoritative and current: @byollm/server. In short: a catch-all route mounting createHandler (pass a function, not an object — builds run without secrets), a memoized store (MemoryStore to start, the Supabase adapter for production), and your app calling enqueue.

Two names in the examples below are yours rather than ours. getApp() is the memoized accessor from that setup — it is File 2 of the reference, and every block here assumes you wrote it. runOnHostedModel() is your fallback, which we cannot write for you: it is whatever you already do when BYOLLM is not the answer. Neither is exported by any package of ours, and a reader who searches npm for them will not find them.

const job = await getApp().enqueue({
  kind: "llm.generate",
  owner: ownerId,             // who this is depends on your mode — see 4
  payload: { prompt },
});

// On BYOLLM Cloud the fallback is the CATCH, not a callback — see below.
let outcome;
try {
  ({ outcome } = await job.result({ timeoutMs: 120_000 }));
} catch {
  // Your fallback, your choice. Nobody was there, or nobody finished in time;
  // from here those look the same and your users want an answer either way.
  outcome = { outcome: "ok" as const, text: await runOnHostedModel(prompt) };
}
onNoRunner does not fire on BYOLLM Cloud, so this guide does not use it. It runs when the SDK can see that nobody is online — which it does by counting paired devices in your own store. On the cloud lane your users’ devices pair with the relay instead, so that count is not a question this side can answer, and the SDK deliberately does not ask it rather than answering confidently and wrongly. The wait ends at your timeoutMs, which your catch already handles.

If you run your own relay in direct mode, the devices are in your store and onNoRunner works as written — it is the same SDK, and the difference is where the devices are.
Two rules the API makes you keep. Your app never invents whose devices to use — the owner is established by the person, never asserted by a browser and never guessed by you. What that looks like differs by mode, and section 4 is the one part of this guide you cannot skim. And results from devices the user does not own arrive marked untrusted — disclose their origin and never feed them to a privileged step as first-party output.

What about work that takes minutes?

Translation, long summaries, anything where the model thinks for a while. Five things matter, and two of them used to be warnings and are now simply answers.

1. A job may take longer than a lease.        Leases renew while the device works.
2. Pump while a caller is waiting.           The payload is delivered BY the pump.
3. Catch the timeout yourself.               onNoRunner does not fire on this lane.
4. Check `stop` before you trust the text.   "length" means it was cut off.
5. Keep a few jobs per owner in flight.      A device runs its `concurrency` at a time.
1 and 4 were true the other way round until 2026-09-18, and a team building from this page found both by losing work. A job that outlived one lease was requeued and its finished result refused; and stop was sealed by the device, opened here and thrown away, so a truncated answer arrived indistinguishable from a whole one. Both are fixed. They are listed as facts rather than as history because a reader wants to know what is true, but the dates matter if you are on an older server.

On the fifth. A device runs a small number of jobs at once — its concurrency, which its owner sets and which defaults low. Queue twenty jobs for one person and most of them sit queued, and a job waiting behind another is unclaimed, so its TTL is running the whole time. Enough of them and the ones at the back expire having never been offered to anybody. Send a few, wait, send more.

Do I have to call anything while a job runs?

Yes, on BYOLLM Cloud, and this is the one that catches people. Your site does not hold a connection to the device. The relay brokers, and app.cloud.pump() is your side of that conversation: it collects claims, seals and hands over the payload, and picks up finished results.

// While you wait. The payload is delivered BY the pump, not by enqueue.
const job = await getApp().enqueue({ kind: "llm.generate", owner, payload });

const pumping = setInterval(() => void getApp().cloud?.pump(), 2000);
try {
  const { outcome } = await job.result({ timeoutMs: 120_000 });
} finally {
  clearInterval(pumping);
}
A job nobody pumps for never runs, and nothing tells you. When a device claims a job the relay starts a short clock — AWAITING_PAYLOAD_MS, published by @byollm/relay and a matter of seconds — waiting for your site to seal the payload for that device. Miss it and the job goes back to the queue with the work undone.

So a scheduled pump — a cron, a background worker every minute — collects results perfectly well and can never deliver a payload in time. That combination is the trap: enqueue works, the dashboard shows a device, and nothing ever runs. Pump while a caller is waiting.

What state can a job be in?

Seven, and you have to branch on which one arrived. Four are terminal: a result() stops waiting the moment a job reaches any of them.

queued     enqueued, waiting for a device to take it
claimed    a device has taken it; no heartbeat has renewed the lease yet
running    the device confirmed on its heartbeat that it still holds it

ok         terminal — the device returned a result
error      terminal — the device reported a failure
canceled   terminal — you asked it to stop
expired    terminal — nobody ran it in time

A lease that lapses does not end the job. If a device takes a job and then goes quiet, the job returns to queued and is offered again — and its TTL clock restarts, because the TTL measures how long a job has waited unclaimed rather than how long it has existed. A job whose device died would otherwise expire for time somebody spent working on it. Total lifetime is bounded separately by the deadline the job carries, which is absolute.

expired is a state a job can be in and never an outcome a device reports. The three a device can return are ok, error and canceled — a job that expired never ran, so there is nothing to report about it. If you are branching on the outcome rather than on the state, that is the case you will miss.

Why did the answer stop where it did?

Every ok result carries a stop reason, and it is present on every one of them — the whole value is the difference between finishing and being cut off.

end             the model finished on its own
length          the model stopped at its own output ceiling
stop-sequence   a configured stop token ended it
unknown         the adapter cannot tell, and says so

unknown is the default, and never end. An adapter nobody has updated must not be able to claim completion by saying nothing: if absence meant end, every un-updated adapter would report finished answers that were actually truncated. So a missing signal reports itself as missing.

length is the one to handle. It means the answer is cut off mid-thought, and rendering it as a complete reply is the defect this field exists to let you avoid — ask again with more room, or tell the reader it was truncated.

How do I cancel a job?

The handle enqueue returns carries cancel() beside result(), and app.cancel(jobId) does the same thing from anywhere you only have the id.

const job = await getApp().enqueue({
  kind: "llm.generate",
  owner: ownerId,
  payload: { prompt },
});

// The user navigated away, or your own timeout won the race.
await job.cancel();

What happens next depends on whether a device already has it. A job still queued becomes canceled at once, and that state is terminal — a result() waiting on it stops waiting and returns the canceled record rather than timing out. A job a runner already holds (claimed or running) keeps its state for now: the cancellation travels on that runner’s next heartbeat, and the runner reports canceled itself.

Cancelling is not a stop button on somebody else’s computer. On the cloud lane the relay is the only party talking to the device, so a cancellation stops a future result from being accepted — it does not reach into a machine that is already generating. For a moment the relay may still offer a job your site will now refuse. Cancelling a job that already finished changes nothing, and app.cancel() answers null for an id it does not know, so a cancel sent twice or sent late is safe rather than an error.

4. Who owner is — this differs by mode

owner names the person whose devices should do the work. There are two connection modes, and they do not use the same names for people. This guide walks the cloud path — steps 1 and 2 registered you with the hub — and direct mode is described here because the difference is exactly where integrations go wrong. Getting it wrong used to be the one integration mistake that produced no error; it now comes back as a 409 in the time one request takes, which is the difference between a typo you find in a minute and one you find when somebody complains.

Direct mode — devices reach your own handlers, and you are the only party who knows who anybody is. owner is your own user id, taken from your session and never from the client. This is what the example above shows.

Cloud mode — devices reach hub.byollm.cloud, and the hub has its own identity space: rosters, consents and budgets all speak BYOLLM ids. Your own user id means nothing there. owner must be the person’s BYOLLM id, and your job is to learn it from them and store the mapping.

A wrong id is refused at once — on the hub. The relay routes on a consented (site, owner) pair, and enqueue() asks whether the slot can be satisfied before it accepts the job. An id that is mistyped, or belongs to somebody who has not connected your site, matches no route and comes back as slot-unsatisfiable — one sentence, never why, because which service and whose device are theirs and not yours.

On some relays a wrong id is not an error — it is silence. The check is a property of the relay answering rather than of the id being wrong: a relay you run yourself, without a control plane wired to it, accepts the job and lets it expire. Ask yours which it does — the readiness endpoint says.

So validate ids when you receive them, not when you enqueue — the same advice as before, for a smaller reason. A refusal you get in one request is a good backstop and a bad first line of defence: it tells you at send time, which is after the person typed it, and it is the one part of this that depends on which relay you are talking to.

4a. Getting a user’s BYOLLM id

Ask them for it. Every signed-in person can read it on their account page under Your BYOLLM id, with a copy button — add a settings field, have them paste it, and store it against your own user record.

The id names them and authorises nothing: somebody holding it can address work to a person and cannot deliver it, because the consent row is what opens the route. So a pasted id is not a credential you are being trusted with, and a wrong one is harmless. Two things still matter — the id must come from the person’s own account, never inferred or guessed by your app, and they must connect your site on Connected sites before anything will route.

Check the id at the moment they give it to you, so the failure lands where they can fix it rather than inside a feature days later:

import { signSiteRequest } from "@byollm/protocol";

// `siteKeys` and `siteId` are the pair you were issued when you registered
// the site — the same two values every other call on this page signs with.
// No new credential and no new code path.
async function askConnectStatus(owner) {
  const body = JSON.stringify({ owner });
  const issuedAt = Date.now();
  const signature = signSiteRequest(siteKeys, {
    endpoint: "connect/status",
    siteId,
    issuedAt,
    body,
  });

  const res = await fetch("https://dashboard.byollm.cloud/api/connect/status", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "x-byollm-site": siteId,
      "x-byollm-issued-at": String(issuedAt),
      "x-byollm-signature": signature.signature,
    },
    body,
  });

  return res.json();   // { enabled }
}

A person’s BYOLLM id is the owner this call takes — the same value the Connect button hands you in { valid: true, owner }, so the two doors end in one identifier and this helper serves both. Then, wherever they paste one:

const { enabled } = await askConnectStatus(pastedByollmId);

if (!enabled) {
  // Existence-neutral, deliberately — see below.
  return "That id has no devices for you. Check it, or connect this site on byollm.cloud.";
}
runnerAvailability is the wrong call here, and it is the right one in direct mode. On the cloud lane it refuses: it counts runners in your own store, and on this lane devices pair with the relay rather than with you, so the honest answer is unknown rather than none. It reported “no device paired” to every cloud app that asked, which was confident, specific and false. It is still the right call in direct mode, where the devices really are yours.
Keep that message vague, on purpose. A typo’d id and an id belonging to somebody who has not connected you give the same answer, and that is the system working rather than a limitation to route around. Telling the two apart would turn this call into an account-existence oracle: probe a guess, sort the answers, enumerate. Never write “no such account” — and if you ever get an answer that does distinguish them, that is a bug worth reporting to us.

4c. The embedded button, and it stays on the page

Most sites do not build the redirect above by hand — they embed our button, which is a small document served from dashboard.byollm.cloud and framed by you. It reads whether this visitor has your site enabled without you holding any credential of ours, and you cannot restyle a control that makes a claim about somebody’s account.

<iframe
  title="BYOLLM"
  src="https://dashboard.byollm.cloud/embed/button?site=<your site id>"
  width="165" height="34" style="border:0;display:block"
  sandbox="allow-scripts allow-same-origin allow-popups allow-popups-to-escape-sandbox"
></iframe>
Leave it on the page after they connect. This is the one rule that costs a support ticket when it is broken, and our own test site broke it: it embedded the frame while disconnected and drew its own static “BYOLLM Enabled” label afterwards.

A label you draw cannot be pressed, cannot carry the settings gear, and cannot stop being true. Somebody who disconnects your site in another tab is left looking at your word that they are connected, with nothing to click. The frame owns the label, the state and the door — in both states.

It posts you a message when it knows something: { byollm: "status", connected: true | false }. There is no message for I could not tell — a third-party frame often cannot read a first-party cookie, and that is a permanent condition rather than a fault. Silence means unknown, and a site that treated it as false would tell somebody to connect an account they already connected.

4b. The Connect button — once you want this to scale

Pasting an id works and does not scale past people who will do it. Put a Connect BYOLLM button in your settings and the person never sees an id at all:

https://dashboard.byollm.cloud/connect
  ?site=<your site id>
  &state=<opaque, bound to your session>
  &return=<https URL on your verified domain>

They land on a consent screen naming your site and its verified domain, approve, and come back to your return URL with byollm_assertion and your state echoed. The assertion is short-lived, single-use, and means nothing on its own — your backend exchanges it for the id:

import { signSiteRequest } from "@byollm/protocol";

const body = JSON.stringify({ assertion });   // the value from the redirect
const issuedAt = Date.now();

// The same signing scheme your handlers already use to reach the relay —
// no new credential, no new code path.
const signature = signSiteRequest(siteKeys, {
  endpoint: "connect/verify",
  siteId,
  issuedAt,
  body,
});

const res = await fetch("https://dashboard.byollm.cloud/api/connect/verify", {
  method: "POST",
  headers: {
    "content-type": "application/json",
    "x-byollm-site": siteId,
    "x-byollm-issued-at": String(issuedAt),
    "x-byollm-signature": signature.signature,
  },
  body,
});

const { valid, owner } = await res.json();   // { valid: true, owner, site }

Then store owner against your user exactly as you would a pasted id.

Verifying proves identity once. Enablement is a separate question, and you may ask it again. The assertion above is single-use and short-lived: it tells you who just connected, and it is spent the moment you redeem it. The owner you store stays valid forever in the only sense it ever was — it names the same person — and it says nothing about whether they still want you. Somebody who disables you on Connected sites changes nothing about that id.

So ask POST /api/connect/status, signed the same way, with { owner }. It answers { enabled } — a live connection with something mapped to it, which is what decides whether your jobs get answered. Ask it per session or per page load. Do not cache it into a cookie and treat that as current: an enablement badge that outlives the consent it reports is a promise your page is making on somebody else’s behalf.

A refused enqueue is the other half of the same signal, and it costs you no extra request: EnqueueRefused means the relay declined to queue the job, so whatever your page last displayed, it is not enabled now. Downgrade on it. A network failure is not that — unreadable is not disabled, and logging somebody out because your own request failed tells them to redo something they have already done.

Two things this flow will refuse. Your return URL must be on the domain you verified — we will not forward a person anywhere else, because the screen they just read named that domain and the redirect has to land where the promise was. And {valid: false} is one answer for every failure: unknown, expired, already spent, minted for another site. Telling them apart would confirm to anyone replaying a string that a particular person connected at a particular time.

5. Team work: one user’s job on a teammate’s device

Nothing. There is no field for it, and that is the whole design: who may serve a job is the person’s decision, made on their dashboard — your site asks for kinds and purposes. If they mapped a purpose to a teammate’s shared model, that is where it runs; if they mapped their own, it runs there. You are not told which, and you do not choose.

audience is refused on the cloud route. It used to be a field, and it defaulted to “this user’s own devices” — so a site that simply never mentioned it silently broke sharing for every user who had a team, while working perfectly for everyone testing alone. Asking a site to declare something it is deliberately never told is a default in disguise. Remove it; the hub derives it from the mapping, and passing one now fails with the remedy in the message.
await getApp().enqueue({
  kind: "llm.generate",       // the kind your teammates route to the shared model
  purpose: "writing-assistant",
  owner: ownerId,             // still the person whose work it is
  payload: { prompt },
});
You cannot choose the model, and that is the feature. There is no model field on the wire. Which model answers is decided by the owner of whichever device takes the job — so a team shares a fine-tune by each owner routing a kind to it. Pick one kind for shared work and one for private work, and every member’s config follows the same convention.

6. Saying what the work is for

Your site declares the purposes it has — the distinct jobs it does, in your own words — and names one on each job. It never names somebody’s service. Which service answers a purpose is the person’s choice, made once on the consent screen and changeable any time:

await getApp().enqueue({
  kind: "llm.generate",
  purpose: "writing-assistant",   // one YOU declared, not one of their services
  owner: ownerId,
  payload: { prompt },
});

Leave it out and the job takes your site’s default purpose, which is what every job above does. A site that declares no purposes has exactly one, and the consent screen offers it as a single choice named after the site.

A manifest is the purposes themselves, keyed by id. Nothing wraps them — there is no purposes array, and a purpose does not carry a key field, because the key is its id.

1. No manifest at all — start here. Declare nothing and the site is treated as having one purpose covering everything it does. enqueue() works with no purpose, the consent screen shows your site’s name against a single slot, and the person picks one service for it. Most sites never need more.

2. The everyday one, plus something particular. Two purposes: what your site mostly does, and one thing people might reasonably want answered by a different model.

{
  "writing-assistant": {
    "label": "Writing Assistant",
    "description": "Outlining and brainstorming to beat the blank page",
    "kinds": ["llm.chat", "llm.generate"]
  },
  "fact-checker": {
    "label": "Fact Checker",
    "description": "Reviews facts in your non-fiction work and builds the reference list",
    "kinds": ["llm.generate"]
  }
}

3. Several named purposes. Worth it when the jobs are genuinely different work — somebody may want a big model for one and a fast local one for another. This is a real site’s manifest, as it stood on 2 September 2026.

It is a snapshot and it will drift: that site’s owner edits their own manifest whenever they like, so what they declare today is whatever they last saved, not what is printed here. Copy the shape, not the purposes — yours should name the work your site does.

{
  "books": {
    "label": "Books",
    "description": "Reads and parses your existing books for use across the site",
    "kinds": ["llm.generate"]
  },
  "fact-checker": {
    "label": "Fact Checker",
    "description": "Reviews facts in your non-fiction work and builds the reference list",
    "kinds": ["llm.generate"]
  },
  "revenue": {
    "label": "Revenue",
    "description": "Analyzes your sales, revenue, and ad spend performance",
    "kinds": ["llm.generate"]
  },
  "writing-assistant": {
    "label": "Writing Assistant",
    "description": "Outlining and brainstorming to beat the blank page",
    "kinds": ["llm.chat", "llm.generate"]
  },
  "style-trainer": {
    "label": "Style Trainer",
    "description": "Trains a model on your writing style to generate draft content in your voice",
    "kinds": ["llm.generate"]
  }
}
Two rules that cost something to learn late.
A key is an identity, not a name you can revise. It travels on every job and people’s choices are stored against it, so renaming one deletes that purpose and creates another — everybody who had chosen a service for it is unmapped and asked again. The label is the changeable part, and the only thing a consent screen shows.

Adding a purpose reaches everybody already connected.It is not a re-consent — nothing they agreed to has changed — but their connection stops covering everything you ask for until they map the new slot, and their Connected Sites card says so. Jobs naming an unmapped purpose are refused at once rather than waiting.
Copy these from the page, not from a terminal. Every description here is a sentence, so these lines are long — and a terminal that wraps one puts a real newline inside a JSON string, which JSON does not allow. It fails several hundred characters in, at a spot with nothing wrong on either side of it. If you are moving a manifest between machines, move the file.
You declare what it’s for; they decide what runs it. "writing-assistant" is a key in your manifest, and it resolves through the mapping that person authored or it resolves nowhere. A model name, a base URL or a sampling flag would mean the same thing everywhere, and would let your app describe what it wants instead of asking for what it needs. So a purpose you declared is permitted and a description is not — and a purpose nobody has mapped is refused rather than quietly served by something else.

The field is on the job’s routing metadata, never in the payload — so nothing a user types in a prompt can steer their job onto a different service.

A purpose may also say how it expects to be routed, with routing. One of its three values has behaviour today and the other two are reserved, so that a manifest you write now is still valid the day they land.

"user-choice" is the one that works: the person maps this purpose to one of their own services, and that mapping is the consent. It is what every purpose does today and what a purpose that says nothing gets. "site-fixed" is reserved for a site that pays for and pins the service, where the person is told and consents to that rather than to a mapping. "user-first-with-fallback" is reserved for the person’s service if they have one and the site’s if they do not.

Reserved means accepted and not yet acted on. The schema takes all three and nothing branches on the value, so a purpose marked "site-fixed" behaves exactly like "user-choice" until the lane that reads it ships. That is deliberate: a field you cannot write until the day it works is a field that makes every manifest need editing on that day. When a lane does land, a hub that has not implemented it treats the purpose as "user-choice" and says so in a field, rather than quietly serving something the person did not agree to.

Two independent yeses are required before a job lands on someone else’s device: the device’s owner offers that service to their team, and the control plane signs a grant for this particular job — which it does only for somebody on that owner’s roster, running a purpose they mapped. The device verifies that signature itself. A site cannot grant either yes, and cannot see who said them.

7. What your users experience

A user with a connected device uses your AI features on their own compute — their local models or their own subscriptions — at no inference cost to you. A user without one reaches your fallback, which is whatever you already do today: on BYOLLM Cloud that is the catch around result(), because the SDK cannot see devices that pair with the relay and so does not claim to. You never see their model, their keys, or their devices; the protocol handles pairing, sealing, and provenance.

8. Prove it end to end

Pair your own daemon against your deployment, enqueue a real job, and read the result. Then certify: npx --package @byollm/conformance byollm-certify ./my-target.js runs the open conformance kit against your integration — the same checks the reference implementations pass.

Integration reference — paste this into your AI coding session

One page, complete and current: the files, the env vars, the rules, and the mistakes. Written to be handed to an LLM assistant verbatim. Device-readable index at /llms.txt.

What am I building?

Your site enqueues jobs; your users’ own devices run them. You never see their models or keys. Install npm install @byollm/server@latest. Two job kinds exist: llm.generate (stateless: { prompt, system? }) and llm.chat (history matters: { messages: [{role, content}], system? }). Payloads are text only and strictly validated — no model names, no temperature, no extra fields, ever.

File 1 — mount the protocol

// app/api/byollm/[...route]/route.ts
import { createHandler } from "@byollm/server/next";
import { siteKeysFromEnv } from "@byollm/server";
import { getStore } from "@/lib/byollm";

export const { POST } = createHandler(() => ({
  store: getStore(),
  siteKeys: siteKeysFromEnv("BYOLLM_SITE_KEYS"),
  verificationUrl: "https://your-app.com/settings/runners",
  basePath: "/api/byollm", // matches where Next mounts this route
}));
Pass a function, not an object. next build imports route modules with no secrets present; an object is constructed at import time and fails the build.

File 2 — the app and store

// lib/byollm.ts
import { ByollmApp, MemoryStore, siteKeysFromEnv } from "@byollm/server";

let store: MemoryStore | undefined;
export function getStore(): MemoryStore {
  return (store ??= new MemoryStore()); // lazy: module scope runs at build
}

let app: ByollmApp | undefined;
export function getApp(): ByollmApp {
  return (app ??= new ByollmApp({
    store: getStore(),
    siteKeys: siteKeysFromEnv("BYOLLM_SITE_KEYS"),

    // The cloud lane. Omit this block and you are on the direct lane, where
    // daemons reach your own handlers and you keep the runner records.
    lane: {
      // Both of these you wire. The library reads BYOLLM_SITE_KEYS and
      // nothing else — see "Which environment variables" below.
      relayOrigin: process.env.BYOLLM_RELAY_ORIGIN ?? "https://hub.byollm.cloud",
      siteId: process.env.BYOLLM_SITE_ID!,
    },
  }));
}

The lane is this one block, and it changes what the rest of this page means. With lane set, devices pair with the relay and the hub holds consents, rosters and budgets; owner becomes their BYOLLM id, audience stops being yours to state, and the availability question below has no answer. Without it, daemons reach your handlers directly and every one of those is yours.

Production: swap MemoryStore for the Supabase adapter (see the npm README) — same interface, real persistence.

File 3 — enqueue and read

const job = await getApp().enqueue({
  kind: "llm.generate",
  purpose: "writing-assistant",  // one YOU declared — never one of their services
  owner: ownerId,          // DIRECT: your session's user id
                           // CLOUD:  their BYOLLM id — yours means nothing
  payload: { prompt: `Summarize:\n\n${text}` },
});

// On BYOLLM Cloud, nobody-online and nobody-in-time arrive the same way:
// the wait ends at timeoutMs and your catch answers. onNoRunner fires in
// direct mode, where the devices are in your own store.
let outcome;
try {
  ({ outcome } = await job.result({ timeoutMs: 120_000 }));
} catch {
  outcome = await yourHostedFallback(text);
}

yourHostedFallback() is yours — we cannot write it, and no package of ours exports it. It stands for whatever you already do when BYOLLM is not the answer: a hosted model, a cached reply, or telling the person to try later. The point of the block is that on the cloud lane it belongs in the catch, because nobody-online and nobody-in-time arrive there identically.

Who owner is — this differs by mode

Direct (daemons reach your handlers): your own user id, from your session, never from the client. Cloud (daemons reach the hub): the person’s BYOLLM id. The hub’s identity space — rosters, consents, budgets — speaks it, and your id for them names nobody there.

They read it on their byollm.cloud account page and paste it into your settings, or hand it over through the Connect button. Either way it comes from their own account and never from a guess.

On hub.byollm.cloud a wrong id is refused at once, not waited on. Reach for a runnerAvailability() check at paste time and you are solving a problem this relay does not have — and on a relay without the check, where a bad id really does go unanswered until the job times out, that call refuses anyway on the cloud lane. Which you have is a question you can ask.

enqueue() asks whether the slot can be satisfied before it accepts the job, so an id nobody has mapped comes back as a 409 in the time one request takes:

try {
  const job = await getApp().enqueue({ kind: "llm.generate", purpose, owner, payload });
} catch (error) {
  // `slot-unsatisfiable`      nobody has chosen what answers this — needs a person
  // `slot-waiting`            mapped, but nothing can answer right now — needs time
  // `purpose-not-declared`    this purpose is not in your manifest
  if (error instanceof EnqueueRefused) {
    // Existence-neutral, deliberately — never "no such account".
    return "Nobody has chosen a model for that yet. Check the id, or connect this site on byollm.cloud.";
  }
  throw error;
}

Do not reach for runnerAvailability() here. It refuses on the cloud lane — it counts runners in your own store, and cloud-lane devices pair with the relay instead, so the only lane where somebody pastes a BYOLLM id is the one lane where that question has no answer. It stays for direct-lane sites, which own the store it counts.

Saying what the work is for

Your site declares purposes — the distinct jobs it does, in your own words — and names one on each job. It never names somebody else’s service. Which service answers a purpose is their choice, made on the consent screen and changeable any time.

await app.enqueue({
  kind: "llm.generate",
  purpose: "writing-assistant",   // one YOU declared, not one of their services
  owner: ownerId,
  payload: { prompt },
});
There is no service option, and passing one is refused. This page recommended it until 2026-09-16. It was removed before 0.1.0 — a site naming somebody’s service is a site describing what it wants instead of asking for what it needs, and the option is refused outright rather than ignored, because an ignored option is a job running differently than asked with nothing to see.

A manifest is the purposes themselves, keyed by id. Nothing wraps them — there is no purposes array, and a purpose does not carry a key field, because the key is its id.

1. No manifest at all — start here. Declare nothing and the site is treated as having one purpose covering everything it does. enqueue() works with no purpose, the consent screen shows your site’s name against a single slot, and the person picks one service for it. Most sites never need more.

2. The everyday one, plus something particular. Two purposes: what your site mostly does, and one thing people might reasonably want answered by a different model.

{
  "writing-assistant": {
    "label": "Writing Assistant",
    "description": "Outlining and brainstorming to beat the blank page",
    "kinds": ["llm.chat", "llm.generate"]
  },
  "fact-checker": {
    "label": "Fact Checker",
    "description": "Reviews facts in your non-fiction work and builds the reference list",
    "kinds": ["llm.generate"]
  }
}

3. Several named purposes. Worth it when the jobs are genuinely different work — somebody may want a big model for one and a fast local one for another. This is a real site’s manifest, as it stood on 2 September 2026.

It is a snapshot and it will drift: that site’s owner edits their own manifest whenever they like, so what they declare today is whatever they last saved, not what is printed here. Copy the shape, not the purposes — yours should name the work your site does.

{
  "books": {
    "label": "Books",
    "description": "Reads and parses your existing books for use across the site",
    "kinds": ["llm.generate"]
  },
  "fact-checker": {
    "label": "Fact Checker",
    "description": "Reviews facts in your non-fiction work and builds the reference list",
    "kinds": ["llm.generate"]
  },
  "revenue": {
    "label": "Revenue",
    "description": "Analyzes your sales, revenue, and ad spend performance",
    "kinds": ["llm.generate"]
  },
  "writing-assistant": {
    "label": "Writing Assistant",
    "description": "Outlining and brainstorming to beat the blank page",
    "kinds": ["llm.chat", "llm.generate"]
  },
  "style-trainer": {
    "label": "Style Trainer",
    "description": "Trains a model on your writing style to generate draft content in your voice",
    "kinds": ["llm.generate"]
  }
}
Two rules that cost something to learn late.
A key is an identity, not a name you can revise. It travels on every job and people’s choices are stored against it, so renaming one deletes that purpose and creates another — everybody who had chosen a service for it is unmapped and asked again. The label is the changeable part, and the only thing a consent screen shows.

Adding a purpose reaches everybody already connected.It is not a re-consent — nothing they agreed to has changed — but their connection stops covering everything you ask for until they map the new slot, and their Connected Sites card says so. Jobs naming an unmapped purpose are refused at once rather than waiting.
Copy these from the page, not from a terminal. Every description here is a sentence, so these lines are long — and a terminal that wraps one puts a real newline inside a JSON string, which JSON does not allow. It fails several hundred characters in, at a spot with nothing wrong on either side of it. If you are moving a manifest between machines, move the file.

How big can a job be?

One sealed envelope may be at most 6 MiB — that is your prompt after encryption, and the same ceiling applies to the answer coming back. A job over it is refused at the door rather than part-way through.

That number is read from the protocol, not typed here. It is MAX_ENVELOPE_BYTES, imported from @byollm/protocol and rendered — so this page cannot tell you a ceiling the package does not enforce. It has moved once already, from 10 MB to 6 MiB, and a page that had typed the old one would still be saying it.

Import it yourself if you want to check before enqueueing rather than be refused: it is a published export, for exactly that.

What changed, and when

The releases page is the changelog. Every entry is generated from the release note that shipped with the tag, so there is no second list here to fall behind it, and prereleases are labelled as such — a locked version is never mistaken for an alpha.

This page tells you what is true now. When something used to be otherwise and you might still meet the old behaviour — in a tutorial, in your own notes, in a running deployment — it is said where it matters rather than kept in a list you would have to know to check.

Eight things you have to decide, that nothing told you about

Each of these is a choice an integrator makes whether or not they know it exists. They were true and undocumented; the values below were read out of the code rather than remembered.

How result() waits

It polls, every 500 ms, and gives up after five minutes unless you pass timeoutMs. The Supabase adapter substitutes Realtime behind the same interface, so the call you write is the same either way — what changes is whether your database is being asked twice a second.

The ten-second grace you may have heard about is not this. NO_RUNNER_GRACE_MS is how long a sustained no-runner signal has to persist before onNoRunner believes it — a daemon restarting, or one whose heartbeat is a moment late, must not fail every job in flight. It has nothing to do with Realtime, and it never fires on the cloud lane, where onNoRunner does not run at all.

Ten seconds is indicative, not a promise. It is a private constant, and a private constant is one we may change without telling anybody. Do not build a timeout around it — the two numbers above are what a site budgets on, and they will be exported so you can read them rather than copy them.

ttlMs and deadlineAt

A job stops being worth carrying at deadlineAt ?? (claimableAt ?? now) + ttlMs. Pass deadlineAt to say when, or ttlMs to say how long; the default TTL is fifteen minutes.

The clock starts when the job becomes claimable, not when you created it. That matters for anything with a dependsOn: a job blocked for an hour behind another has not spent its fifteen minutes waiting.

dependsOn

A job is not claimable until every job it names has finished. It is how you order work without holding it in your own process — and it is the reason the TTL clock starts where it does.

id is your idempotency key

Pass your own id and enqueueing it twice is one job: the relay answers the second call with what is already routing. It is scoped to your site — the same id from a different site is refused, not answered with yours.

It is also the only handle you keep when enqueue throws, which is what makes tidying up after a refusal possible at all.

There is no streaming

The field exists on the wire and is false, reserved. A job produces one result when it is done. If you are building a chat UI, that is the thing to design around now rather than the thing to wait for.

How many jobs a device runs at once

Two, by default — the device owner may set anything from one to thirty-two. You are not told which, and you cannot ask: a number that describes somebody’s machine is theirs. Design for a queue that may be slow rather than one that is proportional to what you send.

The library retries nothing

There is no backoff, no attempt counter and no retry loop anywhere in @byollm/server. A RelayUnavailable carries retryable so you can tell a draining pod from a bad signature — but the deciding and the waiting are yours.

How does a manifest get registered?

You paste it into the site’s card on Developer Sites, and that is the only door. There is no registration call in @byollm/server, no REST endpoint you can post it to and no CLI command — if you went looking for one, you were right that it is not there.

The editor refuses a manifest it cannot parse and nothing is saved when it does, so a bad paste leaves the previous one intact rather than clearing it.

An edit reaches routing in a couple of seconds, not instantly. The relay does not read your manifest out of the database on every job. It routes from a projection the hub re-reads on a short interval — two seconds by default — so there is a window after you save where a job can still be judged against what you declared before.

You can watch it move. generation at /readyz counts projection changes, not refreshes, so it ticks when — and only when — something you did has landed:

curl -s https://hub.byollm.cloud/readyz | grep -o '"generation":[0-9]*'
Adding a purpose does not connect it for anybody. Every person who has already connected your site sees the new slot unmapped and chooses a service for it before it routes — so a job naming a purpose nobody has mapped is refused with slot-unsatisfiable until they do. That is the design, not a propagation delay, and no amount of waiting clears it.

Does a refused enqueue leave a job behind?

Yes, in your own store, and that is deliberate. The record is created and sealed at rest before the relay is asked, so a relay that is down costs you a routing delay rather than a lost job.

After a 409 that is a real refusal, though, nothing republishes it — publishing happens once, in enqueue. So the row sits in your store as a job that will never route, until whatever expiry your store applies removes it. It is not waiting for a pump and it is not retried.

Pass your own id if you want to tidy up. enqueue throws, so you never receive the handle it would have returned — and with it the cancel that marks the row terminal. An id you chose is one you still have after the throw, so app.cancel(id) is available to you and app.cancel(…) on an id you never saw is not.

A mapped slot whose device is asleep — is that unsatisfiable?

No. It has its own code, and the difference is the only thing you are told.

slot-unsatisfiable   nobody has chosen what answers this
slot-waiting         mapped, and nothing can answer it right now

The question they answer is does this need the person, or only time. slot-unsatisfiable means somebody has to go to their dashboard and choose a model, and no amount of waiting helps — send them there. slot-waiting means the slot may recover with nobody doing anything, so retry later and do not send them to a settings page to fix a laptop that is merely asleep.

slot-waiting carries no cause and no duration, and that is deliberate rather than unfinished. A device asleep, a service withdrawn and an account over its cap arrive as one sentence, because telling them apart would tell you how somebody has spent their day. A duration would do the same thing more precisely.

A relay that cannot see presence answers ok to both. Absent means unknown, and unknown is not a refusal — so where presence is invisible a job for a sleeping device is queued and expires, and your timeout is the answer. That is the older behaviour and it is still the behaviour of a relay without a control plane; ask which one you are talking to.

Does your relay refuse a job nobody can run?

It depends on the relay, and the relay will tell you. A job whose slot nothing can satisfy — an owner who has not connected you, a purpose mapped to nothing — can be refused at enqueue with slot-unsatisfiable, or accepted and left to expire. Both are real behaviours and this page has said each of them, because which one you get is a property of the relay answering rather than of your code.

It is worth asking, because the two are otherwise indistinguishable from outside: a relay that refuses nothing looks exactly like a relay that checked and found nothing wrong. Ask its readiness endpoint:

$ curl -s https://hub.byollm.cloud/readyz
{"status":"ready","generation":5,"satisfiability":"control-plane"}

satisfiability: "control-plane" means unsatisfiable slots are refused at enqueue — this is what hub.byollm.cloud answers. "none" means they are queued and expire, which is what a relay without a control plane wired to it reports. A relay you run yourself is whichever you built.

Write for the second one either way. A refusal at enqueue is a good backstop and a bad only-line-of-defence: it arrives after the person typed the id, and it is the part of this that depends on somebody else’s deployment. Validate ids when you receive them, and let the timeout answer when nothing comes back.

Which environment variables do I need?

# READ BY THE LIBRARY
BYOLLM_SITE_KEYS=...      # npx @byollm/server@latest keygen — run ONCE, keep forever

# READ BY YOU, and passed in — these names are ours only by convention
BYOLLM_RELAY_ORIGIN=https://hub.byollm.cloud   # cloud mode
BYOLLM_SITE_ID=...        # from your site card on dashboard.byollm.cloud
Only the first is read for you, by siteKeysFromEnv("BYOLLM_SITE_KEYS"). The other two are names this page made up. Nothing in @byollm/server looks them up — File 2 above reads them and passes the values into lane, and if you call yours something else it will work exactly as well. Setting them and passing nothing configures nothing, silently, which is the shape this callout exists to stop: a variable that is set, correct, and never read looks identical to one that is working.

Cloud mode is entirely outbound from your site — the hub never calls your endpoints. Register and DNS-verify your domain on the dashboard first; the identity in BYOLLM_SITE_KEYS must be the one you verified, and it is shared by every app on your domain.

The mistakes list (each has failed a real build)

1. Regenerating site keys — every paired device is stranded; keygen runs once. 2. Config object instead of a function in createHandler — build fails without secrets. 3. Pairing against the bare domain when the route lives under /api — daemons append /byollm/... to the origin they are given, so pair against https://your-app.com/api or move the route and drop basePath. 4. Taking owner from the request body — a client can never assert whose devices run a job; and on the cloud route, using your own user id at all, which names nobody the relay knows. On hub.byollm.cloud that is refused at once with slot-unsatisfiable; on a relay you run yourself it may instead be accepted and never routed, which is the same mistake with no error attached — see does your relay check. 5. Rendering community results as trusted — provenance.untrusted is constant on the cloud route and cannot be overridden: every result came from a device the person chose, which may be a teammate’s, and you are not told which — so the disclosure you owe your users is one sentence for every job. 6. Inventing payload fields (model, temperature, tools) — schemas are strict and the daemon re-validates; there is no field on the wire for any of them. 7. Blocking forever on results — always set timeoutMs and handle the wait ending; a user with no device online is the ordinary case, not an error. On BYOLLM Cloud that arrives as the timeout, because the SDK counts devices in your own store and your users’ devices pair with the relay — onNoRunner is the direct-mode form of the same fallback.

Verify the integration

# the open conformance kit
npx --package @byollm/conformance byollm-certify ./my-target.js

Then pair your own daemon against the deployment and run one real job end to end. If both pass, the integration is byollm-compatible — the same bar the reference implementations meet.

What BYOLLM Cloud can and cannot see

The router holds no key that opens anything. That is not a policy — it is the construction, and it is testable.

What can byollm.cloud actually see?

A job’s payload is encrypted on the site’s server to the specific device that claimed it; the result is encrypted on that device back to the site. The hub routes ciphertext and metadata — who may talk to whom, never what they said. There is no configuration of our infrastructure in which we could read a prompt, because no key that opens one ever reaches us.

How do I know I am talking to my own device?

A device joins your account only after you compare its fingerprint on the devices page and approve it. A site’s identity key is pinned by your device at pairing and verified on every job. A changed key is refused loudly — key rotation exists, but only as a signed succession the old key itself authorizes, checked independently by your device and by the control plane.

Consent is the routing table

Nothing routes to your devices from a site that is not on your connected list, and the sentence you read when connecting is the contract — including the honest one about shared devices: the owner of a device can read what runs on it, because every device keeps a local log of its own work, by design. Revoking a consent stops routing as soon as it reaches us, and revocation is recorded as an event, never quietly erased.

What we do see

Routing metadata: opaque account ids, site ids, job ids, sizes, timing, and dispositions. That is what the usage page and the savings estimate are built from — aggregates that cannot name a prompt. When we say a number is an estimate, the page says so, with the price table it came from.

Claims, by their MUST ids

Every sentence above is a named, tested rule in the open spec — not copy. The router reads nothing: RELAY_BLIND. Requests are signed, never bearer-authenticated: REQUESTS_SIGNED_NOT_BEARER. Keys are exchanged and pinned at consent and checked every time: KEYS_EXCHANGED_AT_CONSENT, SITES_LOCALLY_APPROVED. Nothing routes without consent, and rosters are never disclosed: CONSENT_BEFORE_ROUTE, ROSTER_NOT_DISCLOSED. Shared compute is disclosed in plain words before it happens: SHARED_COMPUTE_DISCLOSED. Results name the device that produced them and replays change nothing: PROVENANCE_NAMES_DEVICE, RESULT_IDEMPOTENT. A job can never carry code or select a backend: KIND_TYPED_ONLY, NO_PAYLOAD_ROUTING. Devices serving others are protected by budgets: COMMUNITY_BUDGETS. Every id above opens the registry — the full list of MUSTs, what verifies each, and how — and a documentation page may never claim stronger verification than a check delivers.

What that leaves to whoever operates your relay

Everything above holds no matter whose relay you use. It is construction rather than policy, so it would still hold on a relay run by somebody you have never heard of.

Three things are not like that, and they are promises by a person or a company rather than properties of the software. Availability — a relay that is down routes nothing; your device and your models are unaffected and the work simply does not arrive. Honest metering — the relay counts what it forwarded, and nothing in the protocol lets you audit that count independently, which is worth saying plainly rather than implying otherwise. And any consent or dashboard screen they host: the approval is enforced on your device, but what you were shown before you gave it came from them.

Ours are written down in the terms and the privacy policy. Somebody else's are theirs to state.

How to tell who operates yours

Your device knows. It paired with a specific origin, and it will tell you which one without asking anybody:

byollm status — under who can use this device, every pairing is listed by the origin it was made with.

byollm.cloud is us. Anything else is that operator's relay, and the right way to think about it is how you think about whoever runs your mail server: they can see that traffic exists, they cannot read it, and they are exactly as reliable as the people behind them.

Running your own, honestly

The relay is MIT licensed and you are welcome to run one. Three things make that honest rather than confusing for the people who use it. Run the conformance kit — @byollm/conformance — so a device that speaks the protocol correctly gets the behaviour it expects from you. Name your relay as your own, because that origin is what a person reads in the command above and it should tell them whose it is. And read TRADEMARKS.md: the code is MIT, the name is not, and a relay we do not operate must not be called BYOLLM Cloud.

"BYOLLM conformant" is a claim about code — that a version passes the published kit. It is not a certification of a running deployment and we do not issue one. The kit can tell you an implementation speaks the protocol correctly; it cannot tell you the operator meters honestly or will still be online tomorrow. Those are questions about people, and a badge that appeared to cover them would be doing the opposite of what this page is for.
The protocol, the daemon, the server library and the conformance kit are open source (MIT) at github.com/oftomorrowinc/byollm. Every claim on this page corresponds to a MUST in the spec with a check that enforces it — read them and run them yourself.

Why a job was refused

Every refusal this product can print, in its own words, with what to do about it — the ones a device owner sees and the ones an app gets back. Search the text you were shown: it is here verbatim.

A refused job is not a failed one: the device said no, deliberately, and said why. Each heading below is the exact sentence you will have seen.

Who sees which. The first set is what the owner of a device reads on their own machine, where naming the exact cause is the whole point. The second set is what an app gets back about somebody else’s device, where it deliberately is not.

no backend on this device is configured and healthy for that job kind

The device is connected, and nothing on it can serve that job kind right now. Run byollm services manage on the device and check the service is selected and signed in — a configured CLI that is signed out is configured and not healthy, which is this refusal rather than an error at sign-in time.

Code: no-capability — programs branch on this; the wording above can change, the code does not.

the job is private to its owner and this device is paired to someone else

The job is private to whoever sent it, and the device answering is paired to somebody else. This is not a setting you can widen: private means the owner and nobody else. If you meant to share the device, the job has to be sent by a person the device’s owner has named.

Code: audience-self-other-owner — programs branch on this; the wording above can change, the code does not.

nothing this device can verify says the job's owner may use it

The device has no record that the job’s owner may use it. Usually the owner was never added, or was added on a different device. The device’s owner adds them; nothing the sender does can change this, which is the point.

Code: not-locally-allowed — programs branch on this; the wording above can change, the code does not.

the app restricted this job to named runners and this device is not one of them

The site restricted this job to particular devices and this is not one of them. That restriction lives in the site’s code, not in your settings — if you own the site, look at the runner allowlist it passes; if you do not, the site meant to do this.

Code: not-in-server-allowlist — programs branch on this; the wording above can change, the code does not.

this service is offered to its owner only (`byollm offer <service> team` to widen)

The service exists and is offered to its owner only. Widen it with byollm offer <service> team on the device. The message names the command because the fix is one line and the device is where it has to be typed.

Code: offer-scope-too-narrow — programs branch on this; the wording above can change, the code does not.

subscription-backed models run their owner's work only — this is a protocol rule, not a setting

This one has no fix, by construction. A model backed by somebody’s personal subscription runs that person’s work and nobody else’s — it is a protocol rule, not a preference, and it holds at every tier. Share an API-metered backend or a local model instead.

Code: subscription-self-lock — programs branch on this; the wording above can change, the code does not.

The backend bills its owner per token and they have not agreed to spend it on other people’s jobs. The owner consents on the device; the sender cannot opt in on their behalf, which is the whole reason the question exists.

Code: metered-no-spend-consent — programs branch on this; the wording above can change, the code does not.

this backend is shared but has reached the spend ceiling its owner set

The owner shared this backend and set a spending ceiling, and it has been reached. It is not broken and it is not revoked — it resumes when the owner raises the ceiling or the period rolls over.

Code: metered-ceiling-reached — programs branch on this; the wording above can change, the code does not.

What do the error codes on a thrown error mean?

The refusals above are about a job somebody declined to run. These are about the call itself, and an app meets them on RelayUnavailable, which carries the relay’s own code alongside a retryable flag.

Branch on the class and on `retryable`, not on the code. The SDK already decides retryability for you, and a site that switched on a code it knew would stop matching the day the relay added a new one — which is why EnqueueRefused is keyed on the status class rather than on a list of codes. The codes are here because they appear in your logs and in error messages, and the thing you searched for should be findable.

bad-request — the request was malformed; repeating it unchanged will not help (HTTP 400)

unsupported-protocol-version — the two ends disagree about the contract itself, not about this request (HTTP 400)

daemon-below-floor — the device’s build is older than this relay will serve. Its request was correct and the sender is old, which is why this is not a permission error (HTTP 426)

unauthorized — we do not know who you are (HTTP 401)

forbidden — we know exactly who you are, and the answer is no (HTTP 403)

revoked — the access this call relied on was taken away (HTTP 403)

not-found — no such job (HTTP 404)

not-ready — claimed, but the site has not sealed the payload yet. Keep asking — the job is legitimately still the device’s until its lease says otherwise (HTTP 409)

too-late — the job is over. Not not-found, which says no such job, and not not-ready, which says not yet — this one says it finished, so stop rather than retry (HTTP 409)

clock-skew — the caller’s clock is too far from the server’s to judge a signature’s freshness. The remedy is the machine’s time, not its keys (HTTP 401)

rate-limited — too many requests; this one is worth retrying later (HTTP 429)

server-error — the relay failed, and it says so rather than blaming you (HTTP 500)

this device serves that kind from more than one service and its owner has not chosen which

The device can serve that kind of work in more than one way, and nobody has said which to use by default. Nothing you send can pick for it: the choice belongs to the device’s owner, who makes it with byollm services manage. Until they do, ask for a specific service rather than the default.

Code: default-ambiguity — this one arrives in a refused result envelope as reason; branch on it, never on the sentence.

this device's default for that kind cannot run work for you

The device has a default for that kind and it cannot run your work. Deliberately, this sentence does not tell you why — the reasons are facts about somebody else’s machine, and a message that distinguished them would let anyone map a stranger’s setup by sending jobs at it. The remedy is the same for every reason behind it: the device’s owner can see which one it was, and only they can change it.

Code: default-unusable — this one arrives in a refused result envelope as reason; branch on it, never on the sentence.

The words

Five nouns do most of the work here, and each means exactly one thing. This page is the reference the rest of the docs are checked against.

Every one of these was, at some point, two words for one idea or one word for two ideas — and each cost somebody a wrong decision on a screen where being wrong mattered. They are written down because a vocabulary that lives only in people’s heads drifts, and the drift always shows up on a consent screen first.

Site

An app that sends work. It registers a domain, holds a keypair, and asks a device to run something on a user’s behalf.

A site is never called a service. It was, once — a heading read “Connected services” above a list of sites, two inches from a card that called them sites — and a dashboard lint now remembers that exact phrasing so it cannot come back by habit.

Device

A computer running the daemon, paired to an owner and approved by fingerprint. Not a “machine” — the protocol, the route, the CLI and the hub API all said device while the dashboard said machine, which is the same defect wearing different clothes.

Hardware is still allowed to be hardware in ordinary prose: an M-series Mac, the electricity it uses, the box under a desk. What is fixed is the name of the paired thing.

Service

Something a device runs — one entry in ~/.byollm/config.json, one row in byollm services. It names its adapter type, its base URL, its model, and the kinds it answers.

A service is a thing a device runs. A site is a thing that sends work. They are opposite ends of the same job, which is precisely why calling one by the other’s name was worth a lint.

Kind

What sort of work a job is — llm.generate or llm.chat. A job carries a kind and data, never a model name, a URL, or a flag. Which model answers is the device owner’s decision, not the site’s.

Two services may answer the same kind. When they do, neither is advertised until defaults names the winner — the daemon withholds the kind rather than guessing, because the wrong guess is the metered one.

Audience and offer scope

Two halves of one question, in one vocabulary: private and team.

Audience is set by the site, per job: who this piece of work may run for. Offer scope is set by the device owner, per service: who this route is offered to. A job runs only where both agree.

The first half of that is the direct lane only. On BYOLLM Cloud a site does not set an audience and passing one is refused. Who may serve a job is the person’s decision, made on their dashboard; the hub derives it from their mapping. This entry described one lane without saying so, and a reader on the other one would have gone looking for a field that is not there — see who owner is. Offer scope is unchanged on both: it belongs to the device owner either way.

There is no third value, and its absence is the point. public — a device offering work to anyone at all — was removed from the protocol before 0.1.0, not deprecated and not merely unrouted by the cloud. It was the one value that skipped the device’s own admission check, and an enum with a value that skips verification is a fail-open waiting for the wiring bug that reaches it. Every value that remains requires the device to verify something: private checks the owner, team checks admission.

`team` is enforced by a signed grant, checked on the device. A teammate’s job arrives carrying a grant the control plane signed for that one job, and your device verifies it against the key it pinned when you paired. Being on the roster is what makes the control plane willing to sign; leaving it stops the next job, with nothing to expire and nothing to revoke locally.

Your BYOLLM id

The name the hub knows you by. Rosters, consents, budgets and the ledger all speak it, and a site that wants to send you work has to use it — a site’s own id for you means nothing here.

You can read it on your account page under Your BYOLLM id, with a copy button, and it is echoed on Connected sites as Connected as.

It names you; it authorises nothing. A site still cannot reach your devices until you connect it on Connected sites — so somebody holding your id can address work to you and cannot deliver it. That is why pasting it into a site’s settings is safe, and why a wrong one is harmless rather than dangerous.

Retired words

self and named were the old audience and offer-scope words; they became private and team before 0.1.0, and the schemas no longer contain them. public went the same way and is listed above rather than here, because its removal is a fact about what the protocol guarantees rather than a rename somebody might still type. backends and routes were the old config shape and are now one services map — a daemon meeting either says so by name rather than failing with a schema error.

byollm allow and byollm disallow are gone. They maintained a list on each device of who could use it, kept by hand and separately from the roster the rest of the product already had. A teammate is admitted now by a grant the control plane signs per job, so there is no second list to keep in step and no way for the two to disagree.

Roster sync was the name for the work that would have carried a central roster down to each device. It was never built and is not needed: the roster is read where the grant is signed, so nothing has to be copied anywhere. Any text promising it — or telling you to read team as an upper bound on who may be served — predates 0.1.0.

Questions people ask

Every question here has its own link — open one and the address bar carries it, so you can send somebody the answer rather than the page.

Short and growing. These are the questions whose answers can be pointed at a line of code; the list gets longer as people ask things we have not written down.

Why does `byollm status` say NOT RUNNING when I think it is running?

Two different things produce that line, and the screen tells you which one it found.

“the process that wrote this device’s heartbeat (pid N) is gone” means the daemon was stopped. The beat file on disk records the process that wrote it, and nothing holds that process id any more.

“nothing has written this device’s heartbeat for …” means the opposite: a process is still there and has stopped beating. That is a wedged daemon rather than a stopped one, and restarting it is the answer to the first and not always to the second.

A third line, NOT REPORTING, is neither: the daemon is running and the hub is rejecting what it sends.

Why is my hosted box running an older version of byollm than the one I just released?

Because a box installs byollm@latest when its pod starts, and only then. That is deliberate: pinning an exact version into the image would make every CLI release a fleet-wide rebuild, so a box picks the new version up at its next pod start rather than the moment a release lands.

Which also means latest — not the version a repository happens to name — is what the fleet is actually running.

Why does the box console refuse `ls`?

Because the console on a hosted box runs a restricted shell with a fixed list of commands, not a general one. It is the fence the page describes: a session lands in it and cannot land anywhere else.

The refusal names what it will accept. If you need something outside that list, that is a request to widen the list rather than a bug in the shell.

Why didn't Ctrl-C stop the command in the box console?

Ctrl-C asks the running command to stop, and a command that has stopped answering may never hear it. The usual reason is a CLI sign-in waiting on a reply that never arrives.

Nothing is stuck for good. The console stops a command that has run too long and the prompt returns on its own — the console’s own banner tells you how long it waits, which is the number that moves when the timeout does.

There is no way to detach from a box console. exit ends this console and the box opens a fresh one; it keeps serving either way. To leave, close the window.

Why won't my config load when two services offer the same kind?

One service offering a kind simply serves it. Two or more and byollm refuses to guess: you set defaults for that kind, and the config will not load until you do.

It refuses loudly in your terminal rather than quietly at job time, because the alternative is a job answered by a service you did not choose, three hops away from anything that would tell you.

Why is my device refusing jobs with “no grant arrived”?

A device paired with a control plane expects a signed grant with every relayed job, including its owner’s own. A job that arrives without one is either a relay that dropped it or a version skew, and a device cannot tell those apart — so it refuses rather than guessing in the open direction.

Can I run byollm without pairing it with anything?

Yes — that is direct mode, and it serves your own work only. With no control plane there is nothing that can tell your device who anybody else is, so it will not run a stranger’s job even if a service is offered to a team.

byollm connect <relay> is what changes that.

I enqueued a job on BYOLLM Cloud and nothing ever happens. Why?

Almost always the pump. On the cloud lane a device claims your job and then waits for you to hand over the payload — if nothing does that within about ten seconds the claim lapses, and the job goes round again until it expires. Nothing errors, which is why this looks like silence rather than a failure.

What the pump is and when to call it — and how to keep a long job alive if the work takes minutes.

Why is my job refused with `slot-unsatisfiable`?

Nobody has chosen what answers it. The person has connected your site but has not mapped a service to the purpose and kind you asked for, so there is nothing for the job to route to. Waiting does not help — somebody has to choose, on their own dashboard.

If the slot is mapped and the device is merely asleep you get a different code, and the difference is the only thing you are told: which of the two you have, and what to do about each.

Where do I put my manifest — is there an API for it?

There isn’t one, and you were right to stop looking. You paste it into the site’s card on Developer Sites, and that is the only door — no registration call in @byollm/server, no endpoint, no CLI.

Where the editor is, and how long an edit takes to reach routing — it is a couple of seconds, not instant, and there is a number you can watch.

Does the library retry for me?

No. There is no backoff, no attempt counter and no retry loop anywhere in @byollm/server. A RelayUnavailable carries retryable so you can tell a draining pod from a bad signature, and the deciding and the waiting are yours.

That, and seven other things nothing told you about — including which refusal is worth trying again later.