Where Darkbloom’s paid requests come from: OpenRouter and the Darkbloom API
Updated
Short answer: paid requests reach your Mac from two places: Darkbloom’s own API, and OpenRouter, where Darkbloom has been a paid provider since August 20, 2026. On OpenRouter, Darkbloom competes with other providers of the same model, and OpenRouter’s default routing favors the cheapest stable provider. So how much work reaches the network depends on Darkbloom’s price and reliability on OpenRouter, not only on how popular a model is.
Two ways in
- Darkbloom’s own API (api.darkbloom.dev): developers buy credits and call it directly, at the prices in /v1/pricing.
- OpenRouter: Darkbloom is listed as a provider at openrouter.ai/provider/darkbloom. It started there with a free trial in June 2026 and has been a paid provider since August 20, 2026. Darkbloom publishes a model feed for OpenRouter with the same prices as /v1/pricing, including the cached-input price.
- Either way, the request goes through Darkbloom’s coordinator, which picks a Mac (see how Darkbloom routes requests), and you are paid per token at Darkbloom’s price.
How OpenRouter picks a provider
From OpenRouter’s documentation of its default routing:
- It skips providers that have had significant outages in the last 30 seconds.
- Among the stable ones, it picks from the cheapest, weighted by the inverse square of the price. A provider at half the price is about four times as likely to be tried first.
- The others are fallbacks, tried when the first choice fails.
- If a request sets max_tokens, only providers that support an answer that long are used.
- Customers can turn this off and sort by throughput or latency, or name the providers they want.
Prompt caching and repeat customers
- After a request that used the cache, OpenRouter sends that customer’s next requests for the same model and conversation to the same provider, as long as the provider’s cached price is below its normal input price. The link lapses after 10 minutes without a request.
- Darkbloom prices cached input at half the input price on most models, so it qualifies.
- Darkbloom’s cache sits on each Mac. A provider-filed analysis (GitHub issue #1250) measured a 16.4% cache-hit rate for Darkbloom on Qwen 3.8 27B, against 62–92% for most other providers, and proposes sending follow-up turns to the Mac that already holds the conversation.
Darkbloom’s prices on OpenRouter
OpenRouter’s endpoint list as of September 30, 2026, 01:47 UTC, in dollars per million tokens. Darkbloom’s endpoints listed a maximum answer of 32,768 tokens.
| Model on OpenRouter | Providers | Darkbloom input / output | Cheapest other input | Cheapest other output |
|---|---|---|---|---|
| Gemma 4 26B A4B | 14 | $0.042 / $0.22 | $0.06 | $0.20 |
| gpt-oss-20b | 12 | $0.018 / $0.09 | $0.02 | $0.10 |
| Qwen3.6 35B A3B | 10 | $0.05 / $0.70 | $0.10 | $0.90 |
| Qwen3.8 27B | 16 | $0.05 / $2.20 | $0.0249 | $1.78 |
Price isn’t everything
Issue #1250 found Darkbloom with 0.4% of OpenRouter’s Qwen 3.8 27B tokens on September 29, 2026, although it was the cheapest provider for a typical agent request (30,000 tokens in, 1,000 out). On Darkbloom’s side, 77% of Qwen 3.8 requests completed that day, against 99.9% for Gemma. Low cache hits, a shorter maximum answer and failed requests all push OpenRouter traffic to other providers.
What this means for your Mac
- Because OpenRouter weights by price, demand for a model can move quickly when Darkbloom or another provider changes its price.
- Network-wide failures cost everyone traffic, because OpenRouter steps away from providers with recent outages.
- As one provider you can’t change any of that. What you control is your share of Darkbloom’s traffic: keep your Mac verified, loaded, awake and fast.
Sources
- Darkbloom on OpenRouter: openrouter.ai/provider/darkbloom
- OpenRouter default routing: openrouter.ai/docs/guides/routing/provider-selection
- OpenRouter prompt caching and sticky routing: openrouter.ai/docs/guides/best-practices/prompt-caching
- OpenRouter endpoint prices: openrouter.ai/api/v1/models/qwen/qwen3.8-27b/endpoints (and the same path for each model)
- Darkbloom’s OpenRouter price feed: github.com/Layr-Labs/d-inference/blob/master/docs/reference/pricing-model.md
- OpenRouter share and cache analysis: github.com/Layr-Labs/d-inference/issues/1250
- Paid on OpenRouter from August 20, 2026: x.com/gajesh/status/2090932243309170873
How BloomGauge helps
BloomGauge’s network view shows public Darkbloom traffic, capacity and pricing per model, with weekly patterns, so you can see when a model’s demand moves.
Questions
Where do Darkbloom’s requests come from?
From Darkbloom’s own API, where developers buy credits, and from OpenRouter, where Darkbloom has been a paid provider since August 20, 2026. Both go through Darkbloom’s coordinator, which picks the Mac.
How does OpenRouter choose Darkbloom?
By default OpenRouter skips providers with recent outages and picks among the cheapest, weighted by the inverse square of the price, with the rest as fallbacks. Customers can override this. After a cached request it keeps sending the same conversation to the same provider for up to 10 minutes of inactivity.
Why is Darkbloom’s OpenRouter share low for some models?
A provider-filed analysis (GitHub issue #1250) points to a low cache-hit rate, a 32,768-token output limit and failed requests on Qwen 3.8 27B, even though Darkbloom was the cheapest for a typical agent request.
Related
- How Darkbloom decides which Mac gets a request
- Darkbloom model prices: what each model pays per million tokens
- Which Darkbloom model should I run on my Mac?
- Darkbloom “machine_busy” or “your machine is at capacity” while the Mac is idle
Updated 2026-09-29. Still stuck? Ask in #bloomgauge on the Darkbloom Slack or contact us. BloomGauge is independent and not affiliated with Darkbloom.