Can you run several Darkbloom models at once?
Updated
Yes: by default the Darkbloom provider can keep up to three models loaded at once. Whether it should is another question. For most Macs one well-chosen model earns more than several. Facts below come from Darkbloom’s open-source docs and issues (links at the end); memory figures are worked out from Darkbloom’s own load formula.
How it works
- Advertised models: the provider offers the models in enabled_models in its config, or every model it can serve if that list is empty. darkbloom start --model lets you pick, and the flag can be repeated.
- Loaded models: up to max_model_slots at once (default 3). Only a loaded, warm model gets requests. The base reward needs one model loaded, and which one doesn’t matter.
- Preloading: on every start the provider loads preload_models if set, otherwise the selected models, within the slot limit and free memory. Models that don’t fit load later on demand.
- Idle unloading: with the shipped setting, a model unloads after 60 minutes without requests (idle_timeout_mins; 0 keeps it loaded).
- Darkbloom can also move models itself: its warm-pool controller asks idle Macs to load models they advertise and already have on disk.
- darkbloom switch (0.9.10 and later) replaces the whole selection on a running provider, without a restart.
What a second model costs in memory
Each loaded model adds its full weights (with a 20% loading allowance). A working reserve for activations is charged once, not per model, plus at least 1 GiB for concurrent requests. By default the provider also holds back 4 GB and uses at most 90% of memory, and macOS and your apps need their share. Free memory needed to load each set, from Darkbloom’s formula:
| Models loaded together | Free memory needed | Smallest Mac to try |
|---|---|---|
| gpt-oss-20b alone | 18.0 GiB | 24 GB (tight) |
| Gemma 4 26B QAT alone | 23.9 GiB | 32 GB |
| Qwen3.6 35B A3B alone | 30.3 GiB | 36 GB |
| gpt-oss-20b + Gemma 4 26B QAT | about 37.5 GiB | 64 GB (48 GB too tight) |
| gpt-oss-20b + Qwen3.6 35B | about 43.8 GiB | 64 GB |
| Gemma 4 26B QAT + Qwen3.6 35B | about 47.7 GiB | 64 GB (tight) |
| Qwen3.5 35B + Qwen3.6 35B | about 53.7 GiB | 96 GB |
| gpt-oss-20b + Gemma 4 26B QAT + Qwen3.6 35B | about 61.3 GiB | 96 GB |
Why one model usually earns more
- Gemma needs the Mac to itself. Public Gemma 4 requests only go to Macs whose whole advertised list is Gemma 4. Add any other model and the Mac gets no public Gemma work, even with Gemma loaded.
- Less room for concurrent requests. The cache for requests in flight comes out of whatever memory the weights leave, so a second model means fewer requests at once for the first.
- Models serving at the same time share the Mac’s memory bandwidth. Darkbloom staff advised providers in the Darkbloom Slack in late September 2026 not to load several models, because two models serving at once perform worse.
- Load churn. In GitHub issue #789, a 128 GB Mac advertising five models on three slots went through 692 load and unload cycles in 49 hours. Rarely used models kept pushing out the busy ones, which then had to reload.
- Cold models lose the race. Routing charges a cold or just-loaded model a large time penalty, so a model that keeps getting unloaded and reloaded wins few requests.
Practical combinations
- 24–36 GB: one model. Nothing else fits next to it.
- 48 GB: one model. Gemma 4 26B QAT alone is the usual steady earner; a single Qwen 35B is the alternative.
- 64 GB: one model is still the simple choice. If you want two, pair models that don’t include Gemma, such as gpt-oss-20b with a Qwen 35B, and check memory pressure stays low.
- 96 GB and up: two or three models fit, but the Gemma rule and shared memory bandwidth still apply. Many large Macs do better running one model and switching when demand moves.
- Keep other models downloaded, not loaded. With darkbloom switch, changing models takes a drain and a load, not a restart.
- After any change, run darkbloom status and confirm the model you expect is loaded and serving.
Sources
- Darkbloom hardware requirements (free memory at load, several resident models): github.com/Layr-Labs/d-inference/blob/master/docs/provider/hardware-requirements.md
- Memory model and load formula: github.com/Layr-Labs/d-inference/blob/master/docs/architecture/hardware-support.md
- CLI reference (start, switch, max_model_slots, preload_models, idle_timeout_mins): github.com/Layr-Labs/d-inference/blob/master/docs/provider/cli-reference.md
- Routing and dedicated models: github.com/Layr-Labs/d-inference/blob/master/docs/architecture/routing.md
- Five models on three slots: github.com/Layr-Labs/d-inference/issues/789
- Staff advice on one model: Darkbloom Slack, September 27–28, 2026.
How BloomGauge helps
BloomGauge’s Manager runs one model at a time. With Manager on, it holds the model that has paid best on your Mac over the last 30 days (or, on a new Mac, what pays best on Macs with the same chip and memory), and moves only when at least five Macs like yours have clearly earned more on another model for two hours. Switches go through Darkbloom’s own CLI and must pass memory, idle and temperature checks. It starts observe-only.
Questions
Can Darkbloom run more than one model at a time?
Yes. The provider keeps up to max_model_slots models loaded (default 3), memory permitting. Each extra model adds its full weights to memory, and Gemma 4 only gets public work when it is the only model the Mac advertises, so one model is usually the better choice.
How much memory do two Darkbloom models need?
Worked out from Darkbloom’s load formula: about 37.5 GiB free for gpt-oss-20b plus Gemma 4 26B QAT, about 44 GiB for gpt-oss-20b plus a Qwen 35B, and about 48 GiB for Gemma plus a Qwen 35B. In practice that means a 64 GB Mac or larger.
Should I advertise several models on Darkbloom?
Usually not. Advertising Gemma with anything else loses public Gemma work, a second model takes memory from concurrent requests, and advertising more models than slots can cause constant loading and unloading. Pick one model, keep others downloaded and switch when demand moves.
Related
- Which Darkbloom model should I run on my Mac?
- Darkbloom model won’t load: not enough memory
- How Darkbloom decides which Mac gets a request
- Darkbloom 0.9.10: live model switching and on-demand loading
Updated 2026-09-29. Still stuck? Ask in #bloomgauge on the Darkbloom Slack or contact us. BloomGauge is independent and not affiliated with Darkbloom.