Issue 01 — a VPS for language models · waitlist open

Your own instance.
No reset timer.

One email when instances open. No newsletter, no drip sequence.

How it works

A private LLM instance on EU hardware, for a flat monthly fee. Nothing counts down and nothing resets — the only limit is how many requests you run at the same moment, and you set that with one slider.

No token meter.
No rolling usage window.
No request cut off halfway
through your task.
EU hardware, EU jurisdiction.

01 / The problem

The ceiling isn't the bill. It's the clock.

A flat plan you can still get locked out of isn't flat.

Today — flat plan, rolling window

09:0014:20 — limit reached19:20

You're mid-refactor when the window closes. Nothing is wrong with your usage — you just wait five hours, or open your wallet somewhere else.

On Solheim — concurrency only

09:00no ceiling to reach19:20

The only limit is how many requests run at the same moment. Send more and they queue for milliseconds — the day never stops and the invoice never moves.

  1. 01

    The cap

    Flat-fee assistant plans still meter you through a rolling window. The work stops on someone else's schedule.

  2. 02

    The jurisdiction

    Your code and your customers' data go through a US-jurisdiction provider, and your DPA has to explain it.

  3. 03

    The forecast

    Per-token billing makes next month a guess. Capacity you rent by the month is a line item you can plan around.

02 / The idea

If you've bought a server, you know this product.

One project gets one VPL: one endpoint, one key, one number to size it.

A slider sets how many instances the project runs, and an instance is one request in flight. Three instances means three things at once: your agent, your editor, your test loop. The API is OpenAI-compatible, so Cline, Roo Code, VS Code BYOK and your own backend all talk to it unmodified.

  1. 01

    On a VPS — vCPU + RAM

    Instance count

    One slider, one dial. Inside that size, run as much as you like, all day.

  2. 02

    On a VPS — no request quota

    No usage window

    Nothing counts down and nothing resets. You are never locked out mid-task.

  3. 03

    On a VPS — OS image

    Open-weight model

    Pick and pin the model. Swap it by changing one string, no migration.

  4. 04

    On a VPS — one server, one host

    One project, one VPL

    Keys, usage and limits stay per project — nothing is shared by accident.

  5. 05

    On a VPS — datacenter region

    EU-only region

    Nothing leaves EU jurisdiction — GDPR and AI Act posture by construction.

03 / Migration

Two minutes, two fields.

Nothing to install, nothing to rewrite.

  1. 01

    Create a project

    A project is one VPL: its own endpoint, its own key, its own usage. Name it after the thing it serves.

  2. 02

    Set the instance count

    One slider decides how many requests run at once. Move it up for a week of parallel agents, back down after. That is the whole capacity model.

  3. 03

    Point your tool at it

    Base URL and key into Cline, Roo Code, VS Code BYOK, or any OpenAI SDK. Nothing else in your setup changes.

any openai-compatible client
export OPENAI_BASE_URL=https://api.solheim.eu/v1
export OPENAI_API_KEY=slh_live_9f3c…

# beyond your instance count, calls queue.
# they do not 429 and they do not cost extra.

04 / Sovereignty

EU-hosted, all the way down.

Sovereignty is an architecture, not a checkbox.

  1. 01

    Infrastructure

    GPUs at EU-owned providers in EU datacenters. No US hyperscaler in the serving path, so no CLOUD Act exposure to write into your DPA.

  2. 02

    Models

    Open weights under permissive licenses, pinned per instance and run by us. Your prompts are not training data — not ours, not anyone's.

  3. 03

    Routing

    Requests never leave the region, and this page loads nothing from a third-party origin — no US CDN, no analytics, no fonts phoning home.

The library

  • 01

    Qwen2.5 Coder 32B

    qwen2.5-coder-32b

    128k context
    Apache 2.0
    coding

    The default for day-to-day work: diffs, tests, refactors and tool-calling agents.

More images as capacity lands. On the hardest agentic refactors, frontier closed models are still ahead — we'd rather say so than oversell.

05 / Rate card

One dial: how many at once.

Instances × a flat monthly rate. No tokens, no tiers, no overage.

3
Context window
identical at every size
Models
all of them, included
Beyond your count
queues, never 429

€75 /month · 3 × €25

Join the waitlist at this size

Indicative only — no numbers are final until we've validated real concurrency. Waitlist signups get founder pricing at launch, and discounts belong on annual prepay and instance volume, never on concurrency.

06 / Readership

Who it's for.

Same product, same dial, same price.

  1. 01

    Developers running coding agents

    You live in Cline, Roo Code or VS Code and you hit the window most weeks. Three instances is an agent, an editor and a test loop running at once — for as long as the day lasts.

  2. 02

    Founders who need an inference backend

    Your product needs a private LLM behind it, with a bill that doesn't move when traffic does. Point your backend at the same endpoint and size it the same way.

It is one product, not two. What changes between these is what you point at the endpoint — not the pricing, not the API, not the machine.

One flat fee.
One EU home.

We're onboarding in small batches while capacity comes online, EU-hosted from day one. Waitlist signups get founder pricing.

Stored in the EU · deleted whenever you ask.