Distiller Cloud
Skip to main content
Private beta — opening soon

Own your model.
Not a subscription to someone else's.

Distiller Cloud turns a large open teacher model into a small, specialised LLM — trained on a synthetic dataset built for your task, delivered as a file you download. Run it on your Mac, your RTX box, or your own servers. No ML expertise required.

One email at launch. No newsletter, no sharing with third parties. Unsubscribe in one click.

  • Your weights, downloaded and yours to keep
  • Data and model destroyed within 48 hours
  • Runs locally — no per-token bill

The problem

Cloud inference is a lease you never stop paying.

Every request leaves your infrastructure, gets billed by the token, and runs on a model far larger than your task needs. You optimise prompts for a system you don't control, can't inspect, and can't take with you.

Recurring cost, forever

The bill scales with usage, not with value. Success makes it worse, and the price is set by someone else.

Your data leaves the building

Every prompt is a disclosure. Retention terms change, subprocessors change, and your compliance team notices.

Oversized for the job

A trillion-parameter generalist to classify support tickets. You pay for capability you'll never use.

How it works

From a task description to weights on your disk.

Four steps in a guided interface. No notebooks, no CUDA, no training loop to babysit.

  1. 01

    Configure

    Describe your task and pick an open teacher model. The interface walks you through scope, tone and output format — no hyperparameters to guess.

  2. 02

    Distil

    We generate a synthetic dataset, split it into training and validation sets, and run the distillation. You watch the loss curves; we handle the infrastructure.

  3. 03

    Download

    You get the weights in standard formats — GGUF, safetensors — with the evaluation report. It runs on Apple Silicon, NVIDIA RTX, DGX Spark or your own servers.

  4. 04

    Everything is destroyed

    Within 48 hours your dataset, your run and your model are wiped from our infrastructure. Not archived. Not anonymised. Deleted.

What you get

Small, specialised, and genuinely yours.

You own the weights

The model is a file. Download it, copy it, deploy it, keep it after you stop being our customer. No API key in the critical path, no vendor lock-in, no per-token bill.

Zero retention by default

Your data and your model are destroyed within 48 hours — automatically, not on request. Nothing stays on our servers because keeping it isn't part of the product.

Local inference, low latency

A model sized for one task runs on hardware you already own, answers in milliseconds without a round trip, and draws a fraction of the energy of a frontier API call.

No ML expertise required

Dataset generation, train/validation splits, evaluation — handled. If you can describe the task in writing, you can produce a working model.

Open teacher models

Distil from mature open-weight models. You know exactly what your student learned from, and the licence travels with you.

Evaluation you can read

Every run ships with validation metrics and failure examples, so you can decide whether the model is good enough before it goes anywhere near production.

Why now

The pieces only just landed — all four at once.

Capable local hardware is everywhere

Apple Silicon, RTX cards and DGX Spark put serious inference on desks that had none three years ago. The runtime is already paid for.

Open models grew up

Open-weight teachers are now strong enough to distil from, with licences that permit it. Two years ago the quality gap made this pointless.

The energy maths changed

Inference is the recurring cost of AI, not training. A task-sized model turns a metered expense into a fixed one.

Regulation pushes on-premise

GDPR and the EU AI Act make data locality and traceability a board-level question. A model that runs inside your perimeter is the shortest answer.

Zero retention

We delete everything within 48 hours.

It's easy to promise privacy and keep the data anyway. So we designed the service around not having it: your dataset, your training artefacts and your model are removed from our infrastructure within 48 hours of the run finishing. There is no archive to subpoena and no backup to leak, because retention was never part of the product.

  1. 0h

    Run starts

    Your task description and generated dataset exist only for the duration of the job.

  2. ~2h

    Model ready

    You download the weights and the evaluation report. From here, the model is entirely yours.

  3. 48h

    Everything wiped

    Dataset, checkpoints and final model are deleted automatically. Nothing to opt out of.

  • No training on your data — ever, for any purpose
  • No third-party sharing, no data brokers, no ad tech
  • Deletion is automatic, not a support request
  • EU infrastructure, GDPR-aligned by construction

The only thing we keep is your email, because you asked us to tell you when we launch.

Early access

Be there when it opens.

We're onboarding the first cohort in small batches so we can actually talk to every team. Sign up now and you get in first, with free distillation credits for the beta and a direct line to the people building it.

  • First access when the beta opens
  • Free credits for your first distillation runs
  • A say in which teacher models and formats ship first

No countdown, no fake seat counter. We open the door when the product is good enough, and you'll be among the first to know.

One email at launch. No newsletter, no sharing with third parties. Unsubscribe in one click.

Questions

Before you sign up.

Do I really keep the model after the 48 hours?

Yes. The 48-hour window is about our servers, not your download. Once you've pulled the weights, the file is yours permanently — it keeps working if you stop being a customer, and it keeps working if we disappear.

How good is a distilled small model, honestly?

On one well-defined task, a properly distilled small model gets close to its teacher, and it beats a generic large model that has no context about your domain. On open-ended general reasoning it does not — and we'll tell you when distillation is the wrong tool for your problem.

What hardware do I need to run it?

A recent Mac with Apple Silicon, an NVIDIA RTX card, a DGX Spark or a modest server is enough for the model sizes we target. We ship standard formats — GGUF and safetensors — so the usual local runtimes work out of the box.

What does it cost?

Pricing isn't final, which is exactly why we're asking early subscribers what a fair price looks like. The model is per distillation run, not per token: you pay to produce the model, then inference is free because it runs on your hardware.

Do my data stay private during training?

Your dataset, temporary training artefacts and resulting model are deleted from our infrastructure within 48 hours of a completed run. We do not use customer data to train our own models or share it with third parties.

Do I need to know machine learning or how to code?

No ML expertise is required to create the model. You need to describe the task and assess examples from your domain; the guided workflow handles the technical training steps and delivers a model in standard formats.

How is this different from using an LLM API directly?

An API rents a general-purpose model and charges for each use. Distiller Cloud produces a smaller model for a defined task that you download and run on hardware you control, with no recurring per-token bill.

Appearance