Recurring cost, forever
The bill scales with usage, not with value. Success makes it worse, and the price is set by someone else.
Distiller Cloud turns a large open teacher model into a small, specialised LLM — trained on a synthetic dataset built for your task, delivered as a file you download. Run it on your Mac, your RTX box, or your own servers. No ML expertise required.
The problem
Every request leaves your infrastructure, gets billed by the token, and runs on a model far larger than your task needs. You optimise prompts for a system you don't control, can't inspect, and can't take with you.
The bill scales with usage, not with value. Success makes it worse, and the price is set by someone else.
Every prompt is a disclosure. Retention terms change, subprocessors change, and your compliance team notices.
A trillion-parameter generalist to classify support tickets. You pay for capability you'll never use.
How it works
Four steps in a guided interface. No notebooks, no CUDA, no training loop to babysit.
Describe your task and pick an open teacher model. The interface walks you through scope, tone and output format — no hyperparameters to guess.
We generate a synthetic dataset, split it into training and validation sets, and run the distillation. You watch the loss curves; we handle the infrastructure.
You get the weights in standard formats — GGUF, safetensors — with the evaluation report. It runs on Apple Silicon, NVIDIA RTX, DGX Spark or your own servers.
Within 48 hours your dataset, your run and your model are wiped from our infrastructure. Not archived. Not anonymised. Deleted.
What you get
The model is a file. Download it, copy it, deploy it, keep it after you stop being our customer. No API key in the critical path, no vendor lock-in, no per-token bill.
Your data and your model are destroyed within 48 hours — automatically, not on request. Nothing stays on our servers because keeping it isn't part of the product.
A model sized for one task runs on hardware you already own, answers in milliseconds without a round trip, and draws a fraction of the energy of a frontier API call.
Dataset generation, train/validation splits, evaluation — handled. If you can describe the task in writing, you can produce a working model.
Distil from mature open-weight models. You know exactly what your student learned from, and the licence travels with you.
Every run ships with validation metrics and failure examples, so you can decide whether the model is good enough before it goes anywhere near production.
Why now
Apple Silicon, RTX cards and DGX Spark put serious inference on desks that had none three years ago. The runtime is already paid for.
Open-weight teachers are now strong enough to distil from, with licences that permit it. Two years ago the quality gap made this pointless.
Inference is the recurring cost of AI, not training. A task-sized model turns a metered expense into a fixed one.
GDPR and the EU AI Act make data locality and traceability a board-level question. A model that runs inside your perimeter is the shortest answer.
Zero retention
It's easy to promise privacy and keep the data anyway. So we designed the service around not having it: your dataset, your training artefacts and your model are removed from our infrastructure within 48 hours of the run finishing. There is no archive to subpoena and no backup to leak, because retention was never part of the product.
Your task description and generated dataset exist only for the duration of the job.
You download the weights and the evaluation report. From here, the model is entirely yours.
Dataset, checkpoints and final model are deleted automatically. Nothing to opt out of.
The only thing we keep is your email, because you asked us to tell you when we launch.
Early access
We're onboarding the first cohort in small batches so we can actually talk to every team. Sign up now and you get in first, with free distillation credits for the beta and a direct line to the people building it.
No countdown, no fake seat counter. We open the door when the product is good enough, and you'll be among the first to know.
Questions
Yes. The 48-hour window is about our servers, not your download. Once you've pulled the weights, the file is yours permanently — it keeps working if you stop being a customer, and it keeps working if we disappear.
On one well-defined task, a properly distilled small model gets close to its teacher, and it beats a generic large model that has no context about your domain. On open-ended general reasoning it does not — and we'll tell you when distillation is the wrong tool for your problem.
A recent Mac with Apple Silicon, an NVIDIA RTX card, a DGX Spark or a modest server is enough for the model sizes we target. We ship standard formats — GGUF and safetensors — so the usual local runtimes work out of the box.
Pricing isn't final, which is exactly why we're asking early subscribers what a fair price looks like. The model is per distillation run, not per token: you pay to produce the model, then inference is free because it runs on your hardware.
Your dataset, temporary training artefacts and resulting model are deleted from our infrastructure within 48 hours of a completed run. We do not use customer data to train our own models or share it with third parties.
No ML expertise is required to create the model. You need to describe the task and assess examples from your domain; the guided workflow handles the technical training steps and delivers a model in standard formats.
An API rents a general-purpose model and charges for each use. Distiller Cloud produces a smaller model for a defined task that you download and run on hardware you control, with no recurring per-token bill.
Appearance