Skip to content

Services

Local AI and models of your own

A language model is, in the end, a file. Some are so large that the well known names need data centres for them. Others fit on a machine that stands in your own rack. The gap between the two has become small over the last two years, and it does not sit where most people assume.

The large models know the world. Yours knows your business.

What small models have gained is judgement, not knowledge. They pull the same core out of a paragraph, they follow an instruction, they work a case through cleanly. What they lack is the world's stock of memory. Novels, geography, film history, three thousand recipes. In your business there is no use for any of it either. What is interesting is the room that frees up. That is where the things you pay your people for go in.

Two things this does not mean. It does not mean a small model can do everything the largest one can. For open research, for a genuinely hard piece of reasoning, for a rare language, the large one is better, and we say so when it is. And above all it does not mean you have to choose. The choice stays where it belongs, with you and with the individual task. Most businesses end up running both.

Four layers, and we build all four

A model on its own does nothing. It is one layer in a machine that has four, and the benefit only appears when all four fit together. That is why we do not stop at one of them.

  • The control room

    The surface where you see what is running, what is waiting for your approval and what went wrong. It starts things, follows up and keeps a record. This is software we write for your business, not a tool you have to fold yourself into.

  • The agents

    The small programs that do the work. Each with one task, with exactly the rights it needs for it, and with a record that shows what it based itself on.

  • The model

    This page

    The judgement underneath. This is where a choice sits that most people do not know they have. A public model from the cloud, a provider in Europe, or your own in the building, which knows nothing but your business. All three are fine, and on the third one we go particularly far.

  • The infrastructure

    The server all of it stands on. Open source, in your building or in our rack, with backups, monitoring and a rehearsed way back. Without this layer, digital sovereignty is a statement of intent.

All of it on request. You take every layer from us or bring your own. Anyone who already has a server keeps it. Anyone who only needs a model gets only that. What we do not build is a layer that cannot be swapped out later.

Two lists, and the whole difference lies between them

You trade knowledge of the world for knowledge of your business. It is a good trade, and this is the fastest way to read it.

What a model like that does not know

  • who scored the goal at the Wankdorf Stadium in 1954
  • what chapter twelve of The Magic Mountain is about
  • the history of the Republic of the Congo
  • which films a director made in the eighties
  • how to get a sauce hollandaise right

What it knows once we are done

  • that item 40-2231 was discontinued in March and 40-2240 replaces it
  • that a quote above 50,000 euros goes to Ms Krüger before it goes out
  • what your shop floor means when it calls a part crosswise, and why the delivery note then looks different
  • that customer Meier wants a delivery date without a time window, every time
  • your fifteen years of service reports, so it recognises the fault nobody remembers any more

You do not pay any of your staff for the first list either. For the second one you do.

Three ways your knowledge gets into the model

Knowledge does not get in on one path but on three, and they cost very different amounts. We start at the top and go down only as far as it pays.

  1. Telling it

    The fastest way, and the one most people underrate. Your rules, your tone, your forbidden phrases and a handful of genuinely good examples sit there as text that a person reads and edits. A change is made in an hour and takes effect at once. Everything still being argued about in the business belongs here. An argument can be had in a text, not in a trained model.

  2. Looking it up

    For everything that changes. Price list, stock, yesterday's contract, this morning's minutes. The model gets a search over your own documents and looks things up instead of guessing, and it names the place the answer came from. A price that changes on Monday has changed on Monday, not at the next training run.

  3. Training it in

    For what always holds. How you write, what your forms look like, which steps a type of case has, where the line between two cases runs. That goes into the model itself, and afterwards it no longer needs the long instruction. It gets faster, cheaper to run and more reliable in its own field. For that we hang an extra layer onto an open model instead of building a new one. That is the difference between a weekend and half a year.

Anyone who promises to train your knowledge into the model usually means the second step. We tell you which one we are taking and why, and the most honest answer is usually a mixture of all three.

It gets better while you work with it

A model that stays as it was after the roll-out is a snapshot of your business. Your business is not one. So the same cycle you know from quality management hangs off it, except that here agents turn it.

Every correction is an example
When somebody reworks a draft before it goes out, the difference between before and after is the most valuable information in the whole system. It gets captured instead of disappearing into the sent folder.
An agent collects, a person decides
Proposing is done by machine, taking it up happens after an approval. A system that quietly tops itself up unsupervised will eventually learn the mistake as well.
It is tested against the old model
With cases that were not in the training, and judged by your people. If it does not get better, the old one stays in service. That is what a test is for.
What counts as success is written down first
Without that number every result is a matter of taste, and then whoever says loudest that it got better wins.

What you get out of it

Data protection is the reason most people ask. It is not the biggest one.

Nothing leaves the building
No provider in between, no transfer to another country, no terms of use that change next quarter. What you promise your own customers about their data you can back up with an address instead of a reference to contract clauses.
The cost is a purchase, not a meter
You pay for hardware and electricity, not for every single request. Working with it a lot is not punished. Ten thousand documents overnight cost you a night and no invoice.
It does not go down because something else does
No queue at lunchtime, no status page of somebody else's, and no discontinued model version that rebuilds your process overnight. What runs today runs the same way in two years, if that is what you want.
It runs without internet
On a building site, on the shop floor, out in the field, in a plant behind its own firewall. For some businesses that is the only reason AI comes into question at all.
The model belongs to you
It is a file. You can back it up, copy it, put it on another machine and carry on there. What you trained into it does not belong to a provider, and your knowledge does not become training material for anybody else.
It stays checkable
What the model has learned sits in rules and examples a person can read. A system whose behaviour nobody can explain any more does not belong in a business, however well it answers.
But it is a purchase
A machine with a suitable card, electricity, a place with cooling and somebody to run it. That has to pay against the monthly cost of an account, and where the balance tips is something we work out with you before anything is ordered.

How a project like this runs

Six steps, and the first one is where most attempts fail.

  1. Cut the task down

    Not introducing AI, but one task that happens often enough, is clear enough and eats time today. Plus the number we measure afterwards to see whether it worked.

  2. Collect examples

    From real cases, not invented ones. Where personal data is in them, it comes out first. Five hundred clean examples beat five thousand mediocre ones, and that is the whole craft of this.

  3. Pick the base model

    By task and by hardware, not by how well known it is. There are several open families to take one from, Llama, Qwen, Mistral or Gemma among them. Which of them runs at your place is a question you measure rather than believe, and it can be swapped later.

  4. Train

    We hang an extra layer onto the model instead of computing a whole new one. That is the affordable route, it takes hours rather than months, and it can be undone.

  5. Check

    Against the untouched model and with cases that were not in the training. Your people judge it, not us. If it does not get better, the training was not worth it, and we say so.

  6. Put it into service

    On your hardware or on ours, wired into the place the work turns up, with backups, monitoring and a way back to the previous state.

From our own operation

We decided against training here ourselves

This blog is written by a team of eleven agents. The rules they write by, our tone, our forbidden phrases, our good and our bad examples, sit as text in a database. A person changes them in an hour, and the next draft follows the change. We could have trained a model on all of that instead. We did not, because a school you can read and argue with in an afternoon is worth more for this purpose than a model you can only recompute. That is the first of the three steps above, and for our case it is the right one.

What we do not have, we say as well. We are not running a local model in production at the moment. We are building it for ourselves first, the way we did with our own infrastructure, and once numbers about it stand here, they will carry their date.

It is not worth it for everyone. If a task comes up ten times a month you do not need your own hardware, an account with a provider is cheaper and quicker to have. If there is nobody to look through the examples, you do not get a good model, you get a fast wrong one. And if you need open research across the whole world, a large model is better at that. Then we build both and send each task where it belongs.

Where the model stands is a decision of its own, and it is reversible. In your rack, in ours, or with us first and with you later, once the thing has proved that it carries. The three steps for that are on the home page. And what is said here does not automatically apply to this website. What our own site sends to a model, and where, is in the privacy policy.

Does any of this match what you have in mind?

Then let us talk, even if you are not yet sure what exactly you need.

Discuss a project