Measured, on the machine named below

It runs on a box you already have, and it never phones home.

Two questions decide whether anything gets inside the fence. How big a machine does it need, and what is it going to do to the network. Both are arithmetic, so here they are as arithmetic. Everything below was measured on an Apple M4 with 10 cores, on Python 3.12.12. A speed claim without a machine attached is not a claim.

449 sto compile a 9,511 tag plant, from the folder of exports to a model
0.6 GBthe most memory the compiler asked for, at that same 9,511 tags
3.0 GBthe language model we ship, on disk, on the same box as everything else
20 msa benchmark water network's year of simulated readings, 43 tags, compiled

How much compute

A plant hands over a folder of exports. The compiler reads it, works out what each tag is, what it is attached to and what is simply wrong, and writes a model. Below is how long that takes as the plant gets bigger, measured on plants from 300 to 9,511 tags with every size run in its own process so the memory figure is that size's own.

501002005001k2k5k10k0.010.020.050.10.20.5125102050100200500tags in the plantseconds to compileC-TownHAImeasured slope n^2.232

Log axes, so a straight line is a power law. The measured slope is n^2.232, which is worse than linear and is honest about the duplicate pass: it compares series against each other. At 9,511 tags a whole compile is 449 seconds. The two hollow points are published datasets rather than our own simulator: C-Town water network at 43 tags, HAI testbed at 86 tags, both of which land in well under a tenth of a second. The sweep stops there because building a synthetic plant twice that size needs more memory than the machine has, and a timing taken while a machine is swapping measures the disk.

5001k2k5k10k2050100200500tags in the plantmegabytes the compiler asks formeasured slope n^1.027

What the compiler itself asks for, measured with an allocation tracer in a pass of its own because that tracer more than doubles the runtime and would have spoiled the timing above. It grows as n^1.027, so slightly slower than the plant does, and it tops out at 0.60 GB on a 9,511 tag site. That fits in the spare half of an ordinary industrial PC.

It does not want your cores

Pinned to a single core, with every threading library held to one thread, the same compile takes 2.80 s against 2.81 s unpinned. The difference is inside the noise of the measurement, so the honest reading is that it does not need more than one core at all. That is the useful number, because nobody buys a machine for us. Whatever we need has to fit beside whatever that box is already doing.

9%7%79%reading the clock 9%finding duplicates 0%reading the names 5%describing the values 7%everything else 79%

Where the seconds actually go, from a profile rather than a guess. Time spent inside numpy is pushed up to whichever of our own functions asked for it, because knowing that a fifth of a compile is ndarray.partition tells you nothing and knowing it went on describing the values tells you where to look. Finding duplicates is the largest share and the reason the curve above bends: it compares series against each other, so it grows faster than the plant does.

The part that is a language model

One step reads the paperwork: an instrument index, a datasheet, a 50 page manual. That is the only place a language model runs, it runs on the same box, and it is small. Nothing is sent anywhere, and no article body is ever fed to it, only headings and structured fields. Every number it produces is checked against the fact the data layer computed before it is allowed out.

1252050100gigabytes of model on diskseconds to read the paperworkgemma3:1bqwen3.5:0.8bsmollm2:1.7bqwen2.5:3bllama3.2:3bgranite4:3bgranite4.2:3bphi4-mini:3.8bnemotron-3-nano:4bministral-3:3bgranite4:1bgemma3:4bhf.co/NbAiLab/borealis-open-4b-gguf:Q4_K_Mmistral:7bolmo2:7bgranite4.2:8bministral-3:8b

Local models, measured reading the same paperwork. We ship ministral-3:3b at 2.95 GB, made by Mistral AI in the France under Apache 2.0. It reads the whole set in 80 seconds and phrases a finding in 8.6.

Where the model we ship comes from

We ship ministral-3:3b, made by Mistral AI in France under Apache 2.0. It is also the best reader we have measured on this job, which has not always been true and is worth saying plainly: an earlier version of this page argued that shipping a model from a country our buyers are comfortable with cost us real accuracy, and published the size of that cost. It was not true. The comparison behind it pitted the best Chinese model against one of the weakest models we had, and the trade-off was an artefact of that choice rather than a fact about where models are made. Every model measured stays in the table above, including the ones we did not pick.

What it would have cost to do this in a cloud

Not a competitor's invoice, which we cannot see. A volume, which is multiplication. A plant of 1,469 tags sampled every second is 1,112 GB a year on the wire, a steady 0.28 Mbit a second that never stops. Sampled once a minute instead it is 18 GB, which is the number most people actually have in mind when they say this is cheap. A 20,000 tag site at a second is 15.1 TB a year. Running inside the fence, all of those are zero, and the security assessment has one fewer conduit to argue about.

Which jobs of work this replaces

Seven of them, each with the measurement that backs it and the part we do not do. Two of these are things we are currently bad at, and they are in the list at the same size as the rest.

Work out what each tag is

An integrator or a control engineer opens the drawings and the loop index and writes down, tag by tag, what it measures and what it is attached to. This is the asset model build, and it is the step with no tool behind it.

what it does

On a water network nobody here designed, 39 of 43 tags placed, 93% of quantities and 100% of machines correct against the network's own model.

what it does not

Write the connector for a sector nobody has written one for. With no water connector it places 0 of 43, which is the honest shape of this.

scored 0.866 against 0.526 for name matching, criterion passed

Work out how the plant is wired to itself

Reading the P&IDs and the control narrative to learn which pump feeds which tank, which rectifier serves which line. Where the drawings are stale, asking whoever has been there longest.

what it does

6 of 6 tanks matched to the pump that fills them, on held-out data, from the values alone. The network's own control logic agrees with every one.

what it does not

Tell a pump that fills a tank apart from a pump that merely sits upstream of it. T6 was named and should have been refused.

Find the points that are dead in the field

Nobody, mostly. A tag that has read the same number for a year still appears in the tag list, in the trend and in the report, and is found when someone eventually goes to look at the instrument.

what it does

Seven of the 43 water tags flagged as dead or railed, including both standby pumps' flow and status, from the values alone.

what it does not

Say why it is dead. That is a person with a multimeter.

Find two names for one sensor

Found when two reports disagree, or never.

what it does

Scored on the front page, and this is one of the two criteria that missed.

what it does not

Reach the threshold that was set for it before the run.

scored 0.867 against 0.000 for name matching, criterion missed

Turn raw counts back into engineering values

A person recognises 0 to 27,648 as a Siemens analog word or 4 to 20 as a current loop, and applies the scaling from the datasheet.

what it does

Scored on the front page, and this is the other criterion that missed.

what it does not

Recover a scale for an instrument that has flatlined, or for a point with no sibling anywhere in the plant. Roughly half the misses cannot be solved by anything.

scored 0.671 against 0.000 for name matching, criterion missed

Put every source on one clock

Discovering during an investigation that the historian is in local time without an offset and the SCADA is in UTC, and that an hour in March does not exist in one of them.

what it does

Scored on the front page and it passed.

what it does not

Fix a clock that is wrong rather than differently written.

scored 1.000 against 0.378 for name matching, criterion passed

Say which points cannot be placed, and why

Not a job anyone is given. A tool that cannot place a point usually places it anyway.

what it does

39 of 86 refused on the HAI rig and 1 of 43 on the water network, each with the reason written next to it.

what it does not

Reduce that list. It is the first week of the install, not a defect.

How we know any of this

Everything above is a number, and a number with no interval under it cannot be wrong, which makes it an anecdote with a decimal point. So the criteria on the front page were re-run on fifty held-out plants instead of five, with bootstrap intervals, a paired test against the baseline, and a null hypothesis under the water result. Two of those results are below; all four are in auge_plant/P19_REPORT.md, including the one that failed and why the way I wrote it made failing the only possible outcome.

Refusing is only useful if it refuses the right ones

A supplier that refuses to answer is only worth something if the points it refuses are the ones it would have got wrong. Measured: 0.140 on the aluminium works over 14,844 tags, 0.042 on the grid substation over 12,802 tags, against a limit of 0.5 fixed before the run. In plain terms, when this compiler says it is unsure it is nearly always unsure about exactly the tags it has got wrong. A refusal list of 37 tags is not 37 arbitrary tags, it is close to the 37 a perfect oracle would have picked.

0%25%50%75%100%0%3%6%9%12%how much of the plant it is made to answer forshare of those answers that are wrongAluminium worksGrid substationAluminium excess 0.14 Grid excess 0.04

Make the compiler answer for more of the plant and it has to start including tags it is less sure about. This curve is how wrong the answered set becomes as coverage grows. A system whose confidence meant nothing would be flat. The excess figure is where the area under the curve sits between a perfect ranking of the compiler's own mistakes and no ranking at all: 0 is perfect, 1 is useless.

We built an automatic grader and then threw it away

The explanations need a grader, and the only one that can run inside a fence is a small local model. So we built one, broke 30 real explanations in four known ways, and asked three models from three makers to catch them. qwen3.5:0.8b caught almost every corruption and looked like the clear winner. It also failed 97% of the honest notes: it was not detecting anything, it was answering yes to everything, and a grader that cries wolf catches every wolf. Its real separation between an honest note and a broken one is 0.52, which is a coin. The best of the three, granite4:1b, reaches 0.67, and the two judges that agree most agree on 80% of notes, which sounds respectable until chance is taken out of it and the figure becomes a Cohen's kappa of 0.13. So it is not shipped. A grader that is right two times in three does not belong between a plant and a claim about its equipment. It stays in the repository as a regression test, which is a different and much smaller job. The full write up, including the two faults that turned out to be in our harness rather than in the models, is in auge_plant/P21_REPORT.md.

Are the explanations causal, or only true

The explanations are mostly sufficient (91%, 78%) and mostly not necessary (32%, 22%). What the compiler cites is genuinely enough to reach the answer; the answer would also have been reached without it, because something else carries the same information. That is not a lie and it is not decoration. It is a true statement of one route to the answer, presented as though it were the route. At most 7% of the channels that actually matter go unmentioned, so the explanations over-claim exclusivity rather than substance. The fix is in the evidence writer, not the placement logic: say which route was taken, or say that several agree.

substationsmelter00.20.50.70.9share of sampled explanationsnecessarysufficientboth, a complete account

Exploratory and not preregistered, so nothing here passes or fails. Delete everything an explanation cites and recompile: if the answer does not move, the cited evidence is not what produced it. Then delete everything it does not cite: if the answer survives, the citation was enough on its own. This test found two bugs in itself before it found anything about the compiler, and both are in P20_REPORT.md.

How many hours that is

This is the number every supplier in this category invents, so here is the honest form of it. How long a person takes per tag is not something we can measure, so it is the axis rather than the answer. What is measured is the share: 86.6% of tags placed with asset, quantity and unit all correct, on fifty plants the compiler had never seen, against 52.6% for name matching, which is what a good integrator does by hand with a spreadsheet.

0.512510200122.4244.8367.3489.7minutes a person spends per taghours for one plantby hand, all of itAuge removesname matching removes

One plant of 1,469 tags. Pick your own minutes per tag on the bottom axis. The middle bar is what the measured share removes, and the right hand bar is what a spreadsheet would have removed anyway, so the gap between them is the only part that is ours. The rest is the tags the compiler refuses rather than guesses, and they stay on someone's desk.

Time saved, and where we will not put a number

Every supplier in this category claims hours saved. Ours is held to one rule: a saving is estimated only where a published figure for how long the job takes and a measurement of how much of it we do meet. On two real plants we placed 175 of 212 tags with the right quantity, and refused the rest with a reason. Nobody has published a minutes-per-tag figure that is not a vendor’s (“A 500-tag database that takes two days takes 30 minutes” (Tatsoft, vendor claim)), so the minutes stay a range (Tatsoft’s works out at about two a tag), and a 50-tag spot check is taken off.

tagsminutes a taghours by handhours saved
1 00011713
1 00023326
1 00058365
1 00010167129
10 0001167137
10 0002333274
10 0005833684
10 000101 6671 368

The jobs below take real time, and the console helps with each, but nobody has timed them with and without it, so the hours are shown as what is at stake, not as saved:

What would turn these into estimates is a timed trial in a pilot: the same people, the same jobs, two weeks without the console and two with it, the minutes logged by them, not by us.