You Can't Take a Random Sample from the Future
By Dan Strongin @ 2026-09-29T14:46 (+2)
This is a linkpost to https://whatilearnedinbusiness.substack.com/p/you-cant-take-a-random-sample-from?r=3hfa4&utm_campaign=post&utm_medium=web
How something that knows it is being measured will distort the process or the data when given an unrealistic target, and how distortion relates to what happened with ExploitGym.
Post one in the series ˜When Measuring Makes Things Worse˜. It looked at distortion. Then it covered tampering: reacting to ordinary variation as if it means something. Now I go into a rule to separate ordinary from a signal and what follows from that.
Note: These words are mine. I wrote every sentence and nothing was drafted for me. However I used AI, specifically Claude cowork and Devonthink´s algorithm
- to find articles in my collection,
- for structural feedback on drafts,
- checking sources,
- search,
- and a mechanical pass for punctuations.
- Most of all to be sure that I fully understood the concepts from the ML community as I am writing something from one area of expertise, Quality and Analysis, to another that is not my profession, Machine Learning.
The reason tampering has its own name is that it’s one of two mistakes you can make, and they pull in opposite directions.
The first mistake is treating routine variation as though it were a signal. The number moved, so something must have caused it, so you go and act. This is the scanner. It is also the funnel under rule two. It is what diligence produces when nobody has established what the ordinary range of movement is.
The second mistake is treating a real signal as though it were routine. Something genuinely changed, and you missed it, because you’d got used to the number moving around and this looked like more of the same.
You cannot eliminate both. Any rule for when to act trades one against the other: catch every real change and you will act on a great deal of nothing; never chase noise and you will sit through things you should have caught. There is no setting that avoids both.
What Shewhart built is not a compromise between them. It is a deliberate lean in one direction, and the reason is worth following.
A process has many steps, and every step contributes a little variation of its own. In a training or evaluation run, that includes things as basic as floating-point arithmetic (the same operations on a GPU don’t always combine in the same order, so two identical runs can differ at the bit level) and things as ordinary as how a grader handles an answer that’s almost right. What you observe as the ordinary variation is the accumulated sum of everything like that, seen and unseen. No single one is large enough to find, and none of them is the cause. That’s why chasing an individual result is futile: you’re hunting for a cause that doesn’t exist as a single thing.
A signal is something different in kind. It means one influence has grown large enough to dominate all the rest: to overcome the flow of the many. It needn’t come from outside the process; a step that was contributing its small share can start contributing a large one. But when that happens, there is something to find.
So the thing he built is designed to make signals rare. It is a chart: you plot each result in the order it arrived, with limits worked out from the process’s own past behavior. These are not the specification limits somebody chose. They are a description of what this process does: which is why they’re called natural process limits, and why I’ll keep saying natural limits rather than anything with “control” in it.
The same goes for the chart. Shewhart called it a control chart. Wheeler calls it a process behavior chart, and that’s the name I’d rather use, because the older one suggests the chart controls something, or that a process failing the test has misbehaved. It doesn’t and it hasn’t. The chart describes. It moves away from check this against what we wanted and towards find out what is true: and that’s the distinction the whole argument rests on.
And it is built to make signals rare. The limits are set wide enough that a routine result will almost never be flagged, which means most of what you see is correctly left alone, and it means that when something is flagged you can go looking with reasonable confidence that there’s something there. The cost of that choice is real: a genuine change that isn’t large will not trip it, at least not immediately.
Shewhart’s justification for accepting that cost was economic rather than statistical. Investigating noise is expensive, and it never stops: there’s always another wobble to chase. Missing a real change costs less, because a real change keeps producing signals until it’s caught. So you set the limits to guard against the error that never stops costing you, and accept the one that keeps calling for your attention until you catch it. That’s what I’d most like to hand over, because it changes what a decision rule is for. A rule isn’t there to make you right. It’s there to make the trade explicit, so the same trade gets applied every time: not renegotiated fresh under whatever pressure you’re under that afternoon.
One thing that follows immediately, and I’ll come back to it later: a signal tells you the process changed. It does not tell you the process improved, just because the number moves in a welcome direction.
And the thing that makes tampering worth a whole word: it is not a failure of the measured party. The scanner did nothing wrong. The marble did nothing wrong. The variation those systems produced was ordinary and would have stayed ordinary if left alone. What widened it was the response: mine, just like in the example of the funnel.
Which is why I’ve been careful to keep two failures apart in this post. There is what the measured party does under a demand it cannot meet, which is the thing everyone here is already studying. And there is what the measuring party does in response to ordinary movement, which as far as I can tell is not being studied. My claim is that the second makes the first worse. Tampering widens the distance between the top and the bottom. A wider amplitude makes any fixed demand harder to meet. And a demand that can’t be met is met some other way.
I should be clear about what is already known in the field, because it’s more than I expected when I started reading.
Run-to-run variation on evaluations is documented and quantified. Bouthillier and colleagues showed years ago that benchmark results move substantially with nothing changed but the seed. Madaan’s group and the more recent reproducibility work on reasoning benchmarks have put numbers on it: standard deviations large enough that a good many published comparisons sit inside the noise of their own measurement.
So nobody needs telling that these numbers move. It’s measured, it’s published, and the people doing the measuring are careful about it.
Nor is it only characterized. Madaan’s group compare them plainly: to conclude that one training setup beats another, you would want the performance difference to be larger than the difference seed variance alone would produce. An observed difference weighed against known variation, used to decide whether to believe it.
So the piece I first thought was missing turns out to exist. What I’d say instead is narrower, and it’s about where that comparison gets applied. It governs a choice between two setups, is B better than A, made once, from a finished experiment. What I haven’t found is anything governing the other situation: a process that is running, producing a result this week and another next week, and someone deciding each time whether to touch it. The best evidence I have that this is a real gap rather than my own ignorance comes from a document that was trying to be complete.
In 2024 four researchers (Maxime Riché, Jaime Raldua Veuthey, Harrison Gietz and Edoardo Pona) published an attempt to map the dimensions of AI evaluation practice systematically. It’s a work in progress and they say so themselves; they weren’t satisfied with it and didn’t take it further. But it’s the most thorough attempt I’ve found to lay out what an evaluation is along every axis that matters.
It has three sections. The third is called “Using evaluation outputs.”
Among the properties it lists is precision: whether evaluating several times gives you the same answer. So run-to-run variation is there, named, as a dimension of the practice.
And among the ways it says results can be judged, it offers two kinds: directional, where higher is better, and threshold, where the only question is whether you’re above the line.Those two things never meet. There is nothing in that document, or anywhere else I’ve looked, connecting the fact that a measurement has a variation to the question of what a particular result licenses you to do. Precision is filed as a property of the evaluation. The criteria are filed as ways of judging the number. Nobody has written the sentence that joins them.
That is what I think is missing, stated as precisely as I can manage. Both traditions know that measuring the same thing repeatedly gives you different answers, and both have measured how different. Quality engineering went on to build a rule for what to do about it: one that says when to act and when to leave the thing alone, and applies the same standard to both. As far as I can tell, evaluation research hasn’t gotten there yet.
And I’d point out that “higher is better” is not a neutral way of reading a number. It is an instruction to respond to every movement. Applied to a measurement with a measurement known to move around, it is a formal specification for the funnel.
Since I started writing this I found the paper that comes closest, and it was worth going through carefully, because where it stops is more interesting than where it starts.
Heineman and colleagues at the Allen Institute published a framework last year for reducing uncertainty in language model evaluation. They define noise as the variation in a benchmark score from one training checkpoint to the next, and signal as how far apart different models score on it. The ratio turns out to predict something practical: whether a comparison between two small models will still hold when you scale them up. They ran 465 models over 30 benchmarks to establish it.
They are measuring variation over time, within a single run. Nobody else I’ve read does that.
And they get closer still. Discussing scaling-law predictions, they conjecture that the noise around the model being predicted acts as a floor under the prediction error: if the error you observe is smaller than the noise, it could only have arrived by chance.
That’s the idea. A difference smaller than the variation you already know about carries no information. It’s written down.
Where it stops is what gets done with it. The three interventions they recommend are dropping noisy sub-tasks, averaging across checkpoints, and switching to a smoother metric. All three improve the instrument. Their conclusion addresses benchmark developers, asking them to build better evaluation tools. And the decision the whole framework serves is a one-time comparison: will this small-scale ranking hold at large scale. Nothing in it speaks to the Tuesday when the number arrives.
There’s one more thing, and I raise it because it’s the most concrete thing I have to offer, not as a criticism of careful work.
A process behavior chart is built in two phases. In the first, you take a stretch of results from a period when the process was running steadily and nothing was being changed, and you use it to work out how much this process moves around when nothing is wrong. That gives you your limits. In the second phase you monitor: each new result is plotted against those limits, and the limits do not move. They’re a fixed reference. That’s what makes them able to show you movement away from a reference.
The paper works out its noise figure from the last thirty checkpoints of the run it’s evaluating. That puts both phases in the same stretch of data: the figure used to judge the run is calculated from the run. If the process was not predictable across those thirty checkpoints, the figure is measuring unpredictability along with ordinary variation, not ordinary variation alone. (This is a common error, enabled by how many of the statistical applications calculate limits automatically, and the distinction is not well understood even among many quality practitioners.)
The authors also say, in justifying how many checkpoints they need, that they treat the checkpoint scores as independent of one another. I’d set aside normality here, which comes up whenever this subject does and matters far less than people expect: Shewhart picked three sigma because it worked in practice, not by deriving it from any distribution, and it holds up on data shaped nothing like a bell curve. What it assumes instead is that the numbers you’re using are all telling you about the same thing: the same process, doing what it usually does, the whole time you were collecting them. A number worked out from thirty numbers in a row describes those thirty numbers only if nothing was changing while they were produced, and whether anything was is a question you answer by plotting them in order and looking, before you calculate anything from them.
They have the training curves. They plot them throughout the paper. What isn’t there is that check.
And it isn’t tidiness. That figure is the denominator under every decision in the paper. If the movement was not predictable throughout, real differences get thrown away as noise. If it’s too small, differences that were nothing get acted on. Both cost something, and the second costs twice: because acting on nothing widens the variation in the first place.
So here is what I’d suggest, and it isn’t elaborate.
Take a fixed configuration (same model, same evaluation, same settings) and run it several times without changing anything. Plot the results in the order you got them. What you’re looking at is how much this thing moves when nothing has been deliberately altered.
One thing to watch for, since the order things happen in is where the evidence lives. Those first runs aren’t in any real order: the model is frozen, nothing is developing, and you could shuffle them without losing a thing. That’s the point of them. They aren’t asking what changed; they’re establishing what the measurement does when nothing changes, which is a fact about your instrument rather than about your model.
From that record, work out the limits, commonly called three-sigma limits. The arithmetic is simple enough to do by hand and Wheeler’s short book gives it in a couple of pages; it isn’t the interesting part, and it uses standard factors rather than calculating from an actual distribution. (I will come back to this later).
Then stop calculating. Order starts carrying information now: results are arriving over time, and something might actually have happened between them. Shewhart had a rule for this: whenever an average, range, or histogram summarizes data, the summary shouldn’t mislead you into an action you wouldn’t take if you saw the data in time order.
Each new result gets plotted against the limits you already have, and those limits don’t move. If a result falls within them, you’ve learned something specific (that nothing has changed, that the movement you’re looking at is the movement this thing always makes) and the right response is to leave it alone. If it falls beyond them, there is something to find, and it’s worth the time to go find it.
One result beyond the limits earns an investigation. One result within them earns nothing.
There’s a third use, and it’s the one that makes this worth doing rather than merely prudent. Once you have limits from a steady period, they become the test for anything you deliberately change.
It runs like this. You have a reason to think something will help. You say so first. A hypothesis. You change that one thing, and then you wait: because a process that has just been changed is unsettled, and what it does in the first stretch afterwards is not yet what it does. When it steadies, you read it against the limits you had. Three answers are possible, and all three are worth having. The results moved and stayed moved: something happened, and you can say what you did to cause it. They moved and came back: whatever that was, it wasn’t a real change. They never moved at all: your change did nothing, and you now know that for the price of finding out, rather than believing it for a year. I’ll give you a case of my own, because it’s the clearest one I have.
Years ago I charted weekly sales at a specialty food business. Every week there was a featured item, promoted in the paper: one week a cheese, the next a roast chicken, the next something else. Fourteen weeks, fourteen different promotions.
The chart was flat. Not flat in the sense of no movement: the numbers moved every week, the way numbers do. Flat in the sense that every week fell within the limits, and the limits never had to change. Fourteen deliberate interventions, and not one of them shifted the process.
Nobody had known that. What everybody had was the same conversation every week about whether that week’s item had done well, conducted on the evidence of a number that had gone up or down for reasons that had nothing to do with the item. Some weeks the promotion “worked.” Some weeks it didn’t. Both readings were wrong, and the chart is what made them visible as wrong.
The thing I’d emphasize is that the chart costs nothing. The promotions were already running. The sales were already being recorded. What was missing was plotting them in order against a line, and until that was done there was no way to tell (not a hard way, no way at all) whether any of it was doing anything.
That’s the answer to the cost objection, and I want to be as clear as I can about it. Establishing a baseline does waste a set of runs without producing a better model. The cost of not spending it is much higher. You pay, indefinitely, in response to numbers that were telling you nothing, or by never finding out which of your changes actually worked. The first bill arrives once and you can budget for it. The second arrives every week and you have no way of knowing how much it cost. .
Someone will point out that they already average across seeds. That’s a better estimate of a mean, and it’s worth doing, but it answers a different question, as covered when discussing robust design. An average tells you where the middle is. It doesn’t tell you whether the thing is steady, and it doesn’t tell you what to do on Tuesday.
There is a question that anyone with statistical training will have already asked, and it’s the right question. Where does three come from? It has the look of something derived. And if it was derived from the normal distribution, then it’s wrong for the data in front of you, because eval scores aren’t normal.The scores are bounded: there is the risk of scores piling up or getting lumpy instead of forming a bell curve.
It wasn’t derived. Shewhart picked three sigma because it worked. It was pragmatic. Neave and Wheeler, going back through Shewhart´s 1931 book, put it plainly: no calculation from the normal distribution or any other distribution entered into the choice of the multiplier. Shewhart called three an acceptable economic value and rested it on evidence that it did the job. He did check afterwards that it behaved sensibly under normal conditions, and under a good many other conditions besides: but as they point out, checking that a choice holds up is a long way from deriving the choice from an assumption. That’s the reading that gets reversed. A value chosen and then found to hold widely is a different kind of thing from a value calculated under an assumption, which inherits that assumption’s fragility.
That’s the history. Here’s why it’s also the right choice, and it goes back to the distinction I drew early on between the two kinds of study.
An experimental study is analyzed once. You set it up, run it, look at the result, and decide: one decision, and whatever chance you took of being fooled by it, you took a single time. A real risk of a false alarm is tolerable there, because you’re accepting it once, in exchange for sensitivity.
Observational analysis isn’t like that. Every new result that arrives is a fresh act of analysis, and there is no last one. A risk you’d take without hesitation on a single occasion is a different proposition when you’re going to take it again next week, and the week after, indefinitely. The limits must be set for a rule that never stops running.
That’s what the three-sigma width buys: an acceptable economic value for a rule that never stops running, not a confidence level, and it was never meant to be one.
One more thing about the limits, because it touches on the same worry. They aren’t produced by a distribution at all. Nothing is fitted, nothing is assumed about shape, and at no point does the procedure ask what family the data belongs to. That wasn’t an oversight: Shewhart tried it. Early work assumed some particular frequency function must describe a controlled process, and the normal law was the first candidate. It didn’t hold, generalizations of it didn’t hold either, and by 1939 he had written the search off entirely: all hopes of finding a unique functional form, he says, are blasted.
His reason is the one that still matters. The functional form of a distribution is independent of the order in which the values occurred, and so it cannot be a criterion of randomness. A distribution throws away the sequence, and the sequence is where the evidence is. So the limits come instead from the variation the process itself produced: how much the results actually moved from one to the next, in the order they arrived. Whatever shape the process has is already inside the numbers you computed them from. Which is why the question of whether your data is normal doesn’t come up here: it isn’t a question the procedure ever needed an answer to.
I’ve put this off twice. Here it is.
The objection is that all of this needs a process that holds still, and a model in training doesn’t hold still. Every checkpoint is a different thing from the one before. That isn’t a flaw in training: it is what training is. So there is nothing steady to chart, and the whole apparatus doesn’t apply to the case I’ve been applying it to.
The objection is right about training. I said so earlier and I’m not softening it now. A training run is not a steady process and it never will be. Mid-run, you cannot say what the model can do.
Training is the design phase. You are building the thing. When I introduced robust design I stated clearly that it belongs to the design phase rather than the operating one. The chart is an operating tool, not a design one. Figuring out what the settings should be (which is what design of experiments does, going back to Fisher´s work on experimental design, and used heavily by Taguchi himself) is a different job from watching what a settled process does once it’s running. Nobody who understands this does a control chart on the design phase, because there’s no stable process yet for a chart to describe.
That doesn’t leave the design phase empty: the robust design method is for that phase, and a demanding one: multiple trials, run on purpose, with the noise varied deliberately ..so you can find the settings that move the least when the noise varies.
Look back at the order I gave previously: build it, then get it steady, then find out what it delivers. The objection is that the first step never becomes steady. That’s correct: and it isn’t a problem, because becoming steady was never the first step’s job.
So where does the chart go? After. Once you have a version you’ve stopped adjusting, as close to minimum variation as robust design can get you, you want to know what it actually does once it’s running. That’s what the chart is for: not comparing candidate settings anymore, but recording how the settings you already chose behave when you leave them alone.
In the design phase you use the same model, same evaluation, same settings: run several times with nothing changed. The model is held still on purpose. What moves is the seed, the ordering, the sampling, the phrasing, whatever the grader does that day. None of that is the model changing. That’s what the measurement does on its own.
Suppose you change the data mix and the score goes from 61 to 64. Is it better? You can’t tell. Not from those two numbers. If the same setup, unchanged, gives you 61 one day and 64 the next, then three points is what this thing does with nothing changed and you’ve learned nothing. If it never moves more than one point on its own, those three points are another story. Same two numbers, opposite conclusions, and what decides it isn’t in either of them.
You can’t read a change unless you know how much movement means nothing. And during training you are making change after change after change, asking that same question every time. So the model never holding still isn’t the reason you can’t do this. It’s the reason you need to.
The objection doesn’t stop at training. A deployed agent faces a task distribution that shifts too. People ask for different things this month than last. So what steady process is there to chart?
You chart what comes out, and you let a real change announce itself. If the tasks shift enough to change what the agent delivers, that shows up: that is what a signal is. On the other hand, if they move but the agent goes on behaving as usual, nothing has happened that anyone should act on. Something upstream moving is not by itself a reason to touch anything.
That matters more than it sounds. It means you don’t have to know in advance which conditions count. You don’t have to list the ways deployment might vary, or decide which shifts are the important ones, or keep a set of declared operating conditions and notice when reality has left them. That’s what people usually reach for, and it fails for a plain reason. An aircraft wing has a handful of conditions carrying most of the variation. A deployed model has a great many, and nobody can say beforehand which will turn out to matter. Watching what comes out needs no such list.
And here is the cost, plain as the day. During a training run you cannot say what the model’s capability is. What you have is a history, and a history tells you where something has been, not where it’s going. Every claim built on comparing mid-run numbers (this configuration beat that one, this data mix is better, the curve is bending this way) is a claim about a process that isn’t holding still, measured with an instrument nobody has checked .
That is why there is only one thing worth doing. Run the same configuration several times over and see how far the number moves when nothing has been changed on purpose. That gives you the size of the ordinary movement: and any difference smaller than that is evidence of nothing.
It doesn’t make the comparisons sound. Two configurations three points apart, when the same configuration run twice comes back four points apart, told you nothing about the configurations. You still don’t know which is better. What you know is that you never knew, and that the afternoon spent explaining the three points was wasted. That is less than you might want and a good deal more than nothing.
None of this argues against scaling laws. They predict an aggregate at a scale, they’ve been tested against it, and they were never built to tell you what Tuesday’s number means.
I came across an attempt at an answer to this: don’t score the answer, score the reasoning. Grade the steps. That’s process supervision.
My first observation is that the name covers two different practices. Grading intermediate steps against a score is one. Deciding what the steps should be (what the parts are, what each is for) is the other. Grade the steps and the steps become what gets pushed on, and what you’re reading afterwards is what scores well rather than what happened. Which costs you the reason you were reading them in the first place. That problem has a wider form and I’ll come back to it later.
What needs to be optimized is not the steps but the whole. If a purchasing manager is measured on what he pays for materials he may do his job well in the sense that he gets the price down, but the cheaper material jams the line twice a shift. His number improved. The plant’s number got worse. Nobody did anything wrong, and that’s the point: he was measured on his part, and his part is not the whole.
A supplier’s general manager once described a disturbing version of the same failure to me. A client of his, a large food company he’d worked with for years, outsourced purchasing to an outside firm at a markup that looked lower than what the company had been paying. On paper, it was an unambiguous win. What the arrangement didn’t consider was the cost of logistics: the original supplier had folded it into its delivery contracts. It became a separate charge, billed to the logistics department. Purchasing’s number improved, but logistics got a lot worse. The purchasing managers were treated as heroes. The head of logistics almost lost his job. But nothing had been fixed. The cost hadn’t gone away or even shrunk. It had moved from one line on a spreadsheet to another, and for a year nobody noticed the difference.
That failure is old enough to have a name in my field. It’s called suboptimization, and it’s what happens when you improve the pieces separately: each one gets better against its own measure and the thing they belong to almost always gets worse. The remedy has been known as long as the failure. You improve the system, not the pieces. Deciding what the steps should be is a decision about the whole. Scoring the steps sets each part chasing its own number.
Ought’s agenda post, ‘Supervise Process, not Outcomes,’ focuses on the process steps, not scoring them. Their proposal is that each component gets judged on how well it fills its own role: not on the global score. That’s not a new number to optimize against; that’s deciding what the parts are and what each is responsible for. My objection doesn’t touch it.
They also describe the failure. What they are talking about is suboptimization, described without being named. Which is what I keep running into in the reading I’ve done on LessWrong.
One difference between their case and mine that I don’t want to paper over. They’re concerned with long-horizon tasks where outcomes aren’t available at all: decades-out forecasting, policy, research. No score exists to be gamed, so process is all you have. I’m concerned with the opposite situation: short-horizon tasks where outcomes arrive constantly, cheaply, and get read for more than they contain. Different problems. What the two have in common is the answer, which is that judging the work by the number it produced was never the only option.
Building a process and scoring one are different acts. You can decide what the parts are, what each is responsible for, and how they fit: that’s design. Or you can score a process that is running and read the numbers. That is the operating regime, where you need to know what movement means before you act on it. Confusing them is where you risk scoring a process and believing you’ve built one.