How to Validate Run, Idle, and Down Signal Accuracy Before Rollout

Validate machine state accuracy on four machines before you trust it on a hundred.
Every monitoring approach infers states from a signal, and every approach can be wrong in specific, predictable ways. A machine idling with its hydraulics energised can read as running. A slow cycle can read as a stop. A warm-up can read as production. None of that argues against monitoring. It argues for a short, structured accuracy test, numeric pass criteria agreed in advance, and only then a rollout. This guide sets out that test, and what it can and cannot validate.
Signal accuracy validation at a glance
- Test before scale. Two weeks on four representative machines beats a fleet-wide rollout you have to re-baseline.
- Ground truth comes from known events, not from a second sensor that can be wrong the same way.
- The failure modes are specific: false running, missed micro-stops, warm-up counted as production, setup counted as downtime.
- Write numeric pass criteria before you see results, so the threshold is not negotiated afterwards.
- Sample across shifts, not just across machines. The same crew running the same changeover shows a 57% shift-to-shift spread in Guidewheel's 2026 Factory Uptime Report.
| What you are testing | Failure mode you are hunting | Acceptable result |
|---|---|---|
| Running vs idle | Machine energised but not producing, read as running | State matches known activity on at least 95% of sampled minutes |
| Down detection | A real stop missed, or registered late | Every stop over the threshold detected, inside an agreed lag |
| Micro-stops | Short stops absorbed into running time | Stops at or above your defined minimum are captured consistently |
| Cycle counting | Cycles missed or double-counted at speed or on product change | Count within an agreed tolerance of a manual count |
| Warm-up and setup | Non-productive energised time counted as production | Classified to the state the plant wants, consistently |
| Product mix | Accuracy holds on the light product as well as the heavy one | No material accuracy drop across the product range tested |
Define the states before you measure anything
Most accuracy disputes are definition disputes wearing a technical costume. Two people disagree about whether the number is right because they never agreed what the state means.
Write down, on one page, what each state means for this plant.
- Running. Producing sellable output. Decide explicitly whether a machine cycling on scrap counts.
- Idle. Powered and available but not producing. This causes the most argument, because on machines that keep auxiliaries energised it looks like running from the electrical side.
- Down. Not available. Split planned from unplanned if your reporting needs it, and say who decides which.
- Warm-up. Energised and heating or pressurising before the first good part. It should almost never count as production.
- Setup and changeover. Energised, often cycling, producing nothing sellable. Decide whether this is its own state or a flavour of downtime.
- Variable load. A machine whose draw legitimately swings during normal production. List these before testing, because they will produce the most confusing results.
Get sign-off from an operator, a supervisor and whoever owns the OEE number. If those three disagree, no sensor will resolve it, and the test will simply relocate the argument.
With the states agreed, the next question is which of them the signal can actually establish on its own.
Know what the signal can and cannot establish
A validation test can only validate what the signal is capable of carrying. Deciding that boundary first is what stops a team from testing the wrong thing and then calling the result a failure.
Current-based monitoring reads what a machine draws and infers activity from the electrical signature. That inference is strong for availability and machine state: running, idle, down, and the timing of the transitions between them. It is not a source for quality judgements, reject counts, or job and order context. Those come from an operator, a quality system, or an integration, and they should be excluded from the accuracy test rather than counted against it.
Controller-based monitoring inverts the trade. It can read part counts, alarms and process values directly, and it is limited instead by which assets have a controller worth integrating. The comparison of PLC monitoring and current sensors sets out where each is stronger, and the no-PLC monitoring buyer's guide covers what a current signature can establish on assets that were never integrated.
Before the test starts, split your field list in two: fields the signal should establish on its own (machine state, availability, timing of transitions) and fields that need a person or a system (reject counts, job context, quality judgements). Validate the first list. Specify the second. A test that holds a current sensor responsible for reject counts will fail for reasons that have nothing to do with accuracy — and that distinction also tells you exactly what your ground truth has to cover.
So before the test starts, split your field list in two: fields the signal should establish on its own, and fields that need a person or a system. Validate the first list. Specify the second. A test that holds a current sensor responsible for reject counts will fail for reasons that have nothing to do with accuracy.
That split also tells you what your ground truth has to cover.
Build ground truth from known events
Ground truth means events you independently know happened, at a time you independently know. Comparing one inferred signal against another inferred signal proves nothing except that two systems agree.
Three sources work, and two used together beat one.
Direct observation. Someone stands at the machine with a timestamped log for two to four hours and records every state change. It is unglamorous, and it is the strongest evidence you will get. Run one window per machine, on the shift where behaviour is most variable.
Deliberate induced events. With the line's agreement, stop the machine at recorded times. Run a changeover. Leave it idling. Run the light product. Because you chose the timestamps, they are not in dispute, which makes this the fastest way to measure detection lag.
Existing records with known timestamps. Production records, job start and finish times, maintenance logs, any existing counter. Weaker than the first two, because those records carry their own error, but useful for volume.
Log everything to one sheet: machine, timestamp, event, observer, product. Keep it after the test. It becomes the reference for every later argument about whether the system has got worse, and without it those arguments have no resolution.
With definitions agreed and ground truth in hand, the test itself is mechanical.
Run the test
- Pick four representative machines, not four easy ones. One high-volume workhorse. One old asset with no controller. One machine with wide variable load. One where operators already dispute the numbers.
- Baseline what the plant believes today. Record current OEE, downtime hours and top reasons, and how each is produced. You are testing a signal, and you are also about to learn how wrong the existing picture was.
- Collect two weeks of continuous data. One week rarely covers product mix, a changeover cycle and the shift patterns where behaviour differs. Include a weekend if the plant runs one.
- Overlay ground truth on the recorded states. For each observation window, line the logged events up against the system's timeline. Count agreements minute by minute rather than eyeballing a chart.
- Classify every disagreement. False running, missed stop, late detection, misclassified warm-up, missed cycle, double count. The pattern matters more than the total. Twelve errors of one type is a threshold to adjust. Twelve errors of nine types is a deeper problem.
- Tune thresholds, then re-test on held-back data. Adjust the running threshold, the minimum stop duration, the idle boundary. Then validate against a window you did not use for tuning, otherwise you have fitted the test rather than fixed the signal.
- Test the product range explicitly. Run the lightest product the machine makes. Separation between running and idle is most likely to narrow there.
- Score against the criteria you wrote in advance, per state, not as one blended number.
Sample across shifts as deliberately as you sample across machines. Guidewheel's 2026 Factory Uptime Report, built on 75 million machine-minutes captured by clip-on sensors, found that the same crew running the same changeover shows a 57% shift-to-shift spread. A validation window that samples one shift will either mistake that spread for sensor error or miss it entirely.
Step eight only works if the criteria genuinely predate the data.
Set pass and fail criteria before you see the data
Write these before the test starts. A threshold agreed afterwards is not a threshold, it is a negotiation.
A workable default set, which you should adjust to the plant:
- State agreement: at least 95% of sampled minutes match ground truth on running versus not running.
- Stop detection: 100% of stops longer than your defined minimum, inside an agreed detection lag.
- Micro-stop floor: state the shortest stop you require to be captured. If the plant needs 30-second stops and the configuration resolves two minutes, that is a scope decision to take now, not an argument to have later.
- Cycle count: within an agreed percentage of a manually verified count over a full shift.
- Reason completeness: the share of downtime minutes carrying an operator reason. This measures adoption rather than the sensor, and it usually decides whether the deployment is useful at all.
- No silent degradation: accuracy on the light product within an agreed distance of accuracy on the heavy product.
Set the micro-stop floor against what your losses actually look like. Guidewheel's downtime analysis puts mechanical breakdowns at roughly 72 minutes per event, staffing issues at roughly 197 minutes, and material or supply delays at roughly 119 minutes. Events of that length are easy for any configuration to catch. The threshold decision is really about the short stops underneath them, which is where the disputed minutes live.
Also write the fail action. If state agreement lands at 88%, does the rollout stop, proceed on a subset of machines, or proceed with a recorded caveat on the affected asset class? Deciding this in advance is what stops a middling result from becoming a stalled project.
What good and bad results look like
A good result is rarely 100%. Expect a small number of genuine edge cases, concentrated on a specific asset or a specific product, with an explanation you can articulate. That is a working system with known limits.
A bad result has no pattern. Errors scattered across states, machines and shifts with no common cause usually mean the state definitions are wrong rather than the signal. Go back to the definitions page.
A specific, useful failure looks like this: the 1998 press reads idle as running whenever the hydraulic pump stays energised between parts. One asset class, one clear cause, one threshold to change. Finding it on four machines costs a fortnight. Finding it on a hundred costs the credibility of the whole programme.
A result that looks too good deserves a second look. If the system agrees with ground truth on every minute of every window, check that the observation windows were not all taken during steady production. Accuracy during a clean run is not the claim being tested.
Once accuracy passes, the remaining risk is organisational rather than technical. The plan for piloting beside existing PLC and SCADA controls covers the ownership, security and scale-or-stop decisions that follow, and automated downtime reason tracking covers the operator workflow that turns a validated signal into a reason you can act on.
To design a validation window around your own assets and product mix, talk to the Guidewheel team.
Frequently asked questions
How long should a signal accuracy test run?
Two weeks of continuous data on four representative machines, with at least one direct observation window per machine. One week rarely covers the full product mix, a changeover cycle, and the shift patterns where behaviour differs. Those are where the informative errors live.
Why does a current sensor sometimes show a machine as running when it is idle?
Because the machine is still drawing power. Hydraulics, heaters, pumps, drives and control cabinets stay energised between parts on many assets, so the electrical signature of idle can resemble light production. This is a threshold and definition problem rather than a sensor fault, and it is the most common finding in a validation window.
What accuracy should we accept before rolling out?
Set it before you test. A reasonable default is at least 95% agreement on running versus not running, full detection of stops above a defined minimum duration, and cycle counts within an agreed tolerance of a manual count. Agreeing the number in advance matters more than the number itself.
Can we validate against our existing OEE numbers?
Use them as context, not as ground truth. Existing numbers usually come from manual logs with their own error, and a common outcome of validation is finding the old picture was further off than the new one. Compare against observed events and induced stops instead.
What if accuracy is good on new machines but poor on old ones?
That is a normal and useful result. Treat the affected asset class separately: adjust its thresholds, decide whether its accuracy is sufficient for the decisions it feeds, and record the limit openly. Scaling a known limitation with eyes open is fine. Discovering it after a hundred installs is not.