Short answer: A tolerance is the stated amount a part may differ from its target size and still work. Testing is the proof that it stayed inside that limit. Precision engineering pairs the two. Write the allowance on the drawing, measure against a traceable standard, then publish what the measurement showed.
No workshop has ever made two identical parts. A balance-staff pivot turned to twelve hundredths of a millimetre is never exactly that, and the next one differs again. Precision engineering starts by accepting this fact. Then it does the one useful thing left.
It decides how much difference is allowed, writes the number down, and checks. The allowance is called a tolerance. The check is called inspection, or testing. Between them they hold up jet engines, pacemakers and the watch on your wrist.
Watchmaking is a good place to see the method work, because its allowances are unusually small and its tests are unusually public. The logic belongs to every trade that makes things, including the ones that now make software.
What a tolerance actually says
An engineering drawing never asks for a part to be about a size. It states a nominal dimension, meaning the target, and a tolerance beside it. A shaft might be drawn as 10 millimetres with an allowance of plus or minus one hundredth of a millimetre. Anything inside that window passes.
The window is a design decision in its own right. It must be tight enough that the part works with every neighbour it meets, and loose enough to be made repeatably at a sane price. A tolerance is not permission to be sloppy. It is a precise statement of how imperfect a part may be.
A tolerance is not permission to be sloppy. It is a precise statement of how imperfect a part may be.
Limits, fits and clearance
Parts work in pairs, so tolerances are chosen in pairs. The gap between a shaft and the hole it turns in is the clearance, and the pairing is called a fit. A bearing made too tight will seize. Made too loose, it rattles and wears fast.
The international system of limits and fits, published as ISO 286, gives these windows standard grades and codes. Two workshops on different continents can then read one drawing the same way, which is what makes parts interchangeable. In a watch movement the clearances are extreme, since balance pivots run in jewelled bearings at a few thousandths of a millimetre. Our walkthrough of how a mechanical watch movement works follows those parts in motion.
Why tighter is not automatically better
Cost does not rise in a straight line as a window narrows. It curves upward and then climbs steeply. Halving an allowance can mean finer tooling, slower cutting, more inspection and a higher scrap rate.
So precision is spent where it buys function and saved everywhere else. A case back can be loose by a tenth of a millimetre without consequence. An escape-wheel tooth cannot. That trade-off is followed further in our essay on the philosophy of precision.
The instruments that decide
A limit means nothing without something to measure against it. Workshops keep a ladder of instruments, each one finer and slower than the last.
- Vernier calliper: reads external and internal sizes to about two hundredths of a millimetre, and is the fast first check.
- Micrometer: measures to one thousandth of a millimetre, which is one micrometre, using a calibrated screw.
- Gauge blocks: hardened steel or ceramic blocks of certified length, stacked to build a known reference on the bench.
- Measuring microscope: reads the tiny profiles of watch parts optically, without touching and deforming them.
- Coordinate measuring machine: probes a complex shape point by point and compares the cloud of points against the drawing.
Each instrument has a resolution, meaning the smallest step it can report. Resolution is not the same as accuracy. An instrument can print four decimal places and still be wrong in the second one.
Go and no-go gauges turn measuring into a decision
Once limits exist, an inspector rarely needs the exact size. They need to know which side of two lines the part falls on. That insight produced the go/no-go gauge, a pair of fixed references used in sequence.
The go gauge must slide on. The no-go gauge must not. A part that passes both tests is inside its window, and the check takes a second rather than a minute. The gauge does not say how big the part is. It says the only thing the assembly line needs to know.
Every measurement needs a chain behind it
Measurement has its own science, called metrology. Its central idea is traceability. The gauge on a bench is calibrated against a laboratory standard, which is calibrated against a national standard, which realises the international definition of the unit.
Those definitions belong to the SI system, maintained under the International Bureau of Weights and Measures. An unbroken chain of comparisons links the workshop to the definition. That is why a millimetre in Geneva and a millimetre in Seoul are the same millimetre.
Time shows the chain most clearly. The SI second is defined by the caesium-133 atom, as 9,192,631,770 cycles of a particular radiation. Atomic clocks realise that definition at institutes such as the Physical Measurement Laboratory at NIST. Every timing machine in every watch workshop is judged, at the end of that chain, against atomic time. The route to that definition is told in our short history of measuring time.
Temperature, the variable nobody sees
Metal moves with heat. Steel grows by roughly eleven micrometres per metre for every degree Celsius, which is enough to swallow a fine tolerance. Dimensional measurement therefore has a reference temperature of 20 degrees Celsius, set by ISO 1.
A dimension quoted without its temperature is only half a fact. The same rule applies to any number produced under conditions, which is the point of every test regime that follows.
How a finished watch is tested
Parts pass their own inspections, then the assembled watch faces a second layer of tests. Each one has a published standard behind it.
| What is tested | Standard | Conditions | What passing proves |
|---|---|---|---|
| Rate, for chronometers | ISO 3159, applied by COSC in Switzerland | 15 days, 5 positions, 3 temperatures | Mean daily rate between minus four and plus six seconds |
| Water resistance | ISO 22810, with ISO 6425 for divers' watches | Defined pressures, plus condensation and immersion checks | The case seals at the depth printed on it |
| Shock resistance | ISO 1413 | Two impacts that simulate a fall onto a hard floor | The watch keeps running within a stated rate change |
| Magnetic resistance | ISO 764 | Exposure to a defined magnetic field | The movement still runs and holds its rate afterwards |
Chronometer testing, day after day
The chronometer regime is the strictest of the four and the easiest to check. Every submitted movement is measured individually, not sampled from a batch. It runs for fifteen days while the laboratory changes its position and its temperature.
Position matters because gravity pulls on the balance differently as the watch turns. Temperature matters because heat changes the elasticity of the balance spring. Fifteen days matter because one good day proves nothing. Who runs the test matters too, a point covered in what Swiss Made teaches us about quality standards.
Pressure, shock and magnetism
Water resistance is tested with pressure rather than depth, and dive watches face stricter individual checks than everyday models. The ratings are widely misread, which is why we unpack them in watch water resistance ratings explained.
None of these results is permanent. Gaskets harden, lubricants migrate and shocks accumulate over years. A quality claim is a measurement with a date on it, so periodic pressure checks belong in any serious watch care routine. The honest answer to an old measurement is a new one.
Testing a million parts without measuring them all
Testing every unit works for chronometers. It is impossible for screws stamped out by the million. Mass production therefore measures the process instead of the product.
Two kinds of variation
Walter Shewhart worked out the method at Bell Laboratories in the 1920s and gave it a tool, the control chart. His insight was that variation comes in two kinds, and that confusing them makes quality worse.
Common-cause variation is the ordinary noise of a stable process, and reacting to it only adds disturbance. Special-cause variation is a signal that something has changed, such as a tool wearing or a material batch drifting. A control chart plots samples over time so the two can be told apart early.
Capability and sampling
The companion idea is process capability. It compares the natural spread of a process against the width of the tolerance it must hit. A process that barely fits its window will leak defects at the smallest disturbance. A capable one fits with room to spare, so ordinary variation never threatens the specification.
Sampling plans such as ISO 2859-1 then define how many pieces to inspect from a batch and how many defects end it. Note what such a plan promises. It gives a stated confidence about a batch, not a guarantee about every piece. Honest quality work says which of the two it is offering.
Tip: Inspecting harder at the end of a line is the most expensive way to buy quality. Capable processes and early checks are cheaper, because the cheapest defect is the one that was never made.
Tolerances for behaviour, not dimensions
Software has no dimensions to mismeasure. It does have behaviour, and behaviour takes tolerances too. A response-time budget, a maximum failure rate and an accuracy floor are all windows with limits written down in advance.
Seen that way, the engineer's kit maps across almost part for part. A unit test is a go/no-go gauge for one property. A regression suite is re-inspection after a change to the process. Continuous integration is the inspection gate, moved to the point where a defect is cheapest to catch.
The varied conditions of chronometer testing have equivalents as well. Five positions and three temperatures become many devices, load levels and hostile inputs. Property-based testing and fuzzing exist to explore the conditions a designer never thought to specify.
Published thresholds crossed over intact. Speech recognition is judged by word error rate, a defined metric explained in what word error rate really means, and service reliability runs on explicit error budgets. Our AI accuracy topic page collects the rest of those measures.
Setting a tolerance you can defend
Whether the artefact is a pivot, a gasket or a pipeline, the same five questions size an allowance honestly.
- What does the part have to do? Function sets the window. Habit and copied drawings do not.
- Who is harmed by the error? A dive watch and a desk clock deserve different budgets, because the consequences differ.
- Does the error stack up? Tolerances add along an assembly, so a chain of loose parts can fail while every piece passes.
- Can we actually measure it? An allowance finer than your instrument's uncertainty is a wish, not a specification.
- What does the next step cost? If halving the window doubles the price, say so and stop.
Notice that the third question is the one most often skipped. Stacked tolerances explain many assemblies that fail inspection with no single guilty part. Careful designers therefore run a stack-up calculation before the drawing is released. The watch movements topic page shows how many stacked parts one small case has to hold.
Key takeaways
- A tolerance is a written allowance: the amount a part may differ from its target and still work.
- Gauges beat opinions: go/no-go gauges turn a delicate measurement into a fast, unambiguous decision.
- Traceability makes numbers real: every bench gauge is calibrated back to the SI definitions through an unbroken chain.
- Conditions belong with results: ISO 3159 runs 15 days, in 5 positions, at 3 temperatures, for exactly that reason.
- Behaviour has tolerances too: tests are gauges, continuous integration is the inspection gate, and error budgets are the window.
Tolerance and testing are often described as limits on engineering. They are better read as the conditions that make it possible. Both are admissions, and writing them down is what turns an admission into a promise.
An industry that states its imperfections precisely can state its performance credibly. That stance travels well beyond the workshop, whether the object is a balance staff, a sealed case or a deployment pipeline. Decide how wrong is acceptable, measure honestly, then let the numbers argue. The rest of the argument waits in the Craft & Precision hub.