Abstract
Key comparison measurements serve as an ultimate tool of quality assurance of results. Whenever the inter-comparison results indicate inconsistency, the participating laboratory needs to take the corrective actions. Practically, the systematic errors involved in the measuring system confines the achievable accuracy. Therefore, the corrective action involves either empirically determine the influences afresh or intuitively reassigns these error values. Alternatively, an analytical method based on inter laboratory comparison results is proposed. The novelty of the proposal is considering task-specific errors in the model that is used for the analysis of interlaboratory comparison results. Without accounting the uncertainties of task-specific errors, the analysis grows complicated and even sometimes it is not feasible. To supplement the proposed method, task-specific errors due to the imperfect geometry of ring gauge, practical inability in implementing the measurement, and unattended environmental influences are explored. The proposed method is demonstrated using some internal diameter key comparison data. The systematic errors responsible for the outlier in the measurement comparison are clearly distinguished.
Keywords
Introduction
The mutual recognition arrangement encourages free trade between countries by removing technical barriers. Mutual recognition establishes the basis for a “product tested once and accepted everywhere” principle (CIPM MRA documents:policy @ https://www.bipm.org). The international quality infrastructure aims to assure comparability of measurement results. Consensually, the designated laboratory called national metrology institute (NMI) of different economies participates in the interlaboratory measurement comparisons to demonstrate their equivalence of routine calibration services offered by to clients (CIPM MRA documents: guidance on comparison @ https://www.bipm.org). Participant NMIs measure the same artifact using the same apparatus and methods as routinely applied to client gauges. The measurement values reported by participants are compared with the agreed /assigned reference value (
On the contrary, the analytical model required for the corrective action in response to the key comparison outcome has to be specific to the given measurement (Dobaczewski et al., 2014). Rukhin (2009) expressed that accounting the exact nature of all the participant laboratory measurement techniques needed in the formulation of a finite population sampling is difficult. Conceptually, the corrective action belongs to the jurisdiction of the participant. The work carried out (Chunovkina, 2008) for improving the estimation of bias of intercomparison assume that the probability density function (PDF) that encodes the size of its bias is available and results are consistent. Otherwise, the uncertainty about their biases cannot be reduced. Moreover, the consistency of data relies on the results provided by the participating laboratories including their quoted uncertainties. Often, we do not have the knowledge of PDF of each participant. Therefore, the benefit of intercomparison is not optimally extended for the corrective action using only statistical methods.
In this work, a model is devised to bridge the measurement model and the interlaboratory comparison analytical model to ease the above difficulties. Grabe (1987) also suggests the metrological statistics that account for the systematic bias in the model used for the analysis of intercomparison. His result advocates that the error budget can easily be established and conserved. Then, one can assess the reasons for a given uncertainty. It was also pointed out that these simplifications are intimately connected with a modest increase in measurement uncertainties. To fill these gaps, the concept of the task-specific error is applied here. The authors explored a common minimum task-specific uncertainty due to misalignment, form deviations of artifacts and constant environmental influence involved in the ring gauge measurement. After the evaluation of the degree of equivalence, a common minimum uncertainty is appended to all the results reported by participants and their respective systematic uncertainty components are adjusted accordingly. The authors also investigated deformation of a steel ring gauge of size 4 mm due to atmospheric pressure and gravity using finite element analysis. The proposed method is described using some artificial interlaboratory comparison data. Finally, the revised En values are calculated after readjusting the systematic errors. The revised En values highlighted nearly outliers also. During the redistribution of bias of participant laboratories, the misinterpreted errors and underestimated uncertainty components are revealed.
Proposed approach
The proposed model represents the measurement of internal diameter. The parameter (diameter of ring gauge) is expressed in terms of form imperfection (
The deviation
The probing direction
During inter laboratory measurements, the estimates of the task-specific error
To reduce the offset
Chunovkina (2008) has discussed the cases that reduce the uncertainty
According to Kacker et al. (2007), the revised error is parsed into lab specific bias and task-specific bias
The variation
Obviously, the uncertainty
Since the proposed analysis should be specific to the individual laboratory; the uncertainty of offset
Whenever a participant laboratory has mistaken the systematic errors and/or their associated uncertainty components, it would not be possible to redistribute systematic errors as explained above. The algorithm of the proposed method is summarized here:
Assign systematic errors and their associated uncertainty according to equation (5).
Investigate the given measurement to determine the task-specific error components.
Evaluate the common minimum uncertainty due to task-specific errors.
Reduce the reassigned errors from the bias
Evaluate the uncertainty of refined offset
Evaluate normalized error after readjustment of biases using equation (10).
Analysis of corrective action
Different NMIs have different machines for internal diameter measurement. Even then, the basic principle of internal diameter measurement remains the same for all. Majority of NMIs use single axis length measuring machine (LMM) as follows. The LMM holds the ring gauge on its tip-tilt table. A carriage equipped with a single (two) spherical anvil(s) moves on the guide-ways (Lee et al., 2013). The detector on the carrier reads a reference linear scale to announce displacement of anvils inside the cylinder. Initially, the internal diameter of a master ring gauge is probed. A probing correction is calculated by subtracting the measured displacement inside the hole from the diameter value provided in the calibration certificate of the master ring gauge. The probing correction accounts the probe tip diameter, indentation of a probe and other biases due to imperfections in the probing system. Then, the probe is displaced inside the ring gauge under test. The measured internal diameter will be the sum of displacement of the probe inside the ring gauge and the probing correction. Hence, the uncertainty component due to probing correction (Pahk and Kim, 1995) hikes the overall uncertainty of measurement.
NMIs publish their CMCs on the web portal of the International Bureau of Weights and Measures (BIPM). Ad infinitum, metrologist works to improve CMC by reducing the respective systematic errors (Neugebauer et al., 1997). Experimental quantification of the error components is tedious and impracticable. Neugebauer (1998) comprehensively narrated the various uncertainty components involved in internal diameter measurement. Such knowledge of each influence is required to estimate the undisputable uncertainty of measurement.
Figure 1 shows the strategy to achieve desirable uncertainty of measurement. Scientists can validate error component values in their measuring machine based on their performance in the key comparison. Further, designers can refine the measuring system to reduce the systematic errors that are conceived in the key comparison. The flowchart describes the process of improving the measurement system. An intercomparison data of internal diameter inspired from some key comparison are taken to discuss the proposed method [CCL-K4 Key Comparison, 2000]. The results reported by participants are analyzed by the weighted mean method. The uncertainty components reported by each participant are also given.

Scheme of improving measurement system using intercomparison results.
The key comparison data is analyzed to determine the weighted mean method (
Data take from some key comparison (All numerical are in nm).
Similarly, all the error components and their associated uncertainties can be split for all the laboratories. We (as a third party) cannot have a precise value of systematic error component of each participant. Therefore, we briefly estimated the magnitude of errors. The statistics of systematic errors reported by different NMIs during key comparisons (Banreti, 2010; CCL-K4 Key Comparison, 2000; Chin et al., 2014, Picotto, 2010) are compiled in Table 2. Usually, these perturbations giving rise to unknown offset from the true values have been formally randomized, that is, they are treated as if they were of random origin. Biased estimators thus again become fictitiously unbiased.
Typical systematic errors (nm) reported in some intercomparisons.
The individual laboratory can adjust the systematic error components following some thumb rules as follows. Theoretically, the form deviations of artifacts cause uncertainty equally to all the laboratories participants of key comparison. One can argue that the sign of form deviations will be same for all the labs. But, the individual labs might have assumed different values. The anticipated error components of the chosen inter-comparison results are listed in Table 3.
Adjusted systematic error components (nm).
The associated uncertainty with each error components is taken apart from the respective reported uncertainty components. It should be noted that for each lab, the square root of the sum squared uncertainty component will be less than the reported uncertainty as given in Table 4.
Adjusted uncertainties of systematic error components (nm).
Task-specific systematic errors
The measuring machines should have desirable design characteristics for specific motion, geometry that, if not perfect, will lead to systematic errors in the measurement (Mian et al., 2014). Probing as a means to detect the boundary of an object imposes practical limits on accuracy attainable in the internal diameter measurements. Prominently, the reproducibility of the instrument fall into broad categories viz. reference surface geometry, alignment, and undesirable motion. Reference surface geometry includes the flatness and parallelism of the anvils, the roundness of artifact and the sphericity of the probe balls (Kim et al., 2010).
Misalignment in a vertical plane
According to analytical geometry, the intersection of a plane with a sphere will be a circle. However, the intersection of a plane and cylinder will be a circle provided the plane is normal to the extracted axis of the cylinder. Therefore, the trajectory of the probe should confine to the given measuring plane. Moreover, the extracted datum of the bore of ring gauge must be normal to this plane.
Equation (12) gives the bias
Misalignment in a horizontal plane
Often, two lines are marked on the ring gauge to indicate the direction of measurement. Reproducible results can be ensured by specifying the direction and location of measurement on the face of ring gauge. The translation of probes in this direction should be maintained during the measurement of ring gauge. A microscope can be used to locate these marks when aligning the ring parallel to the translation of the probe. Precisely, the lines do not always pass through the centre of the ring, as shown in Figure 2.

Ring gauge with marks indicating the direction of diameter measurement.
The lateral offset (
Effect of force fields
The force distribution due to atmospheric air pressure (ATP), gravitational pull deforms the ring under test. The location-dependent deviation of gravity and ATP are studied by simulation. A ring gauge of size 4mm is chosen for finite element analysis. The fixed support is set as a reactive force at the base of ring gauge. The ATP is distributed around the ring except for the bottom of the ring gauge. The gravity acts throughout the entire thickness of the material of ring gauge. During finite element analysis, the gravity is acting on each element. The maximum internal diameter variation due to deformation under 10% variation of pressure from reference air pressure (101330 Pa) is about 4 nm. The outer edge of the ring gauge does not change appreciably at the bottom. It was found the increased wall thickness of ring gauge reduces the effect of ATP. Further, the gravity is applied to the entire body of ring gauge keeping the force field due to ATP. The gravity pulls the material layer by layer.
The deformation due to gravity is few tenths of a nanometre for the given sizes of master ring gauges. The combined effect due to gravity, ATP is plotted in Figure 3. The external diameter of the cylinder indenture longitudinally. Similarly, the internal diameter of the cylinder expands by 6 nm.

Simulation of combined deformation due to ATP, gravity.
Results
Often, the results with the lowest uncertainty dominate the key reference value in the weighted mean method. The task-specific errors in the ring gauge measurement are summarized in Table 5. Often, each participant laboratory reports the different magnitude of systematic error ingested with these errors in the key comparison.
Typical value of task-specific systematic errors.
The task-specific uncertainties due to these task-specific errors are given in Table 6. The combined uncertainty is evaluated as 14 nm. Usually, the uncertainty of the weighted mean value will be less than the smallest uncertainty reported among the participant labs. The systematic error value cannot be exclusively known during a given measurement experiment.
Task-specific uncertainty component of internal diameter measurement.
Combined task-specific uncertainty
One needs to perform additional experiments to find the number of systematic errors in each experiment (Stone et al., 2011). The participant laboratories enlarge the uncertainties above their CMC so that the measurement results can be consistent with the KCRV during the draft-A stage of key comparison. Instead, being task specific error are unknown experimentally with respect to magnitude and sign, this component is expected to lie within intervals symmetric to zero. The Eni_adj value of each participant is given in Table 7. It was found that the results of lab 5 are considerably outlier now. It is evident that the lab results are qualified in the inter-comparison due to the large uncertainty claim. Surprisingly, the adverse effect of large uncertainty vanished for lab 13. If we assign positive signed misalignment error value (i.e., 35 nm), lab 13 will qualify with En_adj = 0.96. Theoretically, misalignment results in a lesser value than the actual (nominal) value of the internal diameter of the ring. Therefore, the systematic error component of all labs was set with a negative sign. Individual labs get insight into competitive values of error components. Thus, lab 13 can decide upon the suitable error values in the root-cause analysis. The same is in the case of lab 11; its result will be an outlier with the positively valued temperature compensation.
Results of participant labs.
Conclusion
The proposed approach appears statistically intuitive due to the lack of comprehensive information of each and every participant laboratory. Without the consideration of task-specific uncertainty, the analysis remains intuitive. Since the individual participants have their measurement system at hand; the results of this analysis precisely confirm the compensatable systematic error values and the respectively associated uncertainty. Therefore, the novelty of the analysis involves recognizing inevitable systematic errors. Then, appending the corresponding uncertainty to all participants eliminates the biasing of reference value due to dominant reported uncertainty.
Using this approach, the uncertainty associated with the estimates of the biases is significantly reduced. Thereby, the state of knowledge of the laboratories about their biases is reaffirmed. These benefits of the proposed approach are illustrated by the analysis of internal diameter key comparison data. A common minimum uncertainty is concocted. The deformation due to the variation of atmospheric pressure and gravity are estimated. Task-specific uncertainty component of internal diameter measurement are evaluated. It is also found that the standard uncertainty due to task-specific errors ranges from 11 nm to 25 nm. The results of the proposed analysis discriminated the merely consistent laboratory results. For example, labs 5 and 13 are still outlier despite increasing the uncertainty of their results. The proposed method is easy to practice and fetches the benefits of the intercomparison.
The author(s) declared no potential conflict of interests with respect to the research, authorship and/or publication of this article.
Footnotes
Appendix
Acknowledgements
The authors pay thanks to the Director at National Physical Laboratory, New Delhi for his encouragement.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
