Abstract
Geoscientific datasets can contain individual data for more than 50 different chemical elements. The association between these variables is as important as their individual values. However, it is commonly overlooked that the observed covariance may be overestimated due to correlated errors. Dependent errors arise from many sources, such as the segregation process of minerals associated with these variables during delimitation, extraction, and preparation steps. This study extends a classical model composed of grade-independent (additive) and grade-proportional (multiplicative) errors to a generalised multivariate model that can estimate the real variance, covariance, and correlation from observations affected by shared errors. The use of estimates of the real covariance is recommended when the study objective is to evaluate or estimate the association between processes instead of the association between observations. A numerical example illustrates the bias in statistics and discusses the relevance of considering shared errors in linear regression and kriging.
Keywords
Introduction
Geochemical, environmental, or mining databases commonly contain thousands of samples, each of which may have individual values for more than 50 different chemical elements or species. This abundance of data provides an opportunity to discover a wide range of geological processes within a surveyed area. Multivariate statistics makes it possible to interpret the significance and relationships of tens of variables. Inevitably, all sample measurements are affected by the intrinsic presence of sampling error (Gy 1982). Two or more variables may be affected by the same source of error. The error associated with observations of different variables may be highly correlated. The measured correlation or covariance from observations affected by this type of error can be greater than that of the underlying true process, overestimating the real association between geological processes based on data available.
Covariance and correlation are parameters required by almost all multivariate statistical methods. These statistics are directly affected by the proportion of shared errors in the total error. Ignoring shared errors and assuming the observed statistics are representative of the underlying true process may lead to wrong decisions. There are examples of methods directly impacted by this wrong assumption such as cluster analysis based on the structure of covariance within observations rather than on their means (Ieva et al. 2016), principal component analysis (Pearson 1901), or multivariate geostatistics (Wackernagel 2003).
The total error is divided into four components, shared multiplicative, shared additive, non-shared multiplicative, and non-shared additive errors. Shared errors arise from correlated characteristics of minerals that lead to similar sampling errors and biases on the observations of two or more variables. To our knowledge, no discussion exists in the geosciences related to how to model and infer the components of error or how to estimate the underlying true correlation and covariance from observations affected by shared errors. The developed model of errors in the present study employs the same assumption as the widely known Thompson–Howarth model (Thompson and Howarth 1973, 1976, 1978). The term ‘sample’ has been a source of confusion between the statistical and the geoscience literature. The first usually refers to samples as a set of observations from a population. However, in geoscience literature, all specimens of soil, rocks, stream sediments, etc. are often called ‘samples’. Within this paper, collected specimens are referred to as samples, while each measurement of a given variable of interest at this sample is referred to as an ‘observation’.
The aim of this study is to understand the presence of shared errors and propose an approach to estimate statistics between underlying true processes from observations affected by this type of error. This article is organised as follows: (1) it starts with a section where we study the presence of shared errors and discuss very common situations in which they may occur; (2) next, the classic model of errors is presented and extended to a novel multivariate model to deal with shared errors; (3) the issue of estimating the variance, covariance, and correlation of the underlying true processes from observations is addressed; (4) a solution using pairs of duplicates to estimate the four different types of errors is proposed; (5) a numerical experiment is used to illustrate how bivariate statistics measured from observations with shared errors may diverge from the real statistics; and (6) the paper concludes with examples of applications to which ignoring the presence of shared errors may lead to wrong decisions.
The occurrence of shared errors
Shared errors may arise in a variety of ways in very common geological situations. Suppose there are N variables of interest, k = 1, … , N, and the same sample is assayed to obtain their readings. Some sources of error may be shared within the analysed variables. Figure 1 illustrates four different levels of association between two variables. In Figure 1(a), both variables are associated with the same mineral assemblage structure, so their sampling errors are shared at any scale and granulometry. Figure 1(b–d) illustrates other levels of physical association.
Illustration of four different levels of association between the two variables (red and green). (a) Elements of both variables share the same mineral assemblage structure, so their sampling errors are shared at any scale and granulometry. (b, c) illustrates intermediary associations. (d) Complete independence between minerals, in which case the shared error can be caused only by similar properties that lead to similar segregation.
The heterogeneity of a lot may be greatly affected by the relationship between physical properties of its minerals, leading to a state in which segregation or distribution heterogeneity is the rule, rather than the exception. Therefore, the only method to reduce the segregation factor is to homogenise the material before sampling (Pitard 1993). Figure 2 shows how an incorrectly taken sample may show an association that does not exist in the true geological process. The figure shows the initial homogeneity state (Figure 2(a)) of two different flake minerals (red and green) and the same minerals in a highly segregated state (Figure 2(b)). Three samples were taken from the non-homogenised lot (dashed-line areas in Figure 2(b)). The segregation process leads to a condition in which red and green minerals are simultaneously above or below the average proportion of the lot, creating a false correlation.
A lot composed of gangue minerals (grey) and two types of flake minerals (original). In (a), it is in in situ states of homogeneity; (b) the prepared lot is highly segregated and with three samples delimited (dashed squares). (c) The regression between the proportion of the two minerals in the three samples.
There are many other situations where shared errors occur. These are mainly related to variables associated with minerals that are affected in similar ways in a given segregation process during delimitation, extraction, and preparation steps. Examples of mineral differences that can simultaneously segregate two or more minerals from gangue minerals are (1) different shape and density; (2) magnetic properties; (3) electrostatic properties of liberated constituents occurring as tiny flakes, such as biotite and scheelite; (4) moisture content and/or adsorbing water capability; and (5) minerals adhering differently to the walls of a storage bin or having different angles of repose, etc. (Pitard 1993).
The true association between variables with shared errors may only be correctly estimated if the error model used considers this type of error. The next section presents the classic model of uncorrelated errors to be extended to the multivariate case.
The classic observation model
Observations zi,obs are indirect measurements of the underlying true process of Ztrue. Each individual observation i, i = 1, … .,I, comes from zi,obs = ztrue + εi, where εi in the simplest case is unbiased E{εi} = 0, uncorrelated with the random variable Cov{Zi,true, εi} = 0, and with a known variance Var{εi} = σ² i ,error that may be different for each sample.
When there is a wide range interval in the observations, the sampling error can vary significantly over the range of grades, so the standard deviation
alone cannot properly describe the error behaviour and its precision. In this case, the well-established fact that the average error increases as the grade increase needs to be included (Howarth and Thompson 1976; Thompson and Howarth 1973, 1976, 1978). Thompson and Howarth assumed unbiased data affected by normally distributed errors. The error is part of a function of the grade and part independent. In the resulting model, the unknown underlying true value is affected by additive and multiplicative errors:
: Sampling error variance associated with the observation zi,obs; Xm: Standard normal Gaussian values N(0, 1); Ak: magnitude of additive error component; Ck: magnitude of multiplicative error component.
The scatter diagram between matched pairs of duplicate samples should help to decide which model is more adequate based on the shape of the cloud of points. Figure 3 (green dots) shows the effect of additive errors. When Ck = 0 [Equation (1)], the scatter plot between duplicates shows a constant distance around the x = y line. In the case of multiplicative error (yellow dots), Ak = 0 (Equation (1)), the closer the grade is to zero, the smaller the error. Real datasets are commonly affected by both types of errors.
Behaviour of observations affected by independent (green dots) and grade-proportional errors (yellow dots). The dashed black line indicates the x = y line.
Equation (1) deals with situations where the errors of each variable are independent of the error attached to observations of other variables. We now develop a generalised model to deal with additive or multiplicative errors shared by two or more variables.
The model of observations in the presence of shared errors
The development of the generalised equation is based on adding two new components to Equation (1). (1) The parameter B is the magnitude of additive errors, and (2) D is the magnitude of multiplicative errors. These new components are shared by two or more variables k, affecting them with the same magnitude and proportional deviation between the observed and real value. The resulting model, where the total grade-proportional error is split into shared (D) and non-shared (Ck) and independent errors are divided into shared (B) and non-shared (Ak), is:
Inferring the geological structure of a correlation
The variance of Equation (2) combines the sampling and analytical errors plus the variance of the underlying true process:
Scatter plot between two error-free variables with a correlation r² of 0.63. The variables were perturbed by an error Ck + D = 0.1 with D of 10, 50, and 90%. The measured correlation after adding the error is 0.23, 0.51, and 0.77, respectively.
can be estimated from observations in a given domain based on bilinearity properties. The non-shared additive Ak and multiplicative terms Ck are cancelled and the measured covariance is only affected by shared errors:
can be estimated by the observed variance minus the sampling error variance
. Figure 4(a) shows observations between two error-free variables with a correlation r² = 0.63. The variables were perturbed by a total error Ck + D = 0.1 with D = 0.01, 0.05, and 0.09. The measured r² after adding the error is, respectively, 0.23, 0.51, and 0.77. It is interesting to highlight that the measured correlation between variables with shared errors of 90% is larger than the error-free correlation.

We present how to use pairs of duplicates to estimate the four different types of errors in the next step.
Estimating the error components
An exhaustive discussion about approaches to estimate the presence of shared errors is beyond the scope of this paper. However, some insights about how to infer Ak, B, Ck, and D using duplicate samples are discussed. Duplicates are used to control data quality and process repeatability. The duplicate must be collected at a particular stage of the sampling protocol, such as comminution and sub-sampling, then prepared and assayed following the same rules used for collecting the original sample. Matched pairs of chip samples or pairs of parts of a drill core are the best data to evaluate the combined contribution of all sources of errors, allowing monitoring the variance added by each subsequent stage.
Additive components Ak and B
We can find the value of
by computing the variance of a set of measurements of samples with a null grade, which is possible because grade-proportional components are null, and Ak and B control all the variance.
Multiplicative components Ck and D
The linear regression slope is not affected by additive errors, so it can be used to infer the multiplicative components (Ck and D). Three different regressions allow us to estimate the values of
The slope of the regression line between the variance and the average value of each matched pair is
The regression between the difference Var{zi,1, obs} – Var{zi,2, obs} (computed with duplicates of each variable), and the difference of grades from the original observations has a slope of
The regression between the covariance of measurements and the product of the grade of the original (or duplicate) samples of two observations i. The regression slope is C1 C2 (Figure 5(c)). Regression line and slope equation between variance and data are presented in three different configurations. (a) Y-axis is the variance of each matched pair (duplicate and original) and X-axis their average grade; (b) Y-axis is the difference Var{z1, obs} − Var{z2, obs} measured by the matched pairs of each variable and X-axis the difference of the grade of original observations; (c) Y-axis is the covariance between observations of two variables and X-axis is the product of the grade of their original observations.
,
,
and C1 C2:
(Figure 5(a));
(Figure 5(b));
The three different results from the regression slope are used to define the individual values of each component by solving a linear system of these equations.
Simulation studies
(Tables 1 and 2).
Covariance and correlation between 50, 000 values drawn from N(100, 10) and perturbed individually by an error N(0, 5), resulting in ρ(Xi,1, Xi,2) = 0.62. Next, these values were perturbed by additive errors Ak + B = 10, with different proportions of shared (B) and non-shared (Ak) errors.
Covariance and correlation between 50, 000 values drawn from Ni(100, 10) and perturbed individually by an error N(0, 5), resulting in ρ(Xi,1, Xi,2) = 0.62. Next, these values are perturbed by multiplicative errors Ck + D = 0.1, with different proportions of shared (D) and non-shared (Ck) errors.
Next, the initial values were used to simulate observations Xi,1 and Xi,2, perturbed by additive or multiplicative errors. The total added error is constant to both cases, being Ck + D = 0.1 or Ak + B = 10. We varied the ratios between non-shared and shared components, Ck/D (Figure 4) and Ak/B, using 0, 10, 30, 50, 70, 90, and 100%. Tables 1 and 2 give the results of these simulations in the presence of additive or multiplicative error, respectively.
The non-shared components Ak and Ck do not change the measured covariance, and the higher the error, the lower the measured correlation between variables (Tables 1 and 2). Shared errors have an inverse impact; the bigger the error variance, the closer ρ(Xi,1, Xi,2) is to one since the measured correlation combines the underlying error-free correlation plus the shared-error correlation. In both cases, overlooking the type of error attached to observations can lead to statistics being non-representative of the geological processes of interest.
Applications
The true statistics cannot be directly measured from observations in real-world problems because available data are always affected by errors. Bivariate statistics naively inferred from these observations measure the association between observations but not the variance, covariance, and correlation between the underlying true processes of interest. Two situations in geosciences in which the covariance plays a relevant role in parameter estimation were used to show how the true statistics may be estimated using the developed equations and what needs to be done after we measure shared and non-shared errors and estimate the correct associations between variables Extending the proposed solution to other statistical applications is straightforward.
Linear regression in the presence of shared errors: In a linear regression y = α + Xβ, the ordinary least-squares (OLS) estimates of the intercept α and slope β are obtained from observations X. When observations are error-free, the OLS estimate of the regression slope,
, is:
A simple method of accounting for shared errors between X and Y in linear regression is a direct application of the proposed method, where the equation's slope is ‘debiased’ by replacing the statistics adjusted through available observations with the correct variance (Equation (3)) and covariance (Equation (4)). Next, we use these estimated moments to estimate the regression parameters of Equation (7):
Collocated co-kriging in the presence of shared errors: Collocated co-kriging is commonly used when primary data are present in sparsely distributed points while secondary data are located at all points of the grid being estimated. To avoid matrix instability, collocated co-kriging makes use of the secondary data only at the nodes to be estimated. This solution relies on the Markov-type screening hypothesis, which assumes that these collocated data screen further away data of the same type (Xu et al. 1992). As a practical result of this assumption, we do not need to compute the cross-covariogram between primary and secondary data C12(
In real-world problems, the unknown correlation between secondary data and the underlying true process of interest is assumed to be equal to the correlation between the association ρ12
ρmin is the minimal correlation in which the use of secondary data is worthwhile in estimates. In a case of overestimated correlation, ρ12
Conclusions
Shared errors may arise from the delimitation, preparation, and sub-sampling steps of a sampling protocol, creating a false correlation that is not present in the geological process. This paper proposes a general model that deals with four types of error, shared multiplicative and additive errors and non-shared multiplicative and additive errors. Next, the effect of each type of error in measured bivariate statistics was investigated. A long list of common situations where shared errors can arise during sampling was presented.
There are many applications where the correct covariance and correlation to be considered should measure the real association between the processes, not the association between observations. Constraining models to the correct statistics are important for the greatest prediction accuracy and precision. Our main finding was that the measured correlation and covariance may be substantially under or overestimated by overlooking the impact of the shared errors. When the magnitude of each type of error is known, it is possible to estimate the correct statistics from observations. A method to estimate these magnitudes using quality control data was presented.
Footnotes
Disclosure statement
No potential conflict of interest was reported by the author(s).
