Abstract
Visual Odometry is a key technology for robust and accurate navigation of unmanned aerial vehicles in a number of low altitude applications (<120 m), particularly in environments where access to a global positioning system is possible, but not guaranteed. Navigating via vision alone reduces dependence on a global positioning system and other global navigation satellite systems, enhancing navigation robustness even in the presence of jamming, spoofing or long dropouts. To date, however, most demonstrations of visual odometry are in close proximity to the ground or other structures and are often implemented as a monocular camera combined with inertial sensing, rather than vision alone, to account for scale drift. Stereo visual odometry has received little attention for applications beyond 30 m altitude due to the generally poor performance of stereo rigs for these extremely small baseline-to-depth ratios, otherwise termed long-range stereo. This paper demonstrates stereo visual pose estimation at altitudes of up to 120 m above ground level on a small fixed-wing unmanned aerial vehicle by adapting the traditional stereo visual odometry paradigm to explicitly account for inaccurate triangulation and poorly observed scale from the stereo baseline. In addition, issues related to long-range stereo such as biased sensing are investigated to justify the approach, and a novel bundle adjustment algorithm is presented capable of handling vibration induced structural deformation between the cameras. This is achieved by continually optimizing the stereo transform within a set of inequality bounds. Results are presented demonstrating the algorithm on field-gathered data from a 2 m wingspan fixed-wing unmanned aerial vehicle flying at 30-120 m altitude over a 6.5 km trajectory.
1. Introduction
The recent explosion of interest in low mass unmanned aerial vehicles (UAVs), driven by technological improvements such as increased battery capacity and the availability of cheap open-source autopilots, will revolutionize a number of industries and sciences. Low-cost, rapid-deployment and close-proximity sensing give these smaller UAVs a distinct advantage over manned aircraft and satellite observation in many applications, and large-scale data gathering will soon be in the hands of the average consumer.
While the range of capability, payload capacity and mass of UAVs is broad (from 100 g multi-rotors to 5 ton military aircraft), there exists a distinct niche for UAVs in the range of 2-15 kg operating at altitudes of 30-150 m. These UAVs have a unique advantage in their ease of deployment complemented by a sufficient payload capacity for many scientific and commercial tasks. These tasks are far reaching and include agricultural monitoring, livestock management, search-and-rescue, geographic surveys and infrastructure inspection. Of special note, they have a distinct advantage in natural canyons and mountainous environments, particularly for accurate 3D surveying and reconstruction (Pollefeys and Gool, 2002), due to their manoeuvrability and range. In this 30-150 m altitude range, UAVs can generally avoid obstacles such as trees, power lines and small buildings while remaining below regulated airspace, typically 120 m (400 ft) above ground level (AGL), avoiding stricter rules and the need for more advanced sense-and-avoid technology. Additionally, this altitude region is well-suited to the sensing apparatus of low-mass UAVs, i.e. low-cost cameras in the visual and infra-red spectrum provide good resolution but remain small in terms of size and weight. Indeed, this is the region where many smaller commercial UAVs already operate (Vallet et al., 2011). We can consider these applications as low altitude in the context of general aerospace, but from a visual odometry (VO) perspective the required visual navigation and scene reconstruction tasks mean the problem can be considered a long-range sensing task. While a fixed-wing UAV in the 2-15 kg range can easily exceed altitudes of 600 m (∼2000 ft), for the purposes of this paper the focus is on the lowest range of this category. In contrast, most stereo VO demonstrations in the literature can be loosely considered to not exceed a scene depth of approximately ∼80 m, often much less than this. We directly address situations in this paper with an average scene distance of 120 m, posing the problem as ‘long-range’.
Most small commercial UAVs in this class utilize a combination of the global positioning system (GPS) and inertial sensing to facilitate an accurate estimate of pose (Nielson et al., 1986; Kim and Sukkarieh, 2002; Chao et al., 2010; Sukkarieh, 2004). The two sensing modes complement each other well; GPS provides a low-frequency global position estimate, while inertial sensing provides a high-frequency update on orientation and acceleration for active control. However, inertial sensors are generally not capable of providing accurate estimates over long periods due to sensor bias and unbounded error growth. Without GPS or other external positioning technology to constrain both scale and orientation drift, inertial sensors quickly become ineffective at providing robust position and orientation estimates. In contrast, it is well known that GPSs and other global navigation satellite systems (GNSSs) are ineffective in many urban areas, natural canyons and similar environments where access to the sky is limited. In addition they are easily subject to jamming and interference, and can even be switched off at the whim of commercial and military operators, meaning GNSS systems still remain fundamentally unreliable as a pose sensor. A number of US government agencies have stated that reliance on GNSS alone is insufficient (Dillingham, 2013) and remains a barrier to integration of unmanned aerial systems (UASs) into US national airspace, even in the less-regulated airspace below 120 m (400 ft) AGL. This demonstrates that the commonly implemented methodology of GPS and inertial sensing will ultimately require the integration of additional sensors for redundancy and increased accuracy, if widespread civilian applications are to be realized.
Vision is a viable complement to the GPS/INS pose estimation framework; not only in providing redundancy and improved pose estimates, but also as an independent pose estimator. In the context of an UAV, the addition of vision to the sensor suite means a more accurate and robust system, particularly in cases where a GPS is unavailable or inaccurate. It also means that lower cost and less precise inertial and GPS sensors can be used for a similar level of total accuracy. In addition, visual information can be utilized to provide a unique secondary output; 3D reconstruction of observed environments (Hygounenc et al., 2004; Mirisola et al., 2007). In many low-cost UAVs such visual sensors are already available as they are usually the key sensor for information gathering in most applications.
One potential approach to using vision for pose estimation is to fuse inertial measurement unit (IMU) measurements with pose updates from a VO solution in a filtered framework (Sukkarieh, 2004; Blosch et al., 2010; Weiss et al., 2012) occasionally in addition to a GPS (Bryson et al., 2009). In some cases, inertial data is fused with visual odometry updates in a loosely-coupled framework (Achtelik et al., 2011; Weiss et al., 2011), where the visual system is treated as a black box and pose updates from this system are fused with those from the IMU. Other systems incorporate feature tracks directly with inertial data to produce a tightly coupled framework (Li et al., 2013; Li and Mourikis, 2013; Weiss et al., 2013).
If we consider a purely visual approach, a robust VO system is capable of estimating pose with lower drift than most inertial sensors on their own, and with a stereo rig (two rigidly linked cameras capturing images synchronously) is capable of providing metric (accurately scaled) pose. Stereo VO has been demonstrated in a number of ground based scenarios (Konolige et al., 2007; Warren et al., 2010; Rehder et al., 2012; Lategahn and Stiller, 2012) on trajectories on the order of 10 km. Additionally, the imagery can be used for loop closure in order to build a full vision-only simultaneous localisation and mapping (SLAM) solution (Konolige and Agrawal, 2008; Konolige et al., 2010). In this paper, we focus exclusively on a vision-only system. While we anticipate that a fully integrated system would take advantage of vision, IMU and GPSs in a filtered or bundle adjusted framework, and see this as future work, in this paper we are interested in pushing and establishing the limits of stereo vision as a redundant pose estimator, specifically in long-range/short-baseline conditions in GPS-poor environments.
As discussed, vision systems have been demonstrated on UAVs using multi-sensor fusion algorithms or purely visual schemes, however, the application of VO at relatively high altitudes is not straightforward. A typical landscape is approximately planar when viewed from an altitude of ∼100 m or greater, and this can lead to degeneracies in monocular schemes. For stereo VO, maintaining accurate scale becomes more difficult as range to observed structure grows; i.e. it becomes increasingly, weakly observable from the small baseline-to-depth ratio (the baseline of a stereo rig compared to the average scene depth) (See Table 1). This effect means the utility of a stereo rig degenerates to an approximately monocular perspective with increasing altitude, causing triangulation from the stereo rig to become ineffective below a certain baseline-to-depth ratio, subject to focal length and camera resolution. This is demonstrated in Figure 1, where varying baselines are compared for their depth resolution and pixel disparity with increasing range from the observed scene.
Common depth ranges and baselines for stereo VO.

Depth versus depth resolution (Z-axis steps between discrete depths in pixel disparity) and disparity for selected baselines of a perfect stereo rig with perfect matching and normalized focal length, f = 1200, pixels for each camera.
Additionally, dealing with the potential high-vibration environment of a typical UAV means that the stereo transform (the physical relative rotation and translation between the two cameras of the stereo rig) is not necessarily constant due to physical deformation of the rig between the cameras. Even small changes in the relative orientation between the cameras from the expected values can mean extremely large triangulation error and rapid failure of a stereo VO system, if using a fixed calibration estimated preflight. The ideal stereo rig for a UAV from an engineering perspective is low-mass (hence low stiffness), the opposite to that desired to maintain an accurate stereo transform, and short baseline (low volume), and therefore less effective at long range. However, if the strict dependence on a fixed stereo transform is removed the engineering options are relaxed, meaning a more cost effective implementation and greater reliability of the VO algorithm. This paper presents a novel stereo VO with the ability to perform accurately scaled visual odometry in this scenario without the need for additional sensors.
1.1. Related work
VO has been applied in a variety of environments and on a number of vehicle types (Nistér et al., 2006), including stereo VO in ground vehicles in urban (Lategahn and Stiller, 2012) and natural scenes (Konolige et al., 2007), underwater (Pizarro, 2004) and hovering vehicles (Blosch et al., 2010), lighter than air (Lacroix et al., 2001) and fixed-wing vehicles (Warren et al., 2012). A number of implementations have been described, from purely visual methods with a bundle-adjusted optimizer, to filtered solutions with an intrinsic motion model (Lategahn and Stiller, 2012). More recently, applications have been presented where the distance between sensor and observed scene is significantly increased beyond the normally presented literature (Rehder et al., 2012). This has a unique effect, particularly in filtered solutions (Sibley et al., 2005), of overestimating scene depth and underestimating camera motion due to the non-Gaussian triangulation error which causes poorly scaled solutions.
In airborne applications, VO has received increased attention in recent years with the advent of multi-rotor UAVs and low-cost autopilot systems. Methods include both monocular and stereo VO (Kelly and Sukhatme, 2007; Stefanik et al., 2011; Jung et al., 2003; Lacroix et al., 2001), and a number of methods integrate inertial measurements in a filtered solution as described at the start of this section. Almost all airborne vehicles have an on-board inertial navigation system (INS) meaning that these filtered solutions, utilizing both visual and inertial measurements, are common. Additionally, monocular methods tend to be dominant in airborne applications due to the practical issues of fitting a stereo rig with a reasonably sized baseline to the platform. Weiss et al. (2013) have noted the deficiencies of stereo rigs at small baseline-to-depth ratios, where the utility of two rigidly-fixed cameras can be considered to reduce to an effectively monocular perspective beyond a certain range. Stereo VO methods have been restricted in the air to small-scale trajectories (Jung et al., 2003; Kelly and Sukhatme, 2007; Eynard et al., 2010; Stefanik et al., 2011), with the major demonstration consisting of a maximum altitude of 40 m and a trajectory of 230 m (Lacroix et al., 2001), despite a large (2 m) stereo baseline. A purely visual stereo VO has not been demonstrated over a large trajectory or small baseline-to-depth ratio in an airborne application.
1.1.1. Bundle adjustment and stereo calibration
Bundle adjustment has existed for decades as a specialization of least-squares optimization to the case of projective geometry (Triggs et al., 2000), optimizing 3D structures and camera motion by minimization of reprojection error (Engels et al., 2006). Many large-scale reconstruction techniques leverage this methodology to generate extremely large 3D reconstructions from crowd-sourced datasets (Snavely et al., 2007; Agarwal et al., 2010, 2011). Most of these more general implementations of bundle adjustment attempt to optimize not only scene structure and camera pose, but camera intrinsics including distortions and focal lengths (Pollefeys et al., 1998) due to the fact that each image is captured by physically different cameras with different intrinsic and distortion parameters.
In the more specific application of visual odometry, bundle adjustment is often implemented with fixed intrinsics to reduce computation and improve least-squares convergence (Konolige and Agrawal, 2008; Warren et al., 2010). In cases with a stereo pair, the stereo transform is often also assumed fixed and is dependent on rigid geometry, often calibrated from a standard offline checkerboard method (Zhang, 1999; Bouget, 2010) where the views are strongly overlapping. Some methods exist for calibration with partially overlapping views (Clipp et al., 2009), but many also specifically refer to cameras where the views do not overlap. Carrera et al. (2011) use MonoSLAM as the pose prior to estimating the rigid transform of a nonoverlapping pair, and Lebraly et al. (2011) use an initial hand-estimated calibration optimized via a modified bundle adjustment algorithm similar to the one presented here. Both of the methods used by Carrera et al. (2011) and Lebraly et al. (2011) require either specific motions or known scene geometry to complete the calibration. All these calibration methods, however, require some element of offline preestimation of the stereo transform before VO is performed, where it is then assumed fixed.
Some previous literature has addressed the issue of continuous or discrete recalibration of a generally fixed stereo transform or other intrinsic parameters in online data. Particularly, Dang et al. (2009), who has shown that it is possible to recalibrate the stereo transform online, however this method depends on a known motion model. Alternatively, Pettersson and Petersson (2005) have shown a calibration method using field programmable gate arrays (FPGAs), but constrained the problem by reducing the degrees of freedom. Our method does not require a motion model and calibrates all six degrees of freedom of the stereo transform. Previously, we have shown that even under high-vibration conditions the intrinsics do not significantly change (Warren et al., 2010), and the assumption that they remain so over a long trajectory is valid. Our method is focused on the optimal solution to stereo calibration of cameras with a high degree of overlap, and specifically in the online case with no other requirement other than a previously known estimation of the stereo transform that approximates the current truth. This is a feasible assumption as, unless there is catastrophic damage to a camera rig, the individual parameters of the stereo transform will not change by more than a few tenths of millimeter or degree.
Building on the basic methodology of bundle adjustment, many methods seek to introduce additional efficiencies based on specific problems or alternative parameterizations (Sibley et al., 2009). Some attempt has also been made in recent years to introduce additional objectives, such as nonvisual sensor readings, into the optimization (Michot and Bartoli, 2010) by augmenting the traditional cost function. By determining appropriate weights for each type of objective, a solution can be found that attempts to find the optimum of these multiple sources. Given a robust estimate of variance on sensor measurements the method is straightforward, however, in the projective case there is often no such variance knowledge and the selection of weights is a more difficult task.
This objective-based solution is in contrast to more traditional filtered solutions, which tend to treat the pose updates from VO as a black box in a loosely coupled configuration (Lategahn and Stiller, 2012) similar to an inertial sensor, often ignoring the information rich feature projections given by visual tracking. For this reason, bundle adjustment can be considered as a more accurate solution given the same data (Strasdat et al., 2010), where all sensor measurements are tightly coupled, but at a larger computational cost. With increasing compute power and recent progress in efficiently sparsifying the problem, bundle adjusted methods are approaching real time implementation (Klein and Murray, 2007).
In contrast to an objective-based optimization with an augmented cost function, a constrained optimization attempts to apply rigid bounds to the value of optimization variables. By applying both soft and hard constraints, it is possible to ensure that a poorly observable variable or a poorly initialized solution converges appropriately when an unconstrained solution will not. Lhuillier (2011, 2012) attempted to incorporate these constraints for VO by using a projection-only bundle adjusted pose estimate, then added GPS-based pose constraints and reoptimized the final solution offline. In our work we integrate the constraints into the online bundle adjustment algorithm itself, but with a different parameterization. Instead of implementing a constraint based on a global position estimate, as in Lhuillier (2011), our method implements a constraint on the allowable motion of the stereo transform to account for deformation between the cameras. We implement a bundle adjustment optimization that includes all available cameras in a window of recent frames, and optimize the stereo transform in addition to the base camera poses and scene structure. Using the extremely strong prior that the deformation of a stereo rig is small and will not move more than a few millimeters or tenths of a degree, we implement a cost-based bound on these variables to ensure an adjustable, but tightly constrained stereo transform.
1.2. Contributions
This paper advances the state-of-the-art in vision-only navigation with a new algorithm for stereo VO that is effective at long-range and explicitly accounts for nonrigid stereo rigs (where the relative pose between the two cameras is known to deform), complemented by experimental results. The main contributions are described below.
An examination of the effects on triangulation performance of stereo vision at long-range, including the impact of deformation and non-Gaussian bias as range increases to justify the use of a bundle adjusted optimizer.
A novel modified bundle adjustment algorithm that includes the geometry of the stereo pair as an additional set of optimized variables, plus the inclusion of constraints to ensure accurate scale even under deformation.
A new VO algorithm capable of handling long-range stereo by implementing a novel initialization step to avoid direct stereo triangulation.
Flight test results that apply the VO algorithm to gathered airborne imagery. Results are compared to an INS/GPS solution to demonstrate the relative accuracy over the trajectory.
The rest of this paper is organized as follows; Sections 2 and 3 examine the challenges of stereo at long-range on a UAV. Section 4 details the visual odometry methodology, including the long-range stereo initialization and the specific bundle adjustment implementation used in our algorithm. Section 5 details the UAV used for experimental validation and the dataset gathered by the platform. Finally, Section 6 shows results of the algorithm applied to this visual data. The paper is concluded in Section 7.
2. Characterizing the impact of rigid deformation on calibrated rigs
An accurately estimated stereo transform is the key requirement to achieve accurate triangulation in stereo VO. We highlight these problems for stereo VO in aerial applications by showing that small errors in the estimated stereo transform result in poor triangulation. Typically, in the case of near horizontally aligned rigs with image sensors in a near coplanar configuration, the incoming imagery is subject to both undistortion and rectification steps (the removal of lens distortion, then horizontal pixel row alignment), performed as a single image warp from a combination of the two 1 , to simplify feature matching along pixel rows so that the search space for matches is minimized. In most cases, after undistortion and rectification, the pixel rows are aligned with a maximum vertical alignment error on the order of 0.5 pixels or less and matching is only performed on a small number of adjacent rows of the corresponding image. However, if the physical stereo transform changes due to rig deformation, the pixel rows will be poorly aligned, causing a reduced average number of feature matches between the pair. If the matching distance (the number of rows above and below the current row to search for a match) is relaxed to counteract this deformation, a higher noise in triangulated scene structure will result. Both a reduction in reliable feature tracks and noisier 3D triangulation can adversely affect the robustness and accuracy of any stereo VO algorithm.
For many VO algorithms demonstrated on ground vehicles and on many commercially available stereo heads, it can be assumed that the calibrated estimate of the stereo transform remains pixel-accurate for the lifetime of the dataset or product. In these cases there is insufficient stress and vibration to change the stereo transform, making this a safe assumption. However, in cases such as the high-vibration environment of an aircraft the assumption of no deformation between the cameras can no longer remain. This is a fundamental weakness in stereo VO; the epipolar geometry is always assumed to be accurately calibrated to within a few pixels. Even an extremely small orientation change can mean the epipolar geometry (projection of the ray defining a feature in one image into another) is adversely affected. For triangulation purposes, this deformation can cause a rapid degeneration such that VO is affected, causing unreliable and failed pose updates.
To quantify this error we can examine a simulated rig, subject to deformation, and determine the impact of changing the relative pose of the cameras in terms of the ‘epipolar error’; the Euclidean distance between where a feature is projected in one camera of a stereo pair, and the 2D epipolar line in that image defined by the feature’s projection in the other camera (Figures 2 and 3). Examining the case of horizontally coplanar stereo (Figure 2), it becomes clear that some degrees of freedom on a camera’s position cause more reprojection error than others of the same magnitude. Pitch on one camera, for example (around the base camera’s x-axis), causes a radical deviation in epipolar error compared to say, yaw (around the base camera’s y-axis). Intuitively, this same observation can show that some dimensions are more observable than others, and this has an impact when we are trying to calibrate a stereo rig.

The epipolar geometry of a horizontally aligned stereo rig. On close inspection, a distance error exists between the projected epipolar line of a feature and its actual position in the second image.

The epipolar geometry of a vertically aligned stereo rig. On close inspection, a distance error exists between the projected epipolar line of a feature and its actual position in the second image.
For the purposes of the airborne experiments presented later in this paper, we consider a stereo rig that has vertically aligned cameras (Figure 3), rather than the more often considered horizontally aligned case. We simulate a camera pair with a 0.7 m baseline and other properties similar to the actual cameras presented later in this paper. The pose of the second camera is adjusted so that it deviates from the coplanar, perfect, case and the corresponding epipolar error is examined. The results are quantified by observing a 3D point 100 m along the optical axis of the unmodified camera’s position and projecting back into the ideal and nonideal stereo camera positions. These results are shown in Figures 4 and 5.

Characterizing the reprojection error (distance from the true epipolar line) through modification of the stereo transform parameters for a vertically aligned stereo rig with 0.7 m baseline. The x-axis describes the deviation of the translation (left) and rotation (right) for two selected axes from a known calibration. By increasing this error for a single axis, the deviation from the expected epipolar geometry is characterized through the projection error.

Characterizing the triangulation error (distance from the true 3D point) through modification of the stereo transform parameters for a vertically aligned stereo rig with 0.7 m baseline. The x-axis describes the deviation of the translation (left) and rotation (right) for two selected axes from a known calibration. By increasing this error for a single axis, the Euclidean distance error of the retriangulated point from the observed position is characterized.
As can be seen, a translation error along the y-axis of 50 mm or more (highly unlikely) can cause only a sub-pixel projection error. However, rotational error about the y- (in this case, yaw) axis has a very significant effect. Even a deformation of 0.2° causes a 3-pixel error. For the case of a rig of 0.7 m baseline, such an error is not outside the bounds of reality. If a robust matching scheme has a strict 1-pixel allowable deviation, this will render many features as untrackable.
Progressing further, this projection error manifests itself more readily in 3D triangulation error (Figure 5). For a rotational error of just 0.2° around the most sensitive y-axis, the triangulation error in this example exceeds 30 m. Such triangulation error, over several hundred 3D points, will impact the recovery of a new camera in a pose update step. If, like in most stereo schemes, structure is not triangulated over more than 1 or 2 time-steps, this error can rapidly result in VO failure. In this paper, a ‘near optimal’ VO is performed that tracks and triangulates features over all frames in a large sliding window, rather than across a single stereo pair. Not only does this reduce triangulation error through observations from a large pseudo base-line and multiple image views, but also assists greatly in convergence of the bundle adjusted optimization that can account for stereo deformation.
This rather simple analysis is warranted as it shows that directly recovering sparse 3D structure from a single, deformed, stereo pair with only a small calibration error results in a high degree of triangulation error. More importantly, it also quantifies the sensitivity of the rotational components of the transform in comparison to translation. From this analysis, there is often a single rotational axis that results in the highest quantified error and greatest observability of deformation. We will examine this effect in the results presented later in this paper.
In summary, any deformation of a stereo rig operating at long-range means that a VO algorithm that assumes stereo rigidity will almost certainly fail in only a few frame updates. For this reason it is impossible to perform a ‘standard’ rigid stereo VO on the data presented in this paper, and hence such results are not included.
3. Accounting for long-range bias in stereo vision
For long-range stereo, it is also important to consider the effect of a non-Gaussian triangulation bias on the recovery of scene depth and consequently camera motion. It is well known that in long-range stereo, a depth bias exists for short camera baselines that grows with increasing distance of the triangulated scene. At disparities on the order of 10 pixels or less, triangulation from a single stereo pair tends to follow a non-Gaussian curve with a long tail (Sibley et al., 2005), meaning that the average depth of structure is overestimated and, consequently, relative camera motion is underestimated. In filter-based applications, scene points are often triangulated only from individual stereo pairs (rather than over several frames from the pair). As a consequence, it is important to address non-Gaussian bias induced by this scheme by applying a reduction factor on the depth of triangulated structure, calculated empirically from in-the-field data. The bias is acute in these situations due to the fast marginalization of structure from the filter, causing an artificially short baseline for triangulation. In a bundle adjusted solution (that incorporates observations from multiple views, not necessarily just between the stereo pair) we can observe the effect of this triangulation bias on the solution and hence include an analysis examining this to justify the use of bundle adjustment.
To quantify this effect a Monte-Carlo triangulation simulation is performed with a pair, triplet and quadlet of cameras arranged in a horizontally aligned set, much like a set of traditional stereo cameras. In this simulation, the depth bias can be shown clearly (Figure 6) by observing a set of points at a range of 20-120 m. The simulated cameras share the parameters of the cameras in our experiments (6 mm focal length lens, 1024 × 768 pixel cameras and 0.7 m baseline), and the estimated 3D points are triangulated with 1 pixel Gaussian noise (note the demonstrated bias is in the triangulation, rather than the observation noise). Observing the set of points over the 20-120 m range, we can extrapolate the mean of this bias at each depth and compare the effect with an increasing number of observing cameras. These results (seen in Figure 6) show that even with a small number of additional observations, with an overlapping effect from an increased baseline and a doubling of triangulation constraints, the bias for a feature observed with four cameras is less than 0.5 m at 120 m range. Of course, the biasing effect is most pronounced for a single pair, where a nearly 3 m distance bias exists for features at 120 m range.

Characterisation of range bias (a) for sets of 2, 3 and 4 cameras with a 0.7m baseline between each adjacent camera, and their associated disparities (b). Range bias denotes the triangulation error averaged over 100,000 trials at each observed distance.
This non-Gaussian triangulation bias can be directly observed by plotting the same Monte-Carlo simulation results for retriangulation of a subset of observed 3D points. To exaggerate this bias for visualization, we reduce the baseline of the cameras but triangulate over the same range. The results for two, three and four cameras at 0.2 m spacing between each adjacent camera can be seen in Figures 7a, 7b and 7c. It can be seen that with four cameras, giving a pseudo-baseline of 0.6 m, the triangulation approaches a Gaussian curve, but with two cameras the curve is distinctly non-Gaussian.

Binned distributions of triangulated depth for a point at (a) 40 m, (b) 60 m, (c) 80 m and (d) 100 m from a set of cameras given a Gaussian 1 pixel projection error. Each distribution represents 500,000 trials for sets of two, three, four and five cameras with a 0.2 m baseline between each camera. Independent of distance of a scene point to a set of cameras, the addition of each camera and the extension of the baseline causes the range estimation to approach a Gaussian distribution.
This result highlights the difference between the output of most filtered and bundle adjusted VO solutions and will justify the intended approach presented later in this paper. At long-range, distance bias is an important factor to account for, and this has already been performed explicitly in a number of filtered applications. However, due to the large-pseudo baseline generated in a bundle adjusted solution by the triangulation from cameras through time, where links are not marginalized out quickly, this bias can be safely ignored at long-range. This reflects the overarching paradigm of long-range stereo VO in this paper; rather than relying on the stereo geometry of a rigid camera pair for scene triangulation and pose updates, the problem is treated more akin to a monocular VO where the large pseudo baseline of a camera moving in time is used for more accurate triangulation. Instead, the baseline of the stereo pair serves more as a scale constrainer and additional feature observer.
4. Long range stereo VO with constrained stereo bundle adjustment: Algorithm overview
The investigation performed above shows that the dominant issue for long-range stereo is rig deformation that affects relative camera rotation. Small UAVs often have requirements of small payload weight, so a stiffer or stronger rig may not be possible. The problem of deformation is accounted for directly in the bundle adjustment optimization presented here, allowing relative motions of a few tenths of degrees or millimeters in the geometry without adversely affecting pose. The following section outlines the three major modifications incorporated into the long-range stereo VO routine:
Initialization of scale for small baselines;
Inclusion of the rigid geometry in the bundle adjustment optimization;
The addition of constraints to the optimization to ensure weakly observable parameters such as the translational components of the stereo baseline (and consequently the scale) remain within a known region.
For completeness, the standard parts of stereo VO that are incorporated into the presented algorithm are repeated here. Our long-range stereo VO algorithm follows the general principle of most VO algorithms; iterative updates through feature tracking, scene triangulation, pose extraction and finally, a windowed optimization via bundle adjustment on the most recent frames. The entire algorithm is summarized in Figure 8, however, the loop-closure and subsequent pose-graph optimization is not addressed in this paper.

Overview of the long-range stereo VO algorithm.
Feature ztracking is performed by detecting and matching SURF (Bay et al., 2006) features between sequential frames and across stereo pairs using loosened epipolar constraints. Upright descriptors are utilized for quick matching between the typically high overlap (>90%) imagery. In a deviation from traditional stereo VO, structure is not directly triangulated from stereo pairs due to the noisy estimates at long-range (as previously discussed). Instead, scene structure is triangulated from at least three sequential views. This ensures a wider triangulation baseline and enforcement of accuracy by the inclusion of at least a third view.
Pose updates are performed through the traditional 3-point algorithm (Haralick et al., 1994), inside a MLESAC (Torr, 2000) estimator for robustness, but only include observations from the base camera of the stereo pair to recover the pose. The pose of the secondary camera is then generated from the assumed stereo transform gathered from an offline, on-the-ground, calibration.
To account for long-range and stereo deformation, however, we implement several changes. A special initialization step is required more akin to monocular VO to reduce dependence on a functionally inaccurate stereo transform, while still maintaining the metricity enabled by the stereo rig’s baseline. We derive a bundle adjustment routine that explicitly optimizes over each of the cameras in the stereo pair. This is critical as the deformation of the rig in flight causes the rigid transform between the cameras to continuously change. Additionally, constraints are introduced into the bundle adjustment to ensure parameters such as scale remain within an allowed region, as the effective scale constraint (the baseline of the stereo rig) would have the freedom to drift without them. Conventional bundle adjustment routines that optimize only one of the cameras, and rely on the assumption of rigid geometry to include the second camera, will fail in this scenario. As outlined earlier, this failure will occur within a few frames.
4.1. Multi-camera bundle adjustment
Here, we rederive bundle adjustment taking into account the presence of multiple physical cameras in a stereo (or multi-camera) framework that capture imagery synchronously. The basic bundle adjustment framework as a least-squares implementation is not described here in detail as it is well covered in the literature (Hartley and Zisserman, 2004; Triggs et al., 2000). The modification presented here differs from the literature by taking into account the feature projections for all cameras in a set of two or more cameras. In addition, the transform between the base camera and additional cameras is parameterized in order to optimize it online. Some implementations have developed similar theory, but with differing outcomes. For example, FrameSLAM (Konolige and Agrawal, 2008) includes the stereo transform for feature tracking but implicitly assumes the stereo transform remains fixed.
We formally state the problem as: Given m (j ∈ [1,…,m]) scene points observed at n unique timepoints/locations (i ∈ [1,…,n]) by a single physical camera, the traditional model used for the projection of point j in space (
where

Monocular VO for a single camera moving through an arbitrary motion for four time steps, showing the setup of the monocular bundle adjustment problem.
Bundle adjustment attempts to minimize the sum of squares objective function by modifying the estimated camera poses
where
where
with
We can now extend the single camera case to two unique, rigidly linked cameras (k ∈ [0,1]), synchronized to capture imagery at the same time instant, and express the second camera in terms of the base camera via the stereo transform
additionally separating intrinsics

Stereo VO, with two rigidly linked cameras moving through an arbitrary motion for two time steps, showing the setup of the stereo bundle adjustment problem.
We explicitly note the cost function for purely projective optimization (taking account of the multi-camera case) is
In this implementation, there are a total of 3m+6n+6k variables to optimize, corresponding to scene points, base camera poses and the stereo transform parameters (shared among all poses optimized in a bundle adjustment step) respectively, if intrinsics remain static. In order to generate a generalized, but efficient, implementation we separate these variables into three distinct groups, scene points, shared and independent parameters; allowing the stereo transform to be shared amongst a group of camera pairs in a sliding window (in this case approximately 12-15 camera pairs) but allowing the poses at each time step to remain independent. The transform is shared amongst the the group of pairs to increase observability of the stereo parameters, meaning that extremely small vibrations are averaged out while the lower frequency, larger deformations are accounted for.
This parameterization allows an efficient implementation of the update Hessian. The separation of the variables results in a partitioned parameter vector (
Keeping this parameterization, we split the partial derivatives (the projections with respect to each optimized variable) into three separate groups:
and as a consequence of this and the modified parameter vector the augmented normal equations (equation (4)) have the following representation
Figure 11 shows the structure of the approximate Hessian matrix

An example of the block structure of the Hessian with two sets of two rigidly linked cameras and four points.
with the augmentation of the diagonals by multiplication of the Levenberg Marquadt damping parameter 1+λ resulting in
To solve the system efficiently, we reimplement the Schur complement matrix by making the following substitutions
which gives the simplified set of equations
Multiplying this system by
As seen in Hartley and Zisserman (2004), the top half of this set can be used to find
Then, through back-substitution, the point updates are
4.2. Constrained bundle adjustment
The stereo transform
Due to the rigidity of a well-engineered stereo pair, even under deformation, any movement between the cameras is physically restricted to at most a few degrees or millimeters. This is an extremely strong prior that can be used to constrain the optimization. Here, we take advantage of this to encode a strictly feasible region for some of the parameters that represent this deformation. To demonstrate the implementation of this restricted space for a feasible solution, constrained optimization must first be examined.
4.2.1. Constrained optimization
In many optimization algorithms, including most implementations of bundle adjustment, the optimized variables and the cost function that integrates them is allowed to span the entire space of real numbers. However, in many cases this is neither feasible nor practical. Some optimizations need to account for physical realities (such as a circle needing a non-negative diameter), and hence implement an inequality constraint to define the feasible region for a variable. A general formulation of constrained optimization can be setup as follows
That is, minimize f(x) subject to the constraints ct
(x) = 0 and ct
(x) > 0, where
This can be visualized in Figure 12. In this case the optimal solution is x 1 = 2, x 2 = 1. For the case of linear constraints, the constraint equations are of the form

Minimization of the function
where b is termed the barrier value. To implement a specific region of feasibility, two constraints can be implemented for a single variable, denoting the barrier terms b 1 and b 2 as the lower and upper limits, respectively. Implementing this barrier is straightforward in a generalized optimizer, and we utilize a penalty function to augment the cost term as a method of implementing this barrier. Using equation (6) as a basis, a logarithmic barrier function can be integrated as a soft constraint in an augmented cost function to ensure the optimization of specific parameters remains within specific bounds, i.e.
where μt is termed the barrier parameter and is used to tune the cost as the parameter approaches the barrier (Figure 13). The augmentation of the cost is negative due to the nature of the log-barrier function. We can see what will happen to the cost as the optimizer progresses; as a variable x approaches b the cost term (log ct (x)) grows and the influence of the projective observations on the parameter reduces. A step that takes x beyond b will also yield an infinite cost and hence not be updated in a bundle adjustment step. Instead, a Levenberg-Marquadt optimizer will increase the damping parameter to instantiate a smaller parameter step on the next iteration.

The log barrier cost for varying values of μt , within the barriers p − q and p+q.
Alternative constraint implementations (such as gradient-projection) exist, but we utilize a soft cost-based implementation due to its simplicity. Additional important details for implementing a constrained optimizer, such as the Karush-Kuhn-Tucker (KKT) conditions of optimality and ensuring smoothness can be found in Nocedal and Wright (2000).
4.2.2. Constrained stereo optimization
As stated previously, we know that for a stereo rig the deformations are small, and therefore the variables for rotation and translation between the two cameras can be defined in a strictly feasible region. The stereo transform is parameterized as six variables (x, y, z, roll, pitch, yaw), yielding 12 constraints (two per parameter, defining an upper and lower bound).
From a known calibration, such as that generated from a set of checkerboard images, we implement the feasible region based on the initial value p of a parameter (such as translation along the x-axis) plus a bound ±q (see Figure 13). The resultant barrier values p − q and p+q are synonymous with the barrier value b of equation (17), defining the equivalent barrier terms b 1 and b 2. In this paper, the magnitudes of q for each variable are empirically evaluated. They could alternatively, for example, be estimated via an analysis of material expansion based on temperature or elasticity of the material under load.
The log barrier cost from equation (18) is also integrated into the bundle adjustment Jacobian and used to augment the relevant parameters of the stereo transform
where the first term incorporates the Jacobian with respect to the shared stereo parameters, and the second term incorporates the additional term in the Jacobian generated from the barrier function.
Using this methodology, it is now possible to optimize for 3D scene, base camera poses and the shared stereo transform under strict constraint conditions. This allows the bundle adjustment algorithm to converge appropriately without the stereo parameters leaving the feasible space. Given that the stereo transform defines the scale of a solution via the length of the baseline, constraints that restrict this transform ensure that accurate scale is maintained even when the distance to the scene means the observability of the stereo baseline is weak.
4.3. Long range VO initialization: Scale without stereo triangulation
In order to set-up the iterative VO algorithm, an initial estimate of pose and 3D scene is required. In standard stereo VO, scene points are initially triangulated from the calibrated stereo pair, hence there is no need for a special initialization step. However, we have noticed that at large depths such triangulation is inaccurate and stereo rig deformation may render triangulation impossible due to the poor epipolar geometry. Therefore, a solution that retains or maintains accurate scale is needed for camera pose without initially computing structure from a fixed pair, more akin to monocular pose initialization. A boot-strapping procedure is presented here that ensures that accurate triangulation is achieved from a wide-baseline pair and is not dependent on the geometric stereo transform.
The initialization procedure occurs as follows (Figure 14). Firstly, an essential matrix

The long-range stereo initialization routine. A monocular pose estimate on a set of base camera images is followed by adding the secondary camera and optimizing for the scale constraint κ. This is followed by the application of
We have attempted to perform an initialization from just a triplet of images, solving for the scale by including the stereo baseline in a method of vector addition. However, it is clear from experimentation that this methodology is insufficient as the baseline remains weakly observable with only three observing cameras and hence the estimation of the solution suffers. As highlighted in Section 2, the translation between just two cameras at the small baseline-to-depth ratios in this paper is not particularly sensitive to projection error, and therefore difficult to optimize accurately. It is only with a large set of observations from multiple camera pairs that enough information is available to achieve accuracy. We don’t use a trifocal tensor as there is no implicit method to integrate scale into the calculation that yields a satisfactory linear solution.
To increase the information and ensure robustness, a monocular VO is then performed on the initial arbitrarily scaled pair of base cameras (rather than the stereo pair) extracted from
Importantly, the scale of this initial set of pose estimates is arbitrary, initialized so that the first camera pair has a distance of unity to ensure stable numerical precision, irrespective of the true distance between camera poses. In order to accurately recover scale, the stereo baseline must now be included (in lieu of other scale measurements from an IMU or external reference) by utilizing feature projections from the secondary camera to provide a geometric constraint on the overall scale of the scene. We also introduce an inverse scaling term κ to equation (5), that allows the translational component of the stereo transform

Scaling of the stereo transform via the scale term κ.
Due to the normalization of the monocular initialization, the scale of the scene is arbitrarily set to unity, meaning that the reprojection error on the secondary camera will be high. By optimizing this κ variable in addition to the scene points, base camera positions and stereo transform, any discrepancy in scale of the scene defined by adding the secondary cameras is handled efficiently without rendering the bundle adjustment problem as poorly initialized.
The set of all base cameras in the initial set and their corresponding secondary cameras are then optimized along with scene structure in a batch bundle adjustment that includes the scale term κ and the stereo transform subject to the boundaries mentioned in Section 4.1.
Once the optimization routine has converged and the scale parameter κ is well estimated, the entire scene and camera geometry is scaled by
4.3.1. Simulated initialization experiments
To verify the utility of the modified initialization scheme and of introducing the scale parameter to the optimization, several repeated simulation experiments are performed over two sets of variables:
Modifying baseline-to-depth ratio;
Modifying initial scale incorrectness.
To introduce a control for the new algorithm, the initialization is also performed in each separate experiment using: (a) a standard stereo VO pipeline (that attempts to perform stereo optimization on the first frame-pair); and (b) using the modified initialization scheme without explicitly optimizing scale. In each experiment, a pair of simulated cameras is flown over a simulated 3D scene generated from previously gathered LiDAR data (Butler and Loskot, 2014) (Figure 16), using the projections of the 3D points as simulated tracked features. In all experiments the cameras have a 0.7 m baseline, reflecting the data presented later in this paper, and are flown at speeds ranging from 15 m/s to 50 m/s to ensure image overlap, mirroring the motion of a fixed-wing UAV model and approximating a realistic field scenario. The total linear distance traveled by the cameras in each experiment is approximately 60 m, independent of flight speed and frame coverage. Additionally, for consistency, the number of features tracked per frame is kept approximately equal independent of altitude. Each simulated image has a resolution of 1024×768 pixels and each feature is projected with Gaussian noise of 1.0 pixels variance, and no erroneous matches. The simulation is performed using both the modified and control algorithm implementations for 20 trials for each value of the modified variables, and the final pose error compared to ground-truth recorded at each iteration.

The simulated scene.
In the first simulated experiment, the initialization algorithm is evaluated over multiple trials at a range of baseline-to-depth ratios. The number of initialization frames during the scale optimization is fixed at ten. By varying the flight height from 10 m to 100 m altitude, the baseline-to-depth ratio is decreased from a value of 0.07 to 0.007, an order of magnitude change that covers (at lower altitude) most common stereo VO demonstrations in the literature to (at high altitude) an approximation of the field-data scenario that is covered in this paper.
In the results for this experiment (Figure 17), the Euclidean distance between the final camera in the initialization and that of the ground truth camera is used to define the final pose error. As reflected in the large variance and high average pose error of the control (conventional stereo VO) above altitudes of approximately 40 m, or a baseline-to-depth ratio of ∼ 0.02, the standard stereo pose estimator fails to consistently recover a sufficiently accurate pose. In many cases during simulation, the standard stereo estimator fails before the minimum number of frames is completed due to the poor triangulation and subsequent lack of viable feature tracks, in addition to forming degenerate pose updates that do not follow the known motion of the cameras. As the altitude approaches 100 m, the conventional estimator consistently fails, approaching the limit of error at 60 m. This is reflected in the high error, low variance result at the highest altitude.

Final translational error for varying flight altitudes (in metres and %), calculated from the Euclidean distance between the final camera poses of the VO and ground truth, performed over 20 trials with mean and variance plotted. The dashed line represents conventional stereo VO (Warren et al., 2010) (no explicit computation of κ), the solid line represents modified long-range VO. Note that the limit of the length of the data set is 60 m, reflected in the largest total translation error at this limit.
In contrast, the long-range estimator with scale optimization shows robust performance even at altitudes of up to 100 m, achieving an accurately scaled pose within 5 m over the trajectory in all trials at the highest simulated altitude. In this configuration, according to Figure 1, a 100 m altitude corresponds to a disparity of ∼9 pixels and a depth resolution of 12.5 m from single pairs. As is clear from this analysis, triangulation from a stereo pair for a conventional VO is difficult and will generally fail in the first few frames at long range.
In the second simulated experiment, the camera pair is flown at a fixed 70 m altitude for the same 60 m distance for each trial, with a fixed flight speed of 30 m/s. In this case, however, the optimization step is artificially initialized at a range of arbitrary scales, defined by the Euclidean distance between the first monocular camera pair (see Figure 14). By varying this initial value, the ability of the modified bundle adjustment algorithm to converge in the presence of a poor initial scale is examined. Without the optimization of κ to accommodate the scale of the scene, bundle adjustment is rendered poorly initialized and should only converge if the initial scale closely approximates the truth. In contrast, with the addition of κ to the optimization, bundle adjustment should converge from a wide range of initialization scales. For this experiment, the algorithm with and without scale optimization is run for a range of initial scale errors (a range of 0.25 to 1.5), with 20 trials performed at each step to examine consistency of performance.
In the results for this experiment (Figure 18), the scale error (expressed as the ratio between ground truth and recovered distance) before and after optimization is demonstrated on the long-range VO algorithm. A direct comparison is made between convergence with explicit κ optimization, and without. As seen, the addition of explicit scale optimization ensures the algorithm converges appropriately at all examined scale ratios (0.25↔1.5). However, without the addition of the κ term the standard bundle adjustment algorithm fails to converge reliably except within a range of ∼10% of the true value. This demonstrates the necessity of the addition of the scale term to improve robustness of the algorithm, ensuring a far more reliable convergence at a wide range of unknown initial scale values.

Scale error comparison before and after optimization, and with and without explicit scale optimization; performed over 20 trials with mean and variance plotted. A scale of unity is considered optimal. The dashed line represents bundle adjustment without κ-optimization and the solid line represents bundle adjustment with κ-optimization.
5. Dataset
To evaluate the performance of the algorithm, it is applied on a stereo dataset gathered from a radio-controlled fixed-wing unmanned aerial vehicle (UAV) (Figure 19) with a 2 m wingspan.

The experimental platform showing component layout. The diagram on the fuselage indicates length and orientation of stereo baseline and on-board cameras.

An example set of images from the dataset for (a) the rear camera,and (b) the front camera (see Figure 19), rotated so that the epipolar lines are approximately horizontal. (c) The inset region of the rear image, showing selected features (circles) and the pixel offset of their match in the front image (solid lines). (d) The inset region of the front image, showing the same selected features (circles) and the pixel offset (solid lines), but also highlighting corresponding epipolar lines from the matches for the original, erroneous, calibration (dashed lines), and the optimized calibration (dashed lines intersecting circles).
The aircraft includes an off-the-shelf mini-ITX computer system for logging both visual and inertial data, and a pair of Point Grey Flea2 IEEE1394B colour cameras, rigidly fixed via an aluminium L-bar situated inside the fuselage of the aircraft. The cameras face down towards the terrain with the baseline in line with the fuselage, as seen in Figure 19. Each camera uses a 6 mm lens giving the 1/3” sensor a field of view of approximately 42°× 32°. The cameras are calibrated before flight using a checkerboard pattern to achieve a standard intrinsic parameter calibration and approximated stereo transform between the cameras.
An XSens MTi-G INS/GPS system is used as the ground truth measurement system on the aircraft, with a manufacturer claimed positional accuracy of 2.5 m circular error probability (CEP). Size and weight restrictions prevent the use of more accurate differential global positioning systems (DGPSs), however, the MTi-G itself provides a reasonably accurate estimate of pose over broad scales. The MTi-G unit is rigidly attached to the onboard camera rig, while the GPS receiver is installed in a rigid configuration directly above the front camera.
Data was collected over approximately 5 minutes of flight, at an approximate flight speed of 20 m/s, over a ∼30-100 m altitude range. The path flown is shown in Figure 21. Bayer encoded colour images are logged at a resolution of 1280 × 960 pixels at 30 Hz and later converted to colour for processing. GPS, unfiltered IMU data and filtered INS pose were recorded at 120 Hz from the XSens MTi-G to give ground truth position and orientation comparison. The terrain of the area flown consisted of rural farmland with relatively few significant landmarks such as trees, animals or buildings.

The path flown as indicated by the on-board INS.
6. Experimental results
To demonstrate the utility of the algorithms, pose estimates are generated using two differing methodologies: (a) Stereo VO pose estimation with unconstrained stereo parameter optimization; (b) Stereo VO pose estimation with constrained stereo parameter estimation. For the constrained case, bounds are placed on the stereo transform of ±1 mm in translation and ±0.01 rad (∼0.57°) in rotation. While these numbers are estimated empirically, theoretical analysis can also be performed to examine effects such as thermal expansion or elastic and ductile deformation, if the period and strength of the vibration can be accurately measured. Often this information is unavailable or difficult to predict and therefore must be determined empirically, as is the case here.
Each stereo VO algorithm is applied over the same 9600 frames in the gathered dataset, comprising about 90% of the full trajectory. Data from takeoff and landing is excluded due to the proximity to scene rendering the algorithm unsuitable. In both cases, the long-range stereo initialization scheme is performed on the first ten frame pairs at a three frame subsampling (i.e. 10 evenly spaced frames from a set of 30), and is checked for accurate convergence before proceeding with the stereo VO algorithm.
In Figure 22, the pose estimate over the 6.5 km trajectory is shown for each case compared to ground truth from a top-down perspective to highlight the performance of the respective algorithms. In Figure 22a (left), it can be seen that the trajectory generated from the unconstrained optimizer suffers significantly due to scale drift and poor pose estimation over the trajectory. In Figure 22b (right), the addition of constraints on the stereo transform ensures that scale does not drift, and the pose estimate is significantly improved over the length of the trajectory. It is important to reiterate here that a pose estimator that does not estimate the stereo transform will almost always inevitably fail on the presented data due to the nonrigidity of the stereo rig conflicting with the fixed estimation, causing feature tracking failure. Therefore it is not represented in these results.

The final visual pose estimate in (a) the unconstrained and b) the constrained stereo compared with ground truth, seen from a top-down perspective. The plots are separated to aid visualization.
Due to the large overlap present in Figure 22, Figure 23 highlights the pose error in comparison to ground truth over the trajectory for the constrained and unconstrained cases. In particular, Figure 23 shows the recovered accuracy in orientation, reflecting the rapid roll and pitch motions of the aircraft during flight. These results indicate the ability of the algorithm to inform pose for future online control.

(a) The constrained visual pose estimate error in reference to ground truth separated into the x-, y- and z-axes. (b) The constrained visual pose orientation error in reference to ground truth separated into roll, pitch and yaw.
To highlight the effectiveness of the stereo optimization with and without these constraints, the estimated six degrees of freedom of the stereo transform during flight are graphed in Figures 24 and 25 for both the constrained and unconstrained cases over the trajectory. The implemented constraints are represented by dashed lines. As seen in Figure 24a, without an effective constraint on the translational components of the transform, the configuration drifts rapidly from the initial calibration and scale drifts significantly as a consequence. In several sections of the trajectory the unconstrained estimator fails to converge to a new value at the bundle adjustment stage, meaning that the stereo transform defaults to the original pre-gathered calibration. Due to the relative magnitude of the unconstrained case in Figure 24a, only the constrained translational components are presented in Figure 24b. Note in Figure 24b that the stereo transform is successfully held within the defined boundaries of the constraints, enforcing scale to remain close to the truth.

(a) The stereo transform translational components without constraints and (b) with constraints. Constraints on each variable of the transform are highlighted by a dashed line. Without constraints, the components of the stereo transform drift significantly in a random fashion, negating the important scale constraint the stereo transform provides, and the algorithm often fails to converge, reverting back to the original calibration. With constraints, the translational components are kept within the strict bounds and the scale is maintained.

The stereo transform rotational components with (in reference to Figure 22a) and without (in reference to Figure 22b) constraints. The boundaries of the constraints on each variable are highlighted by dashed lines. Due to the observability of rotational components, there are no consistently large errors in the estimate in both cases. However, note the prevention of transient errors due to the constraints and the growing yaw error during the progression of the dataset.
Due to the observability of the rotational components (Figure 25), no large-scale walk or bias is observed as in the translational components, but some drift is still observable in the yaw component. Note, however, the prevention of transient errors (in estimated roll) by the introduction of constraints that prevent convergence to widely inaccurate estimates due to incorrect convergence of the optimizer. Additionally, note the tendency for the constrained estimator to force values towards the initial calibration value. While this may prevent the recovery of the true value of a particular parameter, it could also avoid bias induced by the coupling of certain transitional/rotational components. On small scales, these different motions are difficult to distinguish and would require ground-truthing of the stereo transform movement to evaluate the effect. It must also be noted that allowing freedom in the recovery of the translational components that the bundle adjustment optimization step exhibits better convergence towards a reduced least-squares error due to the additional flexibility this allows.
6.1. Timing results
Results for the execution time of the algorithm are presented in Table 2. The algorithm was executed in a Windows 7 environment on a desktop PC with 16 GB of RAM and a four-core Intel Core i5 650 CPU at 3.20 GHz. Average processing speed per pair of stereo frames (feature detection, matching, pose extraction, structure triangulation and bundle adjustment) is approximately 1.350 Hz on the 1280×960 resolution images of the presented dataset. While some updates are fast given small numbers of matched features and fast breakout of the bundle adjustment step (i.e no optimization performed), the average time per frame is 741 ms. While online execution time was not the intended focus of this paper, the algorithm is suited to pose estimation in an online fashion given several improvements. With future work on extended threading and leverage of a GPU for feature detection, feature matching and bundle adjustment, and with increased processor speed, the algorithm should approach online operation. Additionally, execution on lower resolution imagery should improve processing speed dramatically.
Timing results from constrained stereo VO.
Of special note, the bundle adjustment method convergence speed is dependent on the proximity of parameters to the barrier, which can cause some slowdowns in execution time. In future, an alternative method to the log-barrier inequality constraint, such as gradient-projection (Nocedal and Wright, 2000), will be implemented to ensure optimal results and faster execution speed even with parameters approaching their inequality constraints.
7. Conclusion
We have presented a robust long-range stereo VO algorithm for pose estimation of a fixed-wing UAV. Utilizing a modified bundle adjustment method, deformations in the stereo transform were accounted for to ensure accurate triangulation of structure and robust pose estimation where traditional rigid-stereo VO methods fail. A novel pose initializer was developed and presented that allows accurate scale estimation at long range without the requirement that structure is triangulated from potentially poorly calibrated stereo pairs. This initialization ensures robustness where a traditional rigid-stereo initialization would fail.
The presented algorithms were applied to a field-gathered dataset of stereo data from a fixed-wing UAV subject to vibration-induced deformation. The algorithm successfully demonstrated the novel initialization and optimization steps where a traditional stereo method was guaranteed to fail, validating the technique.
The presented VO algorithm is suited to online operation provided that sufficient compute power is available, with loop-closure optimization possible in a ‘best-effort’ output depending on the frequency of loop-closure events and additionally threading capability. The current VO implementation performs offline at around 1-2 Hz on a desktop PC, partly due to the typically high cost of the bundle adjustment stage, accounting for 45% of the computation time. This large fraction can be attributed to the log-barrier cost function for the variable constraints. If a variable approaches a barrier during the optimization, the step function will be forced to take small optimization steps, causing inefficiencies in computation. Additionally, the bundle adjustment stage is performed entirely on the CPU and is not heavily optimized to take advantage of additional sparsity other than the presented theory. By performing this stage on a GPU and implementing localized speedups, the technique could approach online operation on commodity hardware.
The focus of our work has been in application of theory into a functional system, as the precursor to online operation and the integration with additional sensors for closed-loop control. This implementation will form the next component of our work, which will also include the implementation of more computationally efficient constraint-based optimization methods and bundle adjustment speedups.
Footnotes
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
