Abstract
Eye tracking is a non-invasive method to measure an individual’s cognitive state and visual attention allocation patterns, both of which provide valuable insight into their real-time needs when completing complex tasks. With the availability of low-cost systems and the wealth of available information, eye tracking is ever more common in human factors research and applications. However, access to comprehensive, approachable, and adaptable information on various eye tracking metrics is limited. Therefore, we created a publicly available, goal-driven, interactive guide called Generalized Eye Tracking Metrics (GEMs). It supports the learning of (1) metrics, (2) relevant terminology, and (3) analysis approaches, as it relates to the user’s research questions and experimental setup. Our findings from focus groups, interviews, and usability testing among novices and experts indicated that GEMs is a highly usable information source created by and for eye tracking researchers.
Introduction
Eye tracking is a powerful tool that can provide insight into a person’s cognitive processing during a task in real-time. Per the Eye-Mind and Immediacy Assumptions, there is a direct link between the eyes and the mind, such that whatever the eyes fixate on, the mind processes with no appreciable time delay (Just & Carpenter, 1980). With these two assumptions, as well as many other supporting theories and findings (e.g., Aston-Jones & Cohen, 2005; Posner et al., 1980; Treisman & Williams, 1984), it can be reasonably concluded that a person’s viewing patterns provide valuable information about their cognitive state, task-completion strategy, and/or decision making process. Similarly, understanding how pupil size could serve as a physiological measure of mental effort has long been of interest (Beatty, 1982; Kahneman, 1975). Some work suggests pupil size correlates with the firing of the locus coeruleus-norepinephrine (LC-NE) system, which is the system responsible for cognitive arousal regulation (Aston-Jones & Cohen, 2005; Robison et al., 2023). Researchers also believe individual differences in mental effort and cognitive abilities can be measured per task-evoked pupillary responses.
Given its versatility, objectivity, and unobtrusiveness, as well as the availability of low-cost equipment, eye tracking has steadily gained popularity in nearly all human factors applications (Krafka et al., 2016). Its use has led to improved understanding of cognitive workload, user interface design, within-task learning, human-automation trust, training programs, and workplace protocol (Duchowski, 2017; Sato et al., 2023; Sibley et al., 2019). However, implementation challenges still arise due to the breadth of literature, technical jargon, various metrics, and methodological differences across domains. This can make it difficult to confidently incorporate eye tracking into research and applications. Experienced eye tracking researchers can also struggle to learn and apply new metrics. There currently lacks a user-friendly, digestible, and interactive method to quickly and easily learn about eye tracking metrics, their applicability, and their analysis process.
Our guide, Generalized Eye Tracking Metrics (GEMs), addresses these fundamental needs and more by integrating the basics of eye tracking with relevant literature and practical information in a streamlined, interactive format. The goal of GEMs is to be a highly usable guide when learning about eye tracking metrics and fundamentals, while also providing pragmatic information and resources. At the time of this publication, the primary features of GEMs includes: (1) relevant, digestible information for 17 various eye tracking metrics in a uniform, traceable format; (2) interactive and user-friendly elements designed to engage users and support learning, such as terminology flash cards, reactive buttons, and hidden drop downs; (3) downloadable and adaptable analysis resources for users at any skill level to practice analyzing real eye tracking data at no cost.
Background
Eye trackers typically emit near-infrared (NIR) light in to the eyes to create a corneal reflection on the pupil. Then with its embedded high-speed cameras, it captures where the corneal reflection is on the pupil and the pupil size itself. The location difference between the center of the pupil and the corneal reflection is used to determine and record where a person is looking (Poole & Ball, 2006). Eye trackers come in several forms, including head-mounted, off-the-head/desktop-mounted, or directly integrated styles (e.g., head-mounted displays; c.f. Clay et al., 2019).
Plenty of resources for eye tracking research exist, ranging from textbooks (e.g., Duchowski, 2017; Holmqvist et al., 2011), dedicated peer-reviewed journal articles (e.g., Clay et al., 2019; Mathôt & Vilotijević, 2022), online forums or blog posts (e.g., Pető (2024)), and guides from eye tracker manufacturers (e.g., Tobii, n.d). These sources vary in their accessibility and applicability because they often exist behind a paywall (e.g., journal article), rely on technical jargon, are niche in focus, and/or require prior knowledge of a specific research field. Although crowd-sourced and/or self-published information (e.g., YouTube, Stack Overflow, GitHub, etc.) is often more comprehensible, streamlined, and searchable, it is not formally peer-reviewed and therefore its accuracy is not independently assessed. Even when formally peer-reviewed, these resources may not be updated or upkept (e.g., Dar et al., 2021; van Renswoude et al., 2018). Further complicating the learning process is that many of these resources are siloed, meaning terminology on basic concepts can vary (e.g., the concept of mental workload vs. cognitive load (Kosch et al., 2013) or the various definitions of a fixation). As a result, researchers, especially novices, may feel overwhelmed or even incapable of incorporating eye tracking in to their investigations, let alone drawing informative conclusions thereafter. The vast pool of available informational material does not necessarily translate to better and faster learning. Large Language Models (LLMs) have made reviewing large amounts of information more efficient, but there is currently no simple way to fact check their outputs nor the process in which those outputs were decided (Steyvers et al., 2025).
Regardless of method, learning about eye tracking metrics is time consuming because it requires sifting through vast amounts of information, understanding assumptions, and then deciding what metrics are most applicable for the given end goal. The motivation of GEMs was to publish information about eye tracking metrics in a digestible and applicable format, so that researchers may confidently capitalize on the benefits of this technology.
Approach
Application Description
At present, GEMs is a publicly accessible R Shiny application (Chang et al., 2024) and, at the date of publication, hosted online and updated at least semi-annually at https://0czolw-shannon-mcgarry.shinyapps.io/GEMsSPDMEdits/. The landing page indicates the current version of GEMs, its overarching motivation, and describes some of the included functionalities. It introduces eye tracking as a whole and gives information on the two assumptions that form its basis. The “Metric Finder” asks users to answer questions about their goals, hardware, and experimental setup, and provides relevant eye tracking metric information accordingly. GEMs also includes pages titled “DIY Data Analysis,” “FAQs and Terminology,” and “Metrics Table”.
Metric Finder
On the left-hand side of the landing page is the hallmark feature of GEMs. Step 1 of the Metric Finder asks the user to select questions that are relevant to their eye tracking research goals. Questions were iteratively generated and centered on what a GEMs user may want to know from eye tracking data (e.g., “How distributed is attention across the interface?,” “How much mental effort is the operator exerting?”). To avoid biasing the user to select metrics s/he is already familiar with, questions about specific metrics were not asked. Step 2 asks whether the user will be defining areas of interest (AOIs) and Step 3 requests information about the device’s sampling rate to tailor suggested metrics to technical constraints. Step 4 asks whether the user’s eye tracking system employs a fiducial marker and provides an opportunity to learn more about it (c.f. Coyne et al., 2019). Step 5 asks if the user would like to learn about experimental setup considerations to (1) inform their setup based on the given analysis goals or (2) identify limitations that may be associated with their setup (e.g., differing ambient light levels can influence pupil metrics). Figure 1 shows Step 1–5 of the Metric Finder.

Screenshot of the Metric Finder on the GEMs landing page.
Once Steps 1–5 are completed, the user clicks the “Submit” button to generate metrics and other relevant information based on their inputs. For example, a user who (1) checked the box “How distributed is attention across the interface?,” (2) indicated AOIs will be defined, (3) selected 60 Hz as the eye tracker’s sampling rate, (4) selected “Not sure” when asked about a fiducial marker, and (5) did not want to learn more about experimental setup considerations would be introduced to the following eye tracking metrics: Stationary Gaze Entropy, Fixation Spatial Density, and Coefficient K. For every metric, its definition, high-level description, and a pros and cons table are provided. Depending on the metric, other relevant information, such as formulas, visualizations, or the effect of sampling rate, is also provided.
DIY Data Analysis
On this page, users can access and download sample data and analysis resources to practice data preparation, event detection, and metric calculation. R scripts for the calculation of three commonly used metrics can be downloaded: total number of fixations, average fixation duration, and percent viewing time. A script is also provided for a less commonly used metric, Coefficient K (Krejtz et al., 2016). Many resources exist for preprocessing and event detection (c.f. Dar et al., 2021), but formatted, adaptable code to calculate specific metrics is not widely available. These analysis materials are based on what the authors use for their peer-reviewed research. Additional analysis materials will continue to be uploaded in the future.
FAQs and Terminology
The top of the Frequently Asked Questions (FAQs) and Terminology page has several FAQs and their corresponding answers. Several figures and schematics are included in the answers to foster deeper understanding of concepts without relying on jargon. Below the FAQs is a section of foundational terminology, such as saccade, fixation, etc., presented as flash cards. Users can learn terminology by clicking each flash card to reveal their definition and quiz themselves thereafter.
Metrics Table
This page displays 12 of the 17 metrics included in GEMs as hyperlinks to peer-reviewed information source. They are organized based on the aspect of visual attention they measure (i.e., the spread, directness, and duration of visual attention). This page aims to provide easy access to more technical, peer-reviewed information about each metric.
Implementation
GEMs was built with shiny (Posit, n.d.), shinyjs (Attali, 2021), and shinythemes (Chang, 2021). The shiny package enabled the application to be programed in RStudio and produced to HTML output upon execution. The shinyjs package allowed for the use of reactive buttons (see Figure 1 for examples). The visual theme and details of the app were achieved using the shinythemes package. A fluid layout was used for the interface of GEMs, which was implemented via the fluidPage() function within shinyApp(). In turn, the width of vertical panels adapted to the user’s screen size and create the “Metric Finder” as an interactive part of the landing page. Currently, GEMs is hosted on a no-cost server at shinyapps.io (Posit, n.d).
The “Metric Finder” output was developed from a meta-analysis for each eye tracking metric and tailored to a series of selected research questions. It was selectively presented to the user per the research question they selected (i.e., the indicator variable for the checkbox was switched to “on” in the source code). A pros and cons table for each metric was developed to facilitate the understanding of the metric’s feasibility, robustness, and dependencies.
The “Terminology Flash Cards” were created by leveraging the htmltools dependency included in shiny [i.e., div()]. Each flash card was a defined area that is customized with content and categorized as tags$detail, which meant it was classifying information as hidden until a user clicked a pixel within the custom defined area.
Soliciting Feedback
Scientists and engineers within the Department of the Navy were solicited for feedback after the first version of GEMs was completed (version 1.0.0, created August 1, 2023). These individuals had a range of experience with eye tracking research and applications, from novice to expert. Feedback was solicited via semi-structured interviews and focus groups, which began with general questions about eye tracking research needs and concluded by presenting GEMs. Questions such as: “What is something related to eye tracking that you have heard about, but do not have adequate knowledge of?” and “When you are trying to learn about unfamiliar psychophysiological metrics, what things do you want to know up front? (e.g., pros/cons, hesitations, exemplars, etc.)” were asked. Here, the goal was to identify eye tracking information needs more broadly, which is why these questions did not solely pertain to GEMs.
The findings from these interviews and focus groups were used to improve GEMs by including more general and pragmatic information about eye tracking (version 1.1.0). This ranged from including information on the physiological basis of eye tracking research to hyperlinking the documentation for R packages that exist for eye tracking data analyses. Additional follow-up interviews were conducted, but they were specific to the foundational basis, mechanics, and usability of GEMs. From there, additional changes were made, including reducing the number of questions in the Metric Finder, adding schematics and equations to the output, and reducing the number of clicks to complete a given task.
Through this process, it became clear that a formal usability assessment of GEMs was needed to achieve our goal of providing a highly usable, informative, and pragmatic eye tracking resource. Therefore, we conducted a formal usability assessment on GEMs (version 2.2.0).
Method
Participants
Ten prospective or current eye tracking researchers within the U.S. Department of War completed the usability assessment. These included experimental research psychologists, computer scientists, aerospace experimental psychologists, and engineers. All volunteers had some level of exposure to eye tracking research. It took approximately 30 min to complete all tasks, answer open-ended questions, complete questionnaires, and debrief.
User Testing
Participants were interviewed online via Microsoft Teams. Participants completed a series of three tasks that were immediately followed by open-ended questions about their experience. The goal of these questions was to understand how a user might accomplish a specific task within GEMs and to capture any associated positive or negative experiences. All questions were presented to participants in the same order. Afterward, participants completed an online version of the System Usability Scale (SUS) to ascertain a verified, quantitative measure of usability (Brooke, 1996; Lewis, 2018). At the end of testing, participants were informally debriefed on their experience.
Analyses
An inductive (i.e., data-driven) qualitative thematic analysis was conducted by the lead author on the open-ended survey responses to better understand the user experience of GEMs and extract design elements for areas needing improvement. The six steps for thematic analysis outlined in Braun and Clarke (2006) were followed, which included: (1) familiarization with the data, (2) generating initial codes, (3) searching for themes, (4) reviewing themes, (5) defining and naming themes, and (6) producing a report. Based on the results of the thematic analysis, participant needs were identified and translated into design specifications for future iterations of GEMs (Byrne, 2022). Then, the 10 questions from the SUS that assess usability on a 5-point Likert scale (1 = strongly disagree, 5 = strongly agree) were tabulated. A score above 68 on the SUS indicates above-average usability (Lewis, 2018).
Results
Three high-level themes were identified: (1) general information needs, (2) experimental setup needs, and (3) functional needs. The high-level themes were broken down into tangible design specifications for GEMs. Each design specification was based on qualitative results and anchored to the goal of making GEMs a comprehensive, practical, and usable guide for learning about eye tracking metrics. Table 1 describes each sub-need and how GEMs was iterated in response, leading to its current version (version 2.3.3).
Design specifications per the qualitative thematic analysis.
General information needs consisted of the foundational and applied information about eye tracking metrics and analysis that users considered crucial for inclusion. The tone and delivery of dense eye tracking-related information was also informed by feedback under this category. Feedback that fell into the experimental setup needs category consisted of basic factors that users should understand when designing and running eye tracking studies. Functional needs pertained to how features of GEMs impact the user experience and perceived utility.
System Usability Scale (SUS) Results
To assess the perceived usability of GEMs, the SUS was scored according to Brooke (1996). The results yielded a score of 84, which falls into the 90th percentile of scores (scores range from 0 to 100). Most importantly, the results of the SUS quantified the usability of GEMs to be above-average, substantiating it as a highly usable information resource for eye tracking research and applications.
Conclusion & Future Work
GEMs is an innovative resource that synthesizes information about 17 eye tracking metrics in a single location for novice and experienced eye tracking researchers and practitioners alike. It was built ab initio by leveraging the authors’ many years of experience in eye tracking research. Importantly, GEMs is publicly accessible at no cost to the user, which expands its reach and impact. Additionally, GEMs provides users with structured, concise, and interactive information about eye tracking methodology, which lowers the barrier to entry for learning about this technology and facilitates transparency and standardization across the field. User feedback and testing was instrumental in iterating and improving the design of GEMs, specifically in the areas of general needs, experimental setup needs, and functional needs. The observed SUS score provides evidence that GEMs is addressing a real information need via a user-friendly application.
Limitations
A limited number of metrics were included in GEMs, such that the output of metrics for each research question is not exhaustive. This is mainly a byproduct of the expertise and experience of the present authors. Future work will seek to include a greater number of metrics as the authors gather first-hand experience with them. In an effort to be concise and digestible, it is also possible that information from GEMs was oversimplified, such as the information outputted by the Metric Finder. In an attempt to avoid this, the authors provide multiple definitions of a given metric and suggestions on how to disentangle those meanings (reference the output from the Metric Finder for the question “How efficient is an operator scanning the interface?”). Further, concrete examples were given exclusively within the context of human-automation interaction, even though there are many contexts that benefit or stand to benefit from eye tracking research. Finally, another limitation of the present work is there is no validation of GEMs for a given use case. This was because the current focus was to validate GEMs as a usable information tool. Future work should compare how GEMs informs and equips eye tracking researchers to other methods.
Future Work
There remains considerable work to be done with GEMs, notwithstanding the gaps that still exist when synthesizing eye tracking metrics simultaneously. Future versions of GEMs may incorporate information from the developing research area of eye tracking safety. Finally, to continuously improve its usability, feedback from users could be solicited in real-time. For example, a link to the SUS or a general feedback survey could be placed on the “About” page or prompted upon exiting GEMs to reflexively assess its usability with the goal of remaining highly usable (i.e., at or above the 90th percentile on the SUS).
Footnotes
Acknowledgements
We would like to thank all of the scientists and engineers across the Department of War who so graciously volunteered their time to participate in the formal usability assessment.
ORCID iDs
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Command Decision Making program, led by Dr. Jeffery Morrison, at the Office of Naval Research. It also was supported by the Naval Research Enterprise Internship Program (NREIP) and National Science Foundation Graduate Research Fellowship Program (NSF GRFP) under Grant No. 1842490.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
