Accurate stereoacuity assessment is essential for evaluating binocular vision, yet non-visual factors such as instruction quality and task execution time might influence performance on random-dot stereotest in adults with normal binocular vision.
MethodsIn this randomized controlled study, 100 adults (mean age 26.4 ± 5.17 years) with normal binocular vision were assigned to a control group (n = 40) or an experimental group (n = 60). All participants completed two consecutive administrations of the Random Dot 2 Stereo Acuity Test with LEA Symbols in an immediate test–retest paradigm. In the control group, both administrations followed standard manufacturer instructions. In the experimental group, the second (retest) administration was preceded by a detailed explanation supported by visual examples of possible target shapes. Stereoacuity (in arcseconds) and test completion time (in seconds) were recorded to evaluate the effect of enhanced instruction on retest performance.
ResultsThe control group showed no significant change in stereoacuity between measurements (mean change: 0.03 log10 arcsec, p = 0.125). In contrast, the experimental group demonstrated a significant improvement (mean change: 0.45 log10 arcsec, p < 0.0001), with most participants reaching the test’s finest disparity level (63 arcsec). Changes in response time did not differ between groups (p = 0.063). However, in the experimental group, greater improvements in stereoacuity were strongly associated with longer post-test response times (p < 0.0001).
ConclusionEnhanced instructions with visual cues significantly improved stereoacuity outcomes beyond expected test–retest variability, highlighting the importance of instruction quality and task engagement in stereopsis assessment, which should be carefully considered in both clinical practice and research settings.
Stereoacuity refers to the ability to perceive depth based on binocular disparity and represents one of the most refined aspects of human visual function. It reflects the smallest detectable difference in depth that an individual can perceive and is typically expressed in seconds of arc.1 Accurate assessment of stereoacuity is crucial in both pediatric,2 and adult populations,3 as it provides valuable information about the integrity of binocular vision. Stereoacuity plays a central role in the evaluation and management of binocular vision dysfunctions,4 amblyopia,5,6 and strabismus.7 In addition, stereoacuity is also associated with other vision-related aspects such as performance on motor skills tasks,8,9 sport performance,10 and quality of life.11
Stereoacuity can be assessed at both near and distance, although near testing remains the most common approach in clinical practice.12 Clinicians rely on a variety of commercially available stereotests, each with distinct characteristics that may complicate direct comparisons between results. Some tests, such as the Frisby Stereotest (Frisby Stereotest, Sheffield, United Kingdom) and the Lang Stereotest (Lang-stereotest AG, Küsnacht, Switzerland), are particularly suitable for pediatric populations because they promote engagement and do not require the use of polarized or red–green glasses. In contrast, the TNO stereotest (Laméris Ootech BV, Nieuwegein, Netherland) uses red–green anaglyphic stimuli and therefore requires red–green glasses, while tests such as the Randot Stereotest (Stereo Optical Company, INC, Chicago) and the Titmus Fly Test (Stereo Optical Company, Chicago, USA) are based on polarized stimuli.
Stereoacuity tests can also be classified according to the method used to generate binocular disparity: contour-based tests assess local stereopsis, whereas random-dot tests assess global stereopsis. In individuals with normal binocular vision, no significant differences in stereoacuity are typically found between these two types of tests.12,13 However, in individuals with reduced stereopsis, contour-based methods tend to overestimate stereoacuity compared with random-dot (global) tests.13,14 This discrepancy is primarily attributed to the presence of monocular cues in contour-based tests; and to global stereopsis requiring precise binocular fusion and involving, not only disparity-selective neurons in primary and secondary visual cortex (V1 and V2), but also neurons in higher-order visual areas that integrate local stereoscopic information over larger spatial regions.15
The interobserver and test–retest reliability of stereoacuity tests has been widely investigated. It was found to be high for the Randot Preschool Stereoacuity test (Stereo Optical Co, Chicago, USA) in a population of children with diverse binocular sensory function.16 Mean differences between the test and retest scores were small, 0.021 log seconds of arc, and not significantly different from zero. More recently, Metha and O’Connor (2023), assessed test-retest variability across several stereoacuity tests (TNO, Frisby, Lang Stereopad and Asteroid), and concluded that variability was clinically insignificant for all tests except the Lang Stereopad.17
On the other hand, performance in psychophysical tasks is known to improve with repeated exposure due to perceptual learning mechanisms. In the context of stereopsis, studies using repetitive stereovision training have demonstrated that observers with normal binocular vision can identify the absence or presence of depth faster and more accurately,18 and can significantly improve stereoauity.19 Such improvements may reflect enhanced task understanding, reduced uncertainty, and more efficient search strategies, rather than changes in underlying sensory function.20,21 These findings indicate that stereoacuity measurements obtained with random-dot tests are influenced not only by binocular disparity processing, but also by perceptual, cognitive and attentional factors in contrast to local stereopsis simpler patterns. Stereoacuity thresholds are not fixed; they depend on, among others, familiarity with the task, attention and response criteria.
Despite this evidence, real-world clinical environments may introduce additional sources of variability that are not accounted for in controlled reliability studies. High patient volumes and time constraints in clinical settings can lead to faster or less detailed test administration, potentially affecting patient performance. These non-visual factors may be particularly relevant for tests that require cognitive engagement, attentional focus, or mental interpretation of abstract stimuli, such as random-dot stereograms.22 Importantly, variability arising from these factors may influence not only test outcomes but also subsequent clinical interpretation and decision-making.
Instruction clarity is another variable that can influence test outcomes. Random dot patterns can be inherently difficult to interpret for patients unfamiliar with imagery. If instructions are vague, rushed, or overly technical, patients may not fully understand the task, leading to inaccurate or underestimated results, and potentially generating false positives when measuring high levels of stereoacuity. This is particularly relevant in populations with limited health literacy, cognitive impairments, language barriers, or young children who may not fully comprehend abstract visual concepts without clear guidance. However, the effect of instruction quality and use of visual cues on stereoacuity measurements has not been systematically investigated.
Therefore, the present study aimed to evaluate the impact of instruction quality and time devoted to task completion on global stereoacuity measurements obtained using the Random Dot 2 Stereo Acuity Test with LEA Symbols (Vision Assessment Corporation, Elk Grove Village, USA), a test that assesses both local (countour-based) and global (random-dot) stereopsis.
Material and methodsThis study used a randomized, controlled, two-group experimental design to evaluate the effect of task explanation on performance in a random-dot stereoacuity test. Participants were recruited voluntarily at the Faculty of Optics and Optometry of Terrassa of Universitat Politècnica de Catalunya UPC (Terrassa, Spain). All participants provided written informed consent prior to participation. The study adhered to the tenets of the Declaration of Helsinki and was approved by the UPC review board (approval number: 370–00,709).
Inclusion criteria were age between 18 and 39 years; distance (6 m) and near (0.4 m) visual acuity of 0.0 logMAR or better in each eye, measured with the Bailey–Lovie chart; and absence of manifest strabismus as determined by the cover–uncover test. Individuals with known ocular pathology, binocular vision anomalies, or neurological disease were excluded.
Participants were randomly allocated in a 1:1 ratio to the experimental or control group using simple randomization. No stratification by age, sex, or baseline stereoacuity was performed. All participants first completed the lateral circles section of the Random Dot 2 Stereo Acuity Test with LEA Symbols, a section of the test that assesses local stereopsis from 400 arcsec to 12.5 arcsec in 12 different levels. Only participants with stereoacuity <70 arcsec in the local stereopsis section were included in the study, as values ≤ 70 arcsec are generally considered to fall within the normal range.2 This criteria may appear restrictive considering that the random-dot stereotest used in this study measures global stereopsis at four discrete disparity levels (500, 250, 125, and 63 arcsec). Although earlier studies have shown distinct neural mechanisms for local and global stereopsis,23 more recent research suggests a single underlying mechanism.24 Differences in threshold values may be explained by factors such as monocular cues.25 Consequently, normal performance in local stereopsis does not necessarily predict optimal performance in global stereopsis tasks, and higher threshold values (e.g., 125–500 arcsec) can be observed in individuals with otherwise normal binocular vision.
After the local stereopsis evaluation, participants were allocated to either the control or experimental group, with two trained optometrists assigned, one for each group. Each of these two optometrists had no knowledge of the existence of the other study group. Testing was conducted in two separate rooms under identical lighting conditions for both groups. In both groups the random dot pattern of the test was administered twice, one after the other with a few seconds in between. All measurements were conducted at 0.4 m, with the subjects wearing their habitual refractive correction, and under identical environmental conditions to minimize measurement variability. No practice trials were allowed, and no feedback regarding performance was provided at any time.
Before the measurement, each participant was familiarized with the four shapes presented on the front cover (triangle, circle, square, and plus). Subsequently, the instructions given to the control group and during the first measurement of the experimental group followed those provided in the manufacturer’s manual: “This test consists of four sections (A, B, C, and D), each containing four subsections. In each window, a figure is presented; these figures may be a square, a circle, a house, a heart, or no figure at all.” Fig. 1A illustrates the random-dot stereogram section observed by participants in this test. In the control group, both the initial test and the re-test were administered strictly, without any additional explanation. In the experimental group, the first administration of the test was also undertaken following the instructions provided in the test manual. However, the second administration was preceded by a more detailed explanation of the task and procedure. Subsequently, during the second measurement of the experimental group, the explanation was repeated while printed images of the figures, including the visual cues present in the test, were shown as examples (Fig. 1B). The visual examples provided during instruction were intended to illustrate possible percepts rather than to indicate the spatial location of the target within the test windows. The overall duration of the examiners extended instructions and the participants observation of the examples was one minute.
(A): Random dot-stereograms of the Random Dot 2 Stereo Acuity Test with LEA Symbols with its four subsections: A (500 arcsec), B (250 arcsec), C (125 arcsec), and D (63 arcsec). Each test section includes a single blank shape (B): Examples of the figures shown to the experimental participants as additional visual cues and explanation, printed in DIN A4, subtending a visual angle of ≈ 23°.
The test was terminated when participants provided two consecutive incorrect responses, in accordance with standard test administration guidelines. Test completion time was recorded for each administration using a digital stopwatch with a resolution of 0.01 s. For all measurements, the primary outcomes were the level of stereoacuity achieved, in arcseconds, and test completion time, in seconds. Stereoacuity levels were recorded according to the predefined disparity steps of the test.
Statistical analysisData were analyzed using the Kolmogorov–Smirnov test to assess normality of distribution. Stereoacuity values were transformed to log10 units. The report on this transformed data will use the medians of discrete steps, not continuous conversions. The number of actual octaves was calculated by dividing the pre–post stereoacuity change, derived from the 95% confidence interval of all measurements, by 0.3 log10, as this corresponds to a doubling (one octave) of the stereoacuity threshold. In order to ensure that observed changes exceeded normal measurement variability and therefore reflected true change rather than test–retest noise, test–retest variability was quantified using 95% limits of agreement in stereoacuity thresholds from stable subjects, and converted to octaves.
This approach follows the methodology described by Adams et al. (2008), where changes less than approximately two octaves were considered indistinguishable from test-retest variability based on serial stereoacuity measurements transformed to log units.26 Between-group comparisons were performed using the nonparametric Mann–Whitney U test, while within-group (pre–post) comparisons were conducted using the Wilcoxon signed-rank test. A repeated-measures analysis of variance (ANOVA) was performed to evaluate pre–post changes between groups. Statistical analyses were carried out using SPSS software (version 27.0 for Windows; IBM Corp., Armonk, NY, USA).
The sample size was calculated to detect a minimum clinically meaningful change of 1 octave (0.3 log10 arcsec) in stereoacuity. Based on previous studies reporting test–retest limits of agreement of 0.46 log10 arcsec (corresponding to an SD of approximately 0.23),27 the estimated effect size was Cohen’s d = 1.30. Assuming a two-tailed α of 0.05, a statistical power of 95.78%, and an allocation ratio of 1:1, the required total sample size was 34 participants. The calculation was performed using a two-sample independent t-test in G*Power28 (Heinrich-Heine-Universität Düsseldorf).
ResultsParticipantsA total of 100 participants were included in the study, with 40 in the control group and 60 in the experimental group. The mean age ± standard deviation was 26.4 ± 5.17 years (range: 17–39 years), including 56 women and 44 men. No statistically significant differences were observed between groups in age (Mann–Whitney test, p = 0.091) or sex distribution (chi-square test, χ²=2.41, p = 0.120). Lateral stereacuity in the control group was, mean and standard deviation (±), 17.18 ± 3.80 arcsec (1.22 ± 0.10 log10 arcsec), and 17.04 ± 3.66 arcsec (1.22 ± 0.09 log10 arcsec) in the experimental group, without significant statistical differences, Mann-Whitney test, p = 0.875. No statistically significant differences were found in baseline measures between groups (Mann–Whitney test, p = 0.254).
Effect of additional explanation in stereoacuityTable 1 presents the descriptive results for response time and stereoacuity measurements in both the control and experimental groups. In the control group, the mean change in stereoacuity was 0.03 log10 arcsec (95% CI, 0.00–0.05). In the experimental group, the mean change was 0.45 log10 arsec (95% CI, 0.37–0.53). Specifically, in the control group, 36 participants showed no change in octaves, whereas 4 participants improved by 1 octave. In the experimental group, 15 participants showed no change, 15 improved by 1 octave, 18 improved by 2 octaves, and 12 improved by 3 octaves.
A two-factor repeated-measures ANOVA with time (pre–post), and group (control vs. experimental), revealed a significant interaction between groups, F(1,98) = 34.34, p < 0.001, partial η² = 0.259. Specifically, in the experimental group, 11 participants showed a change of 0.90 octaves. Stereoacuity outcomes for both groups are shown in Fig. 2.
Stereoacuity in the control group showed no statistically significant differences between first, 2.16 ± 0.37 log10 (arcsec), and second, 2.13 ± 0.36 log10 (arcsec), measures (Wilcoxon test, p = 0.125). In contrast, the experimental group showed a statistically significant improvement from 2.23 ± 0.31 log (arcsec) initially to 1.79 ± 0.00 log10 (arcsec) at the second measurement, reflecting that all participants achieved the same stereoacuity level (63 arcsec) in the second measurment, mean difference of 0.44 log10 (arcsec), Wilcoxon test, p < 0.0001.
The analysis of response time revealed no correlation with age at the first measurement across the entire sample, Spearman’s rho(100) =0.02, p = 0.833. The between-group comparison of the pre–post response time difference showed no statistically significant differences (Mann–Whitney test, p = 0.063), indicating that the extended and detailed explanation did not affect task execution time.
The correlation between changes in stereoacuity (pre–post) and changes in response time (pre–post, in seconds) was not statistically significant in the control group, Spearman’s rho(40)=−0.16, p = 0.300. However, a strong and significant correlation was observed in the experimental group, rho(60)=0.67,p < 0.0001, indicating that participants who required more time to complete the second measurement achieved greater improvements in stereoacuity.
To further confirm this finding, participants were divided into those who were “slower” and “faster” based on the pre–post change in response time in seconds. This slow–fast comparison was conducted as an exploratory, post hoc analysis to facilitate descriptive interpretation of the data and was not intended as a predefined categorical threshold or as a substitute for analyses using response time as a continuous variable. Negative pre–post values indicated an increase in response time, and therefore classified participants as slower, whereas positive values reflected a decrease in response time and classified participants as faster. A total of 29 participants were classified as “slower” and 71 as “faster.” A significant difference in stereoacuity change was observed between these groups (Mann–Whitney test, p < 0.0001), indicating an association between increased response time during the second measurement and greater improvements in stereoacuity. A strong correlation between the change in stereoacuity measure with time was found, Spearman’s rho (100)= −0.67, p < 0.0001 (Fig. 3).
DiscussionThe present study demonstrates that providing a more detailed explanation of the task, specifically one that allows participants to visualize the possible outcomes of a random-dot stereogram, significantly improves global stereopsis measurements. Importantly, this improvement does not necessarily mean a change in underlying stereoscopic function or an increase in measurement accuracy. Rather, it is likely driven by other factors such as improved task understanding, response strategy, or perceptual learning. In this context, the visual examples provided in the experimental group may have acted as a brief form of perceptual training, reducing uncertainty about the expected percept and facilitating a more directed search strategy. Such effects are known to enhance performance in tasks requiring the detection and localization of target shapes within visual backgrounds where distractors are present.29
In the experimental group, a statistically significant improvement in stereoacuity was observed in the re-test, when an extended explanation of the task was given (p < 0.0001). In contrast, no statistically significant differences in stereoacuity were found between the two measurements in the control group, where no additional information was provided (p = 0.125). The absence of improvement in the control group suggests that familiarity alone cannot explain the magnitude of change observed in the experimental group.
The magnitude of the observed change further supports this interpretation. The experimental group exhibited a mean change of 0.45 log10 arcsec (95% CI: 0.37–0.53). Substantially greater than the test-retest varibility reported by Fawcett et al. (2000)16 and O’Connor et al. (2010) ,17 with mean differences between measurements of 0.021 and 0.06 log seconds of arc, respectively. In contrast, the control group in the present study showed a mean change of only 0.03 log10 arcsec, consistent with measurement noise rather than true change. Together, these findings demonstrate that differences in how clinicians explain and administer stereotests can affect outcomes to a greater extent than the inherent variability of the instruments themselves.
Previous studies have examined test–retest variability of stereotests in populations with varying binocular status. Adler et al. (2012) evaluated the repeatability of the Randot test in children aged 4–12 years and found that variability increased as stereoacuity worsened, with the poorest repeatability observed in children with intermediate or coarse stereopsis.30 Similarly, Adams et al. (2008),26 investigated the variability of stereoacuity measurements in a sample of 36 subjects with constant strabismus. Stereoacuity was evaluated both at near, with the Preschool Randot and the near Frisby stereotests, and at distance, using the Frisby Davis Distance and the Distance Randot stereotests. In this population variability was particularly high, with 95% limits of agreement between 0.24 and 0.68 log arcsec depending on the test. The authors concluded that, for this population, a change of at least two octaves is required to indicate a true change in stereoscopic ability, whereas smaller variations are attributable to test-related variability. These findings are consistent with the broader literature indicating that poorer baseline stereoacuity is associated with greater measurement variability. Importantly, our results show that even in adults with normal binocular vision, where variability is expected to be minimal, instruction quality can produce changes approaching clinically meaningful thresholds.
An additional finding of this study was the relationship between response time and stereoacuity improvement. Participants who showed longer response times during the second measurement also tended to exhibit greater improvements in stereoacuity. Because these analyses were exploratory and correlational, this finding should not be interpreted causally. One possible explanation is that some participants engaged more carefully with the task during the second assessment, potentially reflecting increased attentional engagement or more deliberate evaluation of the test. However, alternative explanations cannot be excluded, and further studies specifically designed to evaluate the role of response behavior are needed.
Although this effect was observed in adults with normal binocular vision, it may be even more relevant in children or in patients with binocular vision anomalies, who are typically more susceptible to uncertainty and response variability.
It is also important to consider the type of stereotest used. The present study employed a global stereoacuity test based on random-dot stereograms, which minimizes monocular cues and provides a more accurate measure of true stereoscopic function. In contrast, contour-based (local) stereotests, such as the Titmus Fly test or the Wirt circles of the Randot, can be partially solved using monocular cues, leading to potential overestimation of stereoacuity, particularly in patients with amblyopia or strabismus.13,14 Chopin et al. (2019) further demonstrated that even global stereotests can be influenced by non-stereoscopic binocular cues, emphasizing the complexity of accurately measuring stereoacuity and the need for careful test administration.31 Fawcett (2005) also reported only moderate agreement between contour-based and random-dot stereotests, concluding that these methods are not interchangeable.13 These findings reinforce the need for standardized instructions and careful control of testing conditions when assessing stereoacuity.
The influence of instruction quality on test performance has been extensively documented in other areas of vision care, particularly in automated perimetry. Educational interventions such as pre-test training videos have been shown to improve reliability indices and reduce the need for repeated visual field tests.32,33 These findings support the notion that instruction quality is a determining factor of test accuracy and reliability.
Several limitations must be acknowledged. The study included only young adults with normal binocular vision; therefore, the findings cannot be generalized to children or to patients with amblyopia, strabismus, or other binocular vision anomalies. Although stereothreshold values obtained with the random-dot component of the Random Dot 2 Stereo Acuity test were relatively high, the results of the lateral circles test indicate preserved and comparable stereoscopic function in both groups. This apparent discrepancy may be explained by the differing visual demands of the two test components. The random-dot stereogram assesses global stereopsis and relies heavily on binocular correspondence in the absence of monocular cues, making it more sensitive to factors such as visual noise, attentional demands, or task difficulty. In fact, this finding aligns with the purpose of the present study, which demonstrated that providing a more detailed explanation, supported by visual cues, for the random-dot stereogram component of the studied test, results in significantly improved stereoacuity measurements.
In addition, cognitive factors such as intellectual quotient or attentional capacity were not assessed and may have influenced task comprehension and performance. Fatigue, a potential confounder, was minimized by keeping the interval between measurements brief and consistent across participants. Moreover, the use of an immediate test–retest paradigm carries an inherent risk of short-term learning or practice effects, particularly in tasks requiring perceptual interpretation and decision-making. However, this specific limitation is in great measure mitigated by the absence of significant improvement in the control group, suggesting that learning alone does not fully explain the observed effects. Finally, the stereo-acuity test used provides discrete global disparity levels at 500, 250, 125, and 63 arcsec, with a lower limit at 63 arcsec, introducing a ceiling effect that may restrict the detection of further improvements in participants reaching this threshold. Notably, all participants in the experimental group achieved the lowest measurable level (63 arcsec), suggesting that finer stereoacuity thresholds may have been present but could not be assessed by the test.
Another limitation of this study relates to the intrinsic constraints of clinical stereoacuity tests. Many commonly used instruments, such as the Titmus, TNO, or Randot tests, rely on discrete step levels and are therefore subject to a ceiling effect. As a result, the measured stereoacuity does not represent a continuous variable but rather the best threshold available within the predefined test range. Nevertheless, it should be acknowledged that the true magnitude of stereoacuity changes with an addition explanation may be underestimated due to the inherent limitations of the test. Finally, it must be acknowledged that response times differed between groups already at baseline. This suggests the presence of inherent between-group variability unrelated to the intervention. Hence, the results may also reflect differences in task engagement rather than solely changes in stereoacuity performance.
Future studies should extend this work to clinical populations and pediatric samples to determine whether instruction quality has a similar or even greater impact in these groups.
ConclusionsProviding a more detailed explanation, supported by visual cues, of the random-dot stereogram component of the Random Dot 2 Stereo Acuity Test with LEA Symbols results in significantly improved stereoacuity measurements that exceed the expected test–retest variability of the instrument. These findings indicate that the accuracy of stereoacuity assessment is influenced not only by the type of stereotest employed but also by the quality of the instructions delivered prior to testing.
FundingThis research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
The authors have no conflicts of interest to declare.
We thank Gerson-Alexander Aguilar Bocanegra and Andrea Rodríguez Villaverde for their assistance in collecting the empirical data.





