July 16th, 2026
In a Smith-Kettlewell collaboration reaching from San Francisco to Paris and Marseille, Adrien Chopin and colleagues benchmarked the methods scientists use to measure visual perception and found that two, including one they built, are faster, far more accurate, and more reliable.
Highlights
- A Smith-Kettlewell-led study compared the common methods for measuring visual perception and found two clear winners, one of them a new method the team built, psi-marg-grid.
- The standard methods are slow and grow unreliable for the people who matter most: those with low or unusual perception, such as people with amblyopia.
- In simulations, the best methods were far faster and about 11 times more accurate than the standard approach, and almost never missed people who responded randomly or fell outside the normal range, pointing toward quicker, more dependable vision tests.
The measurement problem at the heart of vision science
Amblyopia, commonly known as “lazy eye,” is the most common cause of vision loss in children, affecting an estimated 99 million people worldwide, a number projected to more than double to over 220 million by 2040. Diagnosing and studying conditions like it rests on a deceptively hard task: measuring precisely what a person can see. Measuring perception to a high scientific standard is inherently slow, and the standard methods grow unreliable for the very people whose vision is already poor.
Smith-Kettlewell scientist Adrien Chopin and his colleagues at the Institut de la Vision (Sorbonne Université) in Paris and the Institut de Neurosciences de la Timone (Aix-Marseille Université) in Marseille built a new method, psi-marg-grid, that, alongside one established method, beat every standard way of measuring perception at once: faster, more accurate, and more reliable. To prove it, they pitted the leading methods against one another across thousands of simulations, where psi-marg-grid caught the atypical and random responders that standard methods routinely missed.
The project grew out of a practical frustration. “I was trying to test stereovision in amblyopic patients and it was taking way too much time,” Chopin said. “On top of that, it was not repeatable at all when we were measuring people who had poor perception.” Stereovision, the depth sense he was trying to measure, comes from the brain combining the slightly different images from the two eyes, and it is one of the first abilities amblyopia erodes. And that is exactly where the standard tests falter: the weaker or more irregular a person’s perception, the more trials they demand and the less their results hold up on a retest, the slowness and unreliability Chopin kept running into.
Benchmarking the best way to measure visual perception
Chopin’s frustration carried a quiet challenge to an assumption. When a test keeps failing on the same kind of patient, the easy conclusion is that those patients are simply hard to measure. The team suspected the fault lay elsewhere: what if the measurement methods, not the people, were the weak link? To find out, they turned the usual experiment inside out: instead of measuring more people, they measured the methods themselves. The subjects were a set of simulated observers, computer models that answer a vision test the way real people do, complete with wandering attention and a wide range of individual quirks, but whose true underlying perception is known in advance. Because that true answer was known, the team could score exactly how close each method came to it.
That approach also let them push each method to its limits. Whereas a human volunteer tires after a handful of runs, a simulated observer can repeat a test endlessly, exposing how each method behaves not just on average but for the hardest-to-measure people, including those with low or unusual perception. The comparison rested on the qualities that matter in practice: how repeatable, precise, and accurate each method is, and how reliably it flags someone whose perception is genuinely low.
How perception tests work, from fixed menus to self-adjusting trials
Every method here hunts for the same number: a person’s threshold, the faintest difference they can reliably tell apart. What separates them is the strategy.
The oldest and still most common is the Method of Constant Stimuli. It runs from a fixed menu, a spread of difficulty levels chosen in advance and each shown many times, then reads the threshold off the tally. It is simple and even-handed, but wasteful: it spends as many trials on the too-easy and too-hard levels, which reveal little, as on the informative ones near the threshold.
The newer methods are adaptive. Rather than a fixed menu, they choose each trial on the fly, using everything the person has answered so far to place the next one wherever it will reveal the most. The Bayesian “Psi” family is the leading example, keeping a running estimate of the threshold and updating it after every response to converge in far fewer trials. Two refinements matter here: psi-marginal, which makes the method robust against the occasional “lapse,” when an alert person fumbles an easy trial, and psi-grid, which adjusts the range of possible answers to reallocate all the computation power around the current threshold estimate, in effect refusing to waste effort on answers it has already ruled out.
The new method, psi-marg-grid, is the one Chopin’s team built for this study. It fuses those two refinements, pairing psi-marginal’s tolerance for lapses with psi-grid’s efficient use of computing power. That combination is what let psi-marg-grid and psi-marginal outrun the rest, especially on short tests and on the observers standard methods handle worst.
Which method measures perception best?
Two methods came out ahead on every measure: psi-marginal, rarely in use, and the team’s new psi-marg-grid. For typical observers, they pinned down a person’s threshold about 11 times more accurately than the Method of Constant Stimuli, the most common method, more than three times as precisely, and about 2.5 times more reliably on repeat testing, while reaching that precision in roughly half the trials the next-best methods needed. The study is also the first to measure how well each method catches people who respond randomly or fall outside the normal range, the pattern that can signal a severe or absent perceptual response. There the gap was widest: the standard method missed more than half of these cases; the two leading methods missed almost none.
“The best method seems to be four times quicker than the most commonly used method. It is far more accurate and precise, and it makes zero errors at detecting people with low perception in the simulations.”— Adrien Chopin
The clearest gap showed up with the observers standard tests handle worst. Most methods assume that the easier a target is to see, the better a person does, a steadily rising curve. Some people don’t follow that rule, more often than the textbook assumes (though still the exception): their performance peaks and then falls, a “non-monotonic” pattern that can surface in atypical or impaired vision. When the older methods assumed the textbook curve and met one of these observers, they failed spectacularly, missing the true value by more than 4,000% even after 120 trials. Psi-marginal and psi-marg-grid could be set to expect that pattern, and doing so kept them accurate on these observers at almost no cost when the observer turned out to be typical after all, at least over a full-length test. It is the kind of edge case Chopin had in mind from the start: patients whose vision doesn’t follow the textbook examples, and whom the standard tests serve worst.
What faster vision tests could mean for research and the clinic
Better measurement tools rarely make news, but they raise the ceiling on everything built above them. If perception researchers adopted these methods consistently, studies could run faster and rest on firmer footing. The same gains could carry into eye care, where shorter, more dependable digital vision tests would mean less time in the chair and earlier, clearer answers for clinicians, including for people with amblyopia and other forms of low vision.
That edge is not lost on industry. The team worked with EssilorLuxottica, one of the world’s largest makers of vision tests, who made the reality of test development and adaptation abundantly clear: a shorter test really makes the difference between a test that is adopted or abandoned.
The head-to-head ran on simulated observers rather than real patients, a deliberate first step. But the approach is already proving itself in people: the team has built the new method into a working stereovision test called eRDSv6, where it performs well. The target was clear from the outset: the patients Chopin set out to help, whose vision has always been the hardest to measure, and who stand to gain the most from a test that finally measures it well.
Support vision research at Smith-Kettlewell
Research like this is made possible by supporters like you. Your gift funds independent vision science at Smith-Kettlewell, from amblyopia to low vision.
About the study
“Comparison of methods for quick estimation of psychometric thresholds” was published July 22, 2026 in Frontiers in Neuroscience (open access). DOI: 10.3389/fnins.2026.1760278.
Authors: Adrien Chopin (Smith-Kettlewell Eye Research Institute, San Francisco; Sorbonne Université, Institut de la Vision, Paris); Martin Szinte (Institut de Neurosciences de la Timone, Marseille); Angelo Arleo (EssilorLuxottica, Paris); and Denis Sheynikhovich (Sorbonne Université, Institut de la Vision, Paris).
Funding: Supported by the Chair SILVERSIGHT (ANR-18-CHIN-0002), IHU FOReSIGHT (ANR-18-IAHU-01), and LabEx LIFESENSES (ANR-10-LABX-65), from the French Agence Nationale de la Recherche. Adrien Chopin and Angelo Arleo received partial funding from EssilorLuxottica.
About Smith-Kettlewell
Smith-Kettlewell is an independent nonprofit research institute advancing vision science, accessibility, and innovation to improve understanding, independence, and quality of life. Learn more at ski.org.
Media contact: [TBD]
