When Difference Testing Needs a Quality Check

Using post-hoc diagnostics to whether reference-based discrimination data are reliable enough for decisions, knowledge building and AI-ready reuse.

In descriptive sensory methods such as QDA and Spectrum, data quality is not judged from the product scores alone. Panel performance is monitored, attribute use is reviewed, and calibration is checked over time.
For difference testing, the outcome is often treated more directly: significant difference or no significant difference.
That conclusion is important, but it does not tell the whole story. A test can produce a statistical outcome while still leaving open a more basic question: did the test itself function well enough for the result to be trusted?

Context

reference_mapping
Over time, a large number of reference-based discrimination tests had been conducted across different regions and panels. Each test was designed to assess whether products differed from a trained reference, using an A-not A R / R-index type approach.
Together, these results represented more than a collection of individual project outcomes. They formed a historical sensory data asset with potential value for future decisions, capability building and reusable knowledge systems.
But reuse requires more than storing the outcome of a test.
Before these historical results could be built on, the quality of the evidence had to be assessed. Were the references recognized? Were the scales used as intended? Were the outcomes reliable enough to support decisions beyond the original project?

The Challenge

Difference test results can look deceptively simple: different or not different. But behind that outcome, the quality of the evidence can vary substantially.
If panelists do not recognize the reference, use the response scale inconsistently, avoid uncertainty categories, or confuse the reference with a prototype, the final result may be misleading. A “difference” result may be overstated or hard to interpret. A “no difference” result may reflect weak test sensitivity rather than genuine product similarity.
This becomes especially important when historical data are used beyond the original project: in decision-making, capability building, knowledge systems or AI-enabled interpretation.
Tasting_icecream

The Approach

funnel of data
The analysis added a post-hoc quality layer to reference-based discrimination testing.
Several diagnostics were evaluated together. Scale usage showed whether panelists used the full response structure or relied too heavily on extremes. Reference recognition measures, including recognition d′, assessed whether the blind reference was actually recognized during the test. Test outcomes were then interpreted in light of these diagnostics.
By combining these measures, each study could be classified according to evidence quality. Some results were reliable and decision-ready. Others showed significant differences but weak reference recognition. Some were unreliable because the reference was not identified, and some pointed to possible test errors or method-fit problems.

Results

The analysis made clear which datasets could be used with confidence, which required caution, and where the quality problems were located.
In some regions, the method produced reliable results and useful evidence. In others, the diagnostics pointed to issues with reference training, scale instruction, product variability, panel behaviour, or the suitability of the method itself.
This turned a collection of historical test results into a structured view of evidence quality.
Ice_cream_scoop

Impact

The work showed that difference test data should not be treated as equally reliable by default.
Reference-based discrimination tests can do more than estimate whether products differ. When combined with post-hoc diagnostics such as recognition d′, they can also help assess whether the evidence is strong enough to support a decision.
For organizations building sensory knowledge systems or preparing data for AI-supported interpretation, this step is essential. The value is not simply in having more data. The value is in knowing which data can be trusted.

Scientific foundation

Lee, H.L., van Hout, D., & Lee H.S. (2022). Improving the performance of A-Not AR discrimination test using a sensory panel: Effects of the test protocols on sensory data quality. Food Quality and Preference. 104, https://doi.org/10.1016/j.foodqual.2022.104740
Lee, H. S., Kim, M. A. & van Hout, D. (2024). Sensory Quality Measurement Based on SDT Discrimination. In Discrimination Testing in Sensory Evaluation. Wiley. https://doi.org/10.1002/9781118635353.ch9  
Interested in applying this type of thinking to product change, reformulation, or long-term innovation decisions?