When Difference Testing Needs a Quality Check
Using post-hoc diagnostics to whether reference-based discrimination data are reliable enough for decisions, knowledge building and AI-ready reuse.
In descriptive sensory methods such as QDA and Spectrum, data quality is not judged from the product scores alone. Panel performance is monitored, attribute use is reviewed, and calibration is checked over time.
For difference testing, the outcome is often treated more directly: significant difference or no significant difference.
That conclusion is important, but it does not tell the whole story. A test can produce a statistical outcome while still leaving open a more basic question: did the test itself function well enough for the result to be trusted?
Over time, a large number of reference-based discrimination tests had been conducted across different regions and panels. Each test was designed to assess whether products differed from a trained reference, using an A-not A R / R-index type approach.
Together, these results represented more than a collection of individual project outcomes. They formed a historical sensory data asset with potential value for future decisions, capability building and reusable knowledge systems.
But reuse requires more than storing the outcome of a test.
Before these historical results could be built on, the quality of the evidence had to be assessed. Were the references recognized? Were the scales used as intended? Were the outcomes reliable enough to support decisions beyond the original project?
Difference test results can look deceptively simple: different or not different. But behind that outcome, the quality of the evidence can vary substantially.
If panelists do not recognize the reference, use the response scale inconsistently, avoid uncertainty categories, or confuse the reference with a prototype, the final result may be misleading. A “difference” result may be overstated or hard to interpret. A “no difference” result may reflect weak test sensitivity rather than genuine product similarity.
This becomes especially important when historical data are used beyond the original project: in decision-making, capability building, knowledge systems or AI-enabled interpretation.
The analysis added a post-hoc quality layer to reference-based discrimination testing.
Several diagnostics were evaluated together. Scale usage showed whether panelists used the full response structure or relied too heavily on extremes. Reference recognition measures, including recognition d′, assessed whether the blind reference was actually recognized during the test. Test outcomes were then interpreted in light of these diagnostics.
By combining these measures, each study could be classified according to evidence quality. Some results were reliable and decision-ready. Others showed significant differences but weak reference recognition. Some were unreliable because the reference was not identified, and some pointed to possible test errors or method-fit problems.
The analysis made clear which datasets could be used with confidence, which required caution, and where the quality problems were located.
In some regions, the method produced reliable results and useful evidence. In others, the diagnostics pointed to issues with reference training, scale instruction, product variability, panel behaviour, or the suitability of the method itself.
This turned a collection of historical test results into a structured view of evidence quality.