Most statistical tests require
complete data without missing values. In the WebIDQ workflow, data cleaning is
therefore followed by imputation, which substitutes missing values with
estimates that are as realistic as possible.
However, too many missing
and therefore imputed values make the statistical analysis unreliable. It is not desired to make statistical conclusions based on mainly imputed data. The data
cleaning step ensures that statistical comparisons are performed for metabolites,
where enough valid concentrations are available. While 80% is a widely used
threshold for data cleaning, this value is not strictly defined. Also lower
thresholds, e.g. 70%, or higher thresholds, e.g. 90%, can be used, depending on
the study design and sample size,