Data cleaning: Why is the default threshold set at 80%?

Data cleaning: Why is the default threshold set at 80%?

Most statistical tests require complete data without missing values. In the WebIDQ workflow, data cleaning is therefore followed by imputation, which substitutes missing values with estimates that are as realistic as possible.

However, too many missing and therefore imputed values make the statistical analysis unreliable. It is not desired to make statistical conclusions based on mainly imputed data. The data cleaning step ensures that statistical comparisons are performed for metabolites, where enough valid concentrations are available. While 80% is a widely used threshold for data cleaning, this value is not strictly defined. Also lower thresholds, e.g. 70%, or higher thresholds, e.g. 90%, can be used, depending on the study design and sample size,

Info
For additional information, refer to the WebIDQ user manual > Data cleaning.