Which two characteristics must a researcher consider concerning data quality when ensuring that an analysis is based on a clean data set?
Choose 2 answers.
Correct Answer: B,D
When evaluating whether a data set is clean enough for analysis, a researcher must focus on data quality dimensions that directly affect validity and usefulness. Two important characteristics are uniqueness and relevance. Data elements must be unique to prevent duplicate records from distorting counts, averages, totals, and trend analyses. Duplicate entries can lead to biased results, especially in customer, transaction, or survey data. Relevance is equally important because even accurate data are not helpful if they do not pertain to the question being studied. A clean data set should support the actual purpose of the analysis rather than merely being complete or large. The statement about age is incorrect because timeliness often matters; outdated data may no longer reflect the current environment. The statement that data cannot contain outliers is also too absolute. Outliers may be valid observations and can sometimes reveal important conditions, anomalies, or data-entry problems that require investigation rather than automatic removal. Thus, the best two characteristics are uniqueness and relevance, because both directly support meaningful, accurate, and decision- ready analysis.