Quantifying data entry errors: insights from the HBSC 2022 Lithuania study
| Date | Start Page | End Page |
|---|---|---|
2024-05-29 | 1 | 1 |
Background: Previous research suggests that double entry is one of the most reliable methods for detecting entry errors. However, the effect of usual single-entry method on errors is rarely assessed. Our objective was to identify most common data entry errors and to assess how it relates to different characteristics of items, data entry clerks (DECs), and the data itself.
Methods: The data for our study was retrieved from the HBSC 2022 study in Lithuania. In total, 6730 questionnaires from grades 5–11 were manually entered twice into “Microsoft Excel” database by 27 different DECs. Differences between entries were identified as entry errors and were corrected through a third entry. Error types were identified and clustered into categories. For each category, average impact (average number of entries affected) was calculated. Errors per 1000 entries (error rate) were calculated for each question and its characteristic (e. g. number of response options).
Results: Total error rate was 3.9 (5.6 in the 5-7th grade questionnaire, 2.3 in the 9-11th grade questionnaire). In total, 29 types of errors were identified, with average impact ranging from 1.0 to 5.1 entries. Error rates by question varied from 0.1 to 24.6, error rates by question characteristic varied from 1.1 to 16.9. Depending on the occupation of DEC, error rate per one HBSC questionnaire varied 0.3 to 5.6.
Conclusions: Double data entry method allowed to avoid 3.9 errors per 1000 entries. Occupation of DEC and question characteristics were found to be significant variables associated with differences in error rates."