Testing procedures¶
Errors occur time and again in data processing (→ DataOps) and analysis, whether in simple queries or in complex data pipelines for machine learning (→ MLOps).
In addition to data quality issues, there are other problems that can arise during the various phases of a data science project.
Process errors¶
The most common causes of errors include:
poorly documented, inadequately understood and inconsistently recorded input data with misleading or obscure field names and values that are difficult to interpret
inadequately specified analysis objectives
a lack of validation of input and output data
inadequate verification of the data analysis solution
inappropriate methods
a lack of automated testing of the final analysis pipeline
poorly documented results without a clear understanding of their limitations and conditions of validity
inadequate monitoring of the data pipeline
Our aim is to reduce the frequency and severity of such errors.
Test-driven data analysis¶
In this context, two types of test-driven data analysis can be distinguished:
- Range tests
Sensors are usually calibrated for a clearly defined measurement range only. If these sensors then produce results that fall outside this range, they can no longer be trusted.
- Regression testing
Errors in the data are automatically detected when they deviate from a reference, for example, if the measurement time lies in the future.
Other areas are less readily addressed by software, such as formalisation and implementation errors. Nevertheless, we will discuss some concepts and explain approaches to prevent, or at least reduce, their occurrence.
Phases of a typical data science project¶
Phase |
Error category |
Explanation |
|---|---|---|
|
Formalisation errors |
Data, subject domain or methods were not understood |
|
Implementation errors |
Bug |
|
Application errors |
No data or incorrect data during operation, for example, data that has not been fully updated |
|
Analysis errors |
Discrepancy between the data and assumptions available during development and those processed during operation |
|
Interpretation errors |
Misinterpretation of the results |
Avoiding errors in analytical processes¶
To minimise the various types of errors mentioned above, different approaches are required. The most fundamental of these is a constant awareness of how easily one can be misled in any data analysis. We must constantly bear in mind that the results produced may be nonsense, made to appear plausible only by elegant presentations or the appearance of objective infallibility.
Whilst the probability of error can be reduced across all categories, only three of the error classes can be easily rectified using software:
Implementation errors (bugs) can be rectified through regression testing
Application and analysis errors can be reduced if the data is carefully checked at all stages of the pipeline
Interpretation errors, however, can hardly be detected by software.
‘There are three kinds of lies: lies, damned lies and statistics.’ [1]
Nevertheless, some of the bad practices proposed by Darrell Huff in 1991 [2] should be avoided.
Error Types¶
The creation and maintenance of good metadata, as well as concepts relating to reproducibility, are also relevant here.