Data quality and governance
Data quality
How well data fits its purpose, usually measured as accuracy, completeness, freshness, consistency and uniqueness.
In practice
Measure it per field and per delivery, not as a general promise.
Quality depends on the decision
A dataset can be complete but stale, fresh but duplicated, or structurally valid while containing the wrong products. Set quality criteria for the intended use instead of relying on one overall score. A daily market report and an operational alert may require different freshness thresholds even when they use the same underlying fields.
Example: a price-feed acceptance check
An illustrative check might require a product identifier, numeric price, known currency and observation timestamp. Additional checks could detect duplicate keys or unexpectedly large changes. Passing these rules establishes only the conditions tested. A plausible number could still be attached to the wrong variant, so sample-based comparison with source evidence remains useful.
Make failures actionable
Report quality by source, field and failure type, not only as a blended percentage. Agree whether rejected records are held, corrected or delivered with explicit flags. Keep the validation rule version with the batch so results can be reproduced. Avoid improving a headline completeness score by filling unknown values with guesses; downstream users should be able to distinguish observed, derived and missing information.
In Import.io managed engagements, every delivery is checked against a written schema contract and the source page before it ships.
Managed Services →