Data quality and governance
Data governance
The policies, roles and controls that decide how data is collected, stored, used and retired.
In practice
For web data it covers sources, purposes, personal data and retention.
Decisions and ownership
Governance defines who approves sources and uses, who owns a dataset, who can access it and who handles changes or incidents. A practical policy connects these decisions to the actual collection and delivery workflow. It should make responsibilities understandable to engineers and data consumers rather than existing only as a document separate from operations.
Example: introducing a new source
Before adding a source to a recurring feed, record its intended use, required fields, access approval, retention needs and responsible owner. Check which downstream systems will receive the records. If the schema later gains a new free-text field, review that change rather than assuming the original approval covers every possible value the field could contain.
Evidence that controls are working
Maintain a source inventory, schema versions, access records and a way to trace delivered data back to its collection context. Define how corrections and deletion decisions reach downstream copies. Review policies as sources and uses change. Governance supports consistent decisions and accountability, but the existence of a policy or log alone does not demonstrate that every collection or use is permitted.
In Import.io managed engagements, every delivery is checked against a written schema contract and the source page before it ships.
Managed Services →