← Back to glossary definition

Web scraping and extraction

Data extraction

Also called: Web data extraction

The step that turns raw web content into structured records with defined fields and types, ready to load, query or analyse.

In practice

Extraction quality is measured field by field: a price in the wrong column is worse than a missing one.

From page content to records

Extraction selects the values that matter from source content and assigns them to an agreed schema. A product record might require an identifier, price, currency and availability state. The process needs to preserve the relationship between these fields rather than collecting unrelated text fragments that happen to look like prices or product names.

Example: interpreting a product page

A page can contain a current price, a previous price and prices for recommended products. An extractor must associate the correct current offer with the target product and variant. It may also need to normalise a local number format while retaining the currency and source timestamp. Producing a valid JSON object does not establish that those associations are correct.

Check meaning as well as structure

Validate field types and required values, then review representative records against their source pages. Track missing values explicitly and distinguish unavailable data from extraction failure. Keep source URLs and observation times for troubleshooting. When a page layout changes, test whether fields still refer to the intended elements; a run that returns rows can still be silently wrong.

at import.io

Import.io extractors are built by point-and-click, AI assistance or engineers, with rendering, pagination and schema detection built in.

Web data extraction →

Related terms

Need the data, not just the definition?