Web data extraction
Web scraping that works.Even when sites fight back.
Turn difficult websites into clean, scheduled data, with rendering, anti-bot, proxies and delivery handled.
30-day trial · 5,000 queries · No credit card required
| title | brand | price | list_price | rating | reviews | availability |
|---|
GET api.import.io/v1/extractors/8f2c…/dataschedule daily 06:00 · deliver S3 · webhook · CSV# Same job, from code. Rendering, anti-bot and proxies are handled server-side. import importio client = importio.Client(api_key="io_live_…") # 1. Create an extractor from a URL, with AI schema detection ex = client.extractors.create( url="https://www.homedepot.com/b/Tools-Power-Tools-Drills/N-5yc1vZc27w", schema="auto", paginate=True, render="auto", proxy="residential", ) # 2. Run it and stream rows for row in client.runs.stream(ex.id): print(row["title"], row["price"], row["availability"]) # 3. Schedule and deliver client.schedules.create(ex.id, every="1d", at="06:00", deliver={"s3": "s3://your-bucket/drills/", "webhook": "https://your.app/hooks/importio"})
→ 288 rows · 9 fields · 12 pages · 4.1 s
importio.extract(url="lowes.com/pl/…", schema="auto", paginate=true)
https://mcp.import.io/mcp · works with Claude, Cursor and any MCP client10,000 free successful calls · then $0.0002 per successful callThe web infrastructure behind enterprise data teams and data companiesIncluding three of the leading commerce-intelligence platforms, who run their web capture on it.
One engine.
Four ways in.
Import.io turns difficult websites into structured, scheduled data. Choose the interface; rendering, access, QA and delivery stay underneath.
How to extract data from any website with Import.io.
Four steps. Most first extractors are live in under ten minutes.
Point or describe
Paste a URL, click the fields or describe the data you need.
Refine and test
Confirm the schema, pagination and a live row preview.
Run at scale
Rendering, anti-bot, retries and regional access run for you.
Schedule and deliver
Send typed data by API, webhook, S3, SFTP or warehouse.
Everything a scraping stack needs, in one platform.
The browser, access, proxy, schedule, QA and delivery layers. Operated as one system.
Describe the data.
Point, click or prompt. Import.io detects records, types fields and proposes the schema.
9 fields inferred in 4 sOpen the hard pages.
JavaScript, anti-bot, CAPTCHA and authenticated sessions are handled inside the run.
challenge → handledAppear where the data is.
Datacenter, regional and residential egress with country, city and store context.
residential 190 localesKeep it current.
Pagination, scheduling, retries and typed change events run without a browser farm.
daily 06:00 · price_cutRelease data, not guesses.
Fill rates, distributions and samples are checked against the source before delivery.
7 checks → releasedSend it anywhere.
CSV, JSON or Parquet through API, webhook, S3, SFTP or your warehouse.
manifest + s3:// + webhookHow Import.io compares to the alternatives.
Pick an alternative to spotlight it against Import.io. The labels and first column stay in view while you explore.
| capability | DIY frameworks | Unblocker and proxy APIs | LLM scraping APIs | No-code scrapers | Import.io |
|---|---|---|---|---|---|
| No-code extractor builder | no | no | no | yes | yes, with AI schema detection |
| REST API and SDKs | you build it | yes | yes | varies | yes |
| MCP server for AI agents | no | some | yes | rare | yes, usage-based, 10K free calls |
| Anti-bot, CAPTCHA, residential proxies | bring your own | yes | partial | partial | yes, by plan tier |
| Structured rows, not markdown | you parse | raw HTML | markdown / JSON | yes | typed schema, validated |
| Pagination, chaining, scheduling | you build it | no | partial | yes | yes |
| Change detection and events | you build it | no | rare | some | yes |
| Data QA on every run | no | no | no | no | yes |
| Fully managed service with SLA | no | enterprise only | no | no | yes, fees at risk |
| Product matching and pricing intelligence | no | no | no | no | yes, Aperture |
| Production track record | yours | varies | since ~2023 | varies | since 2012, 500B+ points/month |
| Entry price | free, plus your time | usage-based | usage-based | low monthly | free trial, then $199/mo |
Categories: DIY frameworks such as Scrapy, Playwright and Puppeteer · unblocker and proxy APIs such as Bright Data, Oxylabs and Zyte · LLM scraping APIs such as Firecrawl · no-code scrapers such as Octoparse and Browse AI. Feature availability varies by vendor and plan; see detailed comparisons. Last reviewed September 2026.
Complete product data. Twice as often.
In a published eCommerce test, Import.io returned complete contracted records at roughly 2× the rate of conventional scraping.
Pricing you can read in one screen.
Billed on successful queries only. No card to start. Cancel anytime. Full detail on the pricing page.
Trial
- Full platform access
- Point-and-click extractors
- Docs and community support
- Hard stop, never charged
Standard
- Unlimited extractors and concurrent jobs
- Datacenter proxies
- Scheduling and delivery
- Email support
Professional
- Everything in Standard
- Regional proxies for harder targets
- API access and webhooks
- Screenshots and HTML extraction
- Chat, ticketing and training
Advanced
- Everything in Professional
- Premium and residential proxies
- Multi-user with team management
- Priority live support
Above 500,000 queries a month, heavily protected sources, or dedicated infrastructure: Managed Services with custom SLAs. MCP Web Scraper: 10,000 free successful calls, then $0.0002 per successful call with a spend cap you set. Which plan is right for you.
From prompt to production.
Use the same extraction engine through REST, SDKs or MCP. Rendering, anti-bot and delivery stay server-side.
Connect
Choose API, SDK or MCP.
Authenticate
Create a scoped key.
Call tools
Extract, search, crawl, monitor.
Ship
Schedule and deliver typed output.
# extract a page into typed rows curl https://api.import.io/v1/extract \ -H "Authorization: Bearer $IMPORTIO_API_KEY" \ -d '{"url":"https://www.homedepot.com/b/…/N-5yc1vZc27w", "schema":"auto","paginate":true,"render":"auto","proxy":"residential"}' # monitor and get events curl https://api.import.io/v1/monitor \ -H "Authorization: Bearer $IMPORTIO_API_KEY" \ -d '{"extractor":"8f2c…","every":"1h","fields":["price","availability"], "deliver":{"webhook":"https://your.app/hooks/importio"}}'
import { ImportIO } from "importio"; const client = new ImportIO({ apiKey: process.env.IMPORTIO_API_KEY }); const ex = await client.extractors.create({ url: "https://www.homedepot.com/b/…/N-5yc1vZc27w", schema: "auto", paginate: true, render: "auto", proxy: "residential", }); for await (const row of client.runs.stream(ex.id)) console.log(row.title, row.price); await client.schedules.create(ex.id, { every: "1d", at: "06:00", deliver: { s3: "s3://your-bucket/drills/" } });
// Add Import.io to any MCP client { "mcpServers": { "importio": { "url": "https://mcp.import.io/mcp", "headers": { "Authorization": "Bearer io_live_…" } } } } // tools: importio_extract · importio_search · importio_crawl · importio_monitor // 10,000 free successful calls, then $0.0002 per successful call
What people extract with it.
Retail is where the engine was hardened. It runs everywhere the web publishes data.
eCommerce product data
Product details, prices, promotions, availability by store, rankings, reviews and sellers from any retailer or marketplace.
pricestockreviewsCompetitor price monitoring
Matched SKU by SKU into pricing engines, daily or hourly, with history.
matchingindexDigital shelf and MAP
Content compliance, share of search, MAP violations with evidence, unauthorized sellers.
contentmapsellersFinancial and alternative data
Filings, listings, permits, registries and dockets structured into point-in-time datasets with citations.
pdfpit historyListings and market research
Property, jobs, vehicles, travel rates and ticket inventory, deduplicated to the entity with price history.
listingsdedupeAI agents and training data
Clean, structured web data for RAG, fine-tuning and agent tool use, pulled on demand through the API and MCP.
mcpapijsonScraping fetches pages. Extraction creates usable data.
Import.io handles the operating layer between them: rendering, access, schema, maintenance, QA and delivery.
- 01 · fetchOpen the real page.
Render JavaScript and handle protected access.
- 02 · shapeCreate typed records.
Apply schema, pagination and validation.
- 03 · operateKeep the data moving.
Schedule, monitor, repair and deliver.
Built for production, governed for enterprise.
Developers should want to try it. CTOs should trust it. Procurement should be able to sign it.
Compliance
GDPR and CCPA ready. Data processing agreements. Regional data handling where you need it.
PII detection and removal
Pipelines detect and remove personal and restricted data with sensitivity thresholds you set.
Audit trails
Every run, every record, every delivery logged. Trace any row to the fetch that produced it.
Responsible collection
Rate-aware crawling, robots and terms respected, and a written policy on what will and will not be fetched.
Security
Encryption in transit and at rest, scoped access, team management with roles on Advanced and enterprise.
Reliability
Retries, fallback rendering and typed failures. Monitoring dashboards per run and per source. Since 2012.
Web scraping questions, answered.
Short answers. Longer ones in the docs and on the call.
What is Import.io?
Import.io is a web scraping and data extraction platform founded in 2012 and headquartered in San Francisco. It provides a no-code extractor builder with AI schema detection, a REST API and SDKs, an MCP server for AI agents, and a fully managed service with SLAs, and processes more than 500 billion data points a month for retailers, brands, marketplaces and data platforms.
What is the best web scraping tool for JavaScript-heavy or protected sites?
One that renders pages in a real browser, handles anti-bot and CAPTCHA, includes residential proxies, and structures the result into typed rows. Import.io does all four inside the platform, on plans from $199 a month, and has run the largest marketplaces on the web in production since 2012.
Do I need to code?
No. Point-and-click and prompt-based extractors build a full pipeline without code. API access, SDKs and webhooks are available on Professional and above for teams that prefer to work programmatically.
Can AI agents use Import.io?
Yes. The MCP server at mcp.import.io exposes extract, search, crawl and monitor as tools for Claude, Cursor and any MCP-compatible client. You get 10,000 free successful calls, then pay $0.0002 per successful call with a spend cap you set. Blocked or failed calls are never billed.
How much does Import.io cost?
A free 30-day trial with 5,000 successful queries and no card. Standard is $199 a month billed annually ($249 monthly) for 50,000 queries; Professional $399 ($499) for 200,000; Advanced $699 ($899) for 500,000. Overage is $7, $3.50 and $2.50 per 1,000. Above 500,000 queries a month or for dedicated infrastructure, managed enterprise contracts apply.
Which websites can Import.io extract data from?
Most public websites, including JavaScript-heavy storefronts, marketplaces, listings portals, PDFs and authenticated portals. Standard includes datacenter proxies, Professional adds regional proxies, Advanced includes premium and residential proxies for heavily protected sites. Managed programs handle the rest.
How is Import.io different from Bright Data, Oxylabs or Zyte?
Those are strongest as unblocking and proxy layers that return pages. Import.io returns structured, validated rows with schema detection, pagination, scheduling, QA and delivery built in, plus a no-code builder, an MCP server and a managed service. See the comparison table above and our detailed comparisons page.
How is Import.io different from Firecrawl or other LLM scrapers?
LLM scrapers are built to turn pages into markdown or loose JSON for RAG. Import.io is built to produce typed, matched, scheduled records at retailer scale with anti-bot, residential proxies and QA, and also serves agents through MCP. Both feed AI systems; Import.io is the one you run a pricing program on.
Is web scraping legal?
Collecting publicly available information is legal in most jurisdictions and standard practice across industries, subject to terms of service, rate limits and personal-data law. Import.io operates rate-aware collection, respects robots and terms, detects and removes personal data, and works under data processing agreements.
What happens when a website changes?
Change is detected on the next run through QA against the extractor's baseline. Extractors are repaired with AI assistance and verified against the schema; on managed programs this happens before your next delivery window and you are told about it.
How is data delivered?
CSV, JSON or Parquet to S3, SFTP or a webhook; pulled from the REST API; or returned directly to an agent through MCP. Screenshots and raw HTML are available alongside structured rows.
What does the managed service include?
Feasibility per source, build, internal AI QA, your UAT, production runs, delivery with a manifest, daily monitoring, weekly summary, monthly report and call, a named CSM and SLAs with fees at risk. First delivery is typically two to three weeks from signature.
INTEGRATIONS
Use your extractor results.
Read scheduled results in a spreadsheet, connect a BI report or load the rows from your own code.
See all integrations