Send a URL and the fields you want, as a schema or in plain English. ScrapingBot loads the page, has AI read it and returns your fields as JSON. No CSS selectors to write, and none to fix when the site changes.
curl -X POST https://scrapingbot.io/api/v1/scrape \ -H "x-api-key: YOUR_API_KEY" \ -H "content-type: application/json" \ -d '{ "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html", "ai_schema": { "properties": { "title": { "type": "string" }, "price": { "type": "number" }, "in_stock": { "type": "boolean" }, "available": { "type": "number" }, "upc": { "type": "string" } } } }'
{ "success": true, "status": 200, "html": "<!DOCTYPE html>…", "ai_result": { "title": "A Light in the Attic", "price": 51.77, "in_stock": true, "available": 22, "upc": "a897fe39b1053632" }, "credits_used": 6 }
Write a small JSON schema with the fields and types you want, or ask in plain English with ai_query.
One POST to /api/v1/scrape. Add render_js for pages that build themselves in the browser.
Your fields come back as JSON in ai_result, with the page HTML next to them. Store the rows and move on to the next URL.
import requests API = "https://scrapingbot.io/api/v1/scrape" HEADERS = {"x-api-key": "YOUR_API_KEY"} SCHEMA = { "type": "object", "properties": { "title": {"type": "string"}, "price": {"type": "number", "description": "Price, number only"}, "in_stock": {"type": "boolean"}, }, } product_urls = ["https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"] rows = [] for url in product_urls: r = requests.post(API, headers=HEADERS, json={"url": url, "ai_schema": SCHEMA}).json() if r.get("ai_result"): rows.append({"url": url, **r["ai_result"]}) print(rows[0]) # {'url': ..., 'title': ..., 'price': 51.77, 'in_stock': True}
One request shape for very different pages.
Title, price and stock from store pages, without writing a parser for every shop.
Every item on a page as an array: names, links, locations, prices, dates.
Feed an agent or RAG pipeline fields it can use instead of raw HTML full of markup.
Run the same schema on a schedule and compare the rows to spot changes on public pages.
Two ways. ai_schema is a JSON schema: name each field, give it a type and, if it helps, a short description. You get back an object with exactly those keys. ai_query is a plain-English request, such as "the quotes on this page with their authors", and the AI picks the key names.
In ai_result, next to the page html, so you can keep the source alongside the fields. If the page loads but the fields cannot be extracted, the response has ai_error instead.
No. The AI reads the page text and returns the fields you asked for. When a site changes its layout, the same request usually keeps working, so there are no selectors to fix.
Add render_js: true and the page is loaded in a real browser before extraction. You can also wait for an element with wait_for.
The page fetch plus 5 credits: 6 credits for a plain page, 10 with render_js. On the Starter plan ($49.99 for 275,000 credits) that is about $1.09 per 1,000 plain pages. New accounts get 100 free credits, enough for 16 extractions, with no card.
If the page cannot be fetched (an error or a timeout), the call is refunded automatically. If the page loads but the AI cannot find what you asked for, missing fields come back empty or you get ai_error, and the call counts as a normal page.
Public pages: product pages, listings, articles, directories, docs. It is not for pages behind a login or for private data.
Create an account and the playground opens with a schema filled in. No credit card required.