You fetch a page, search the HTML for the data you can see in your browser, and it is not there. The page was built by JavaScript after it loaded, and a plain HTTP request never runs that JavaScript. This guide shows how to confirm that, how to get the rendered HTML from curl or Python, and how to parse it.
The example is quotes.toscrape.com/js/, a public practice site whose quotes are written into the page by a script. Everything shown about it below was checked against the live page.
Before you start
- A ScrapingBot API key. Create a free account for 100 credits, enough for 20 rendered pages. No card needed.
- curl, or Python 3.9 or newer with
pip install requests. The parser uses only the standard library.
Does the page need a browser?
On the page, each quote sits in a div with the class quote. Fetch the URL directly and count them:
curl -s https://quotes.toscrape.com/js/ | grep -c 'class="quote"'
0
Zero. The server sends a page shell, a copy of jQuery and an inline script. The script holds the quotes as a JavaScript array and writes the markup into the page as it loads:
document.write("<div class='quote'><span class='text'>" + d['text'] + "</span><span>by <small class='author'>" + d['author']['name'] + "</small></span><div class='tags'>Tags: " + tags + "</div></div>");
In a browser that script runs, and the document ends up with ten quote blocks like this one:
<div class="quote"><span class="text">“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”</span><span>by <small class="author">Albert Einstein</small></span><div class="tags">Tags: <a class="tag">change</a> …</div></div>
That gap is the test for any site. Find a piece of text you can see on the page, then look for it in view-source (not the Elements panel in DevTools, which shows the page after JavaScript ran) or in the output of curl. If it is missing from the source, you need rendering. If it is there, a plain fetch at 1 credit is enough and faster.
Render it with one parameter
The Website API takes a GET or POST to https://scrapingbot.io/api/v1/scrape. Add render_js=true and the page loads in a real browser, its scripts run, and you get the resulting HTML for 5 credits. -G with --data-urlencode makes curl build and encode the query string for you:
curl -G https://scrapingbot.io/api/v1/scrape \
-H "x-api-key: YOUR_API_KEY" \
--data-urlencode "url=https://quotes.toscrape.com/js/" \
-d render_js=true \
--data-urlencode "wait_for=.quote" \
-d block_resources=true
The response is JSON, with the rendered page in html:
{
"success": true,
"url": "https://quotes.toscrape.com/js/",
"html": "<!DOCTYPE html><html lang=\"en\"><head>\n\t<meta charset=\"UTF-8\">\n\t<title>Quotes to Scrape</title>…",
"status": 200,
"duration": "…",
"credits_used": 5,
"job_id": "…"
}
status is the status the target site answered with, and the HTTP status of the response mirrors it, so a page that answers 404 comes back as a 404 with success: false. Read the JSON body even when your client reports an error. To keep only the HTML from the command line, add | jq -r .html.
Wait for the right moment
A rendered page is captured as soon as the browser reaches its ready point, which by default is when the HTML has been parsed. Pages that fetch their data afterwards need one of these:
| Parameter | Use it when |
|---|---|
wait_for | You know which element holds the data. Pass a CSS selector such as .quote; the capture happens as soon as it appears. This is the best default. |
wait_browser | The content arrives over several network requests. load waits for the page's resources, networkidle2 for no more than two open connections for 500 ms, and networkidle0 for none. |
wait | Nothing else works. A fixed pause in milliseconds, 0 to 45000, always paid in full. |
block_resources | You only need the HTML. It skips images and other heavy resources so the page loads faster; leave it off if a script depends on them. |
All waiting counts toward the 45-second limit for a request. A request that runs out of time returns 408 and is not charged.
Parse it in Python
The script below renders each page, parses the quotes with the standard library's html.parser, follows the pager's Next link and saves everything to JSON. If you already use BeautifulSoup (pip install beautifulsoup4), soup.select(".quote") does the same job as the parser class in fewer lines.
import json
import time
from html.parser import HTMLParser
from urllib.parse import urljoin
import requests
API = "https://scrapingbot.io/api/v1/scrape"
HEADERS = {"x-api-key": "YOUR_API_KEY"}
def render(url, wait_for, tries=3):
"""Fetch a page in a real browser (render_js, 5 credits) and return its HTML."""
params = {
"url": url,
"render_js": "true",
"wait_for": wait_for, # return as soon as this element exists
"block_resources": "true", # skip images and fonts; we only need the DOM
}
for attempt in range(tries):
res = requests.get(API, headers=HEADERS, params=params, timeout=90)
body = res.json()
if body.get("success"):
return body["html"]
if body.get("fault") == "user" or res.status_code in (400, 401, 402):
raise RuntimeError(body.get("error")) # retrying will not help
time.sleep(2 ** attempt) # timeout, 429 or a site error: try again
raise RuntimeError(f"{url} kept failing: {body.get('error')}")
class QuoteParser(HTMLParser):
"""Collects div.quote blocks and the pager's next link from rendered HTML."""
def __init__(self):
super().__init__()
self.quotes, self.next_href = [], None
self.field = None # "text", "author" or "tag" while inside one
self.in_next = False
def handle_starttag(self, tag, attrs):
a = dict(attrs)
classes = (a.get("class") or "").split()
if tag == "div" and "quote" in classes:
self.quotes.append({"text": "", "author": "", "tags": []})
elif self.quotes and tag == "span" and "text" in classes:
self.field = "text"
elif self.quotes and tag == "small" and "author" in classes:
self.field = "author"
elif self.quotes and tag == "a" and "tag" in classes:
self.field = "tag"
self.quotes[-1]["tags"].append("")
elif tag == "li" and "next" in classes:
self.in_next = True
elif tag == "a" and self.in_next:
self.next_href = a.get("href")
def handle_endtag(self, tag):
if tag in ("span", "small", "a"):
self.field = None
if tag == "li":
self.in_next = False
def handle_data(self, data):
if self.field == "tag":
self.quotes[-1]["tags"][-1] += data
elif self.field:
self.quotes[-1][self.field] += data
def scrape_quotes(start_url, max_pages=10):
url, quotes = start_url, []
for page in range(1, max_pages + 1):
parser = QuoteParser()
parser.feed(render(url, wait_for=".quote"))
if not parser.quotes:
raise RuntimeError(f"no .quote elements on {url}; check the selector")
quotes += parser.quotes
print(f"page {page}: {len(parser.quotes)} quotes from {url}")
if not parser.next_href:
break
url = urljoin(url, parser.next_href)
return quotes
if __name__ == "__main__":
quotes = scrape_quotes("https://quotes.toscrape.com/js/")
with open("quotes.json", "w", encoding="utf-8") as f:
json.dump(quotes, f, ensure_ascii=False, indent=2)
print(f"saved {len(quotes)} quotes to quotes.json")
print(json.dumps(quotes[0], ensure_ascii=False, indent=2))
Run it with python quotes.py:
page 1: 10 quotes from https://quotes.toscrape.com/js/
page 2: 10 quotes from https://quotes.toscrape.com/js/page/2/
page 3: 10 quotes from https://quotes.toscrape.com/js/page/3/
…
page 10: 10 quotes from https://quotes.toscrape.com/js/page/10/
saved 100 quotes to quotes.json
{
"text": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”",
"author": "Albert Einstein",
"tags": [
"change",
"deep-thoughts",
"thinking",
"world"
]
}
Ten rendered pages at 5 credits each is 50 credits for the whole site. Two choices in the code matter on other sites too:
- It fails loudly on an empty page. If the selector matches nothing, the script raises instead of saving zero rows. That usually means the site changed its markup or the wait ended too early, and you want to know.
- It follows the link the page gives it.
urljointurns the relative/js/page/2/into a full URL, and the loop stops when there is no Next link, so you never guess how many pages exist.
Check for the data in the source first
Look back at the diagnosis: the quotes were in the plain HTML all along, inside a var data = [...] array that the script reads from. When a page ships its data as JSON in a script tag, you can skip the browser and read it straight from a 1-credit fetch:
import json
import re
import requests
res = requests.get("https://scrapingbot.io/api/v1/scrape",
headers={"x-api-key": "YOUR_API_KEY"},
params={"url": "https://quotes.toscrape.com/js/"}, # plain HTTP, 1 credit
timeout=60)
html = res.json()["html"]
match = re.search(r"var data = (\[.*?\]);", html, re.DOTALL)
data = json.loads(match.group(1))
for q in data[:3]:
print(q["author"]["name"], "-", q["text"][:60])
Albert Einstein - “The world as we have created it is a process of our thinkin
J.K. Rowling - “It is our choices, Harry, that show what we truly are, far
Albert Einstein - “There are only two ways to live your life. One is as though
Many frameworks do this, for example the __NEXT_DATA__ script on Next.js sites, so it is worth searching the source for a value you need before paying for rendering. The trade-off is fragility: a regular expression tied to a variable name breaks more easily than a CSS selector.
Skip the parser with AI extraction
If writing a parser for each site is the slow part, describe what you want instead. ai_query takes plain English; ai_schema takes the exact fields and types. The extracted data comes back as JSON in ai_result, next to the usual html. It adds 5 credits, so a rendered page with extraction costs 10.
curl -X POST https://scrapingbot.io/api/v1/scrape \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://quotes.toscrape.com/js/",
"render_js": true,
"wait_for": ".quote",
"ai_schema": {
"type": "object",
"properties": {
"quotes": {
"type": "array",
"description": "Every quote on the page with its text, author and tags"
}
}
}
}'
Extraction reads the page's text, up to 50,000 characters, so it suits pages like this one. For large lists you will run every week, a parser is cheaper and fully predictable.
Click, type and scroll first
Some content only appears after an interaction: a cookie banner to accept, a "Load more" button, a search box. js_scenario runs up to 49 browser steps before the HTML is captured, and it always renders (5 credits). Send it in a JSON body:
{
"url": "https://shop.example.com",
"js_scenario": [
{ "click": "#accept-cookies" },
{ "fill": ["input[name=q]", "running shoes"] },
{ "click": "button[type=submit]" },
{ "wait_for": ".results" },
{ "scroll_y": 1500 }
]
}
Add "screenshot": "full_page" while you build a scenario: the base64 PNG in the response shows exactly what the browser saw, which makes a wrong selector obvious.
If the site blocks you
Rendering solves missing content, not blocking. If the response is a 403 or a challenge page, step up one level at a time:
| Mode | Parameters | Credits |
|---|---|---|
| Plain HTTP | none | 1 |
| Rendered | render_js=true | 5 |
| Residential IPs | premium_proxy=true, with or without rendering | 10 |
| Stealth | stealth_proxy=true, always renders | 75 |
Blocked attempts are refunded, so starting cheap and stepping up costs little. Use stealth_proxy only for the sites that need it; it cannot be combined with cookies.
Where to go from here
- The JavaScript rendering reference and the full Website API reference list every parameter and response field.
- For why pages get harder to scrape and what to do about it, read web scraping challenges in 2025.
- Prefer to let an AI assistant do the fetching? Add the MCP server to Claude Code or Cursor, and pages come back as markdown.
Common questions
How do I know if a page is rendered with JavaScript?
Compare the page source with what you see. Open view-source (or fetch the URL with curl) and search for a piece of text that is visible on the page. If it is not in the source, the browser is building it with JavaScript, and you need a rendered fetch.
How do I scrape a JavaScript page with curl?
Call https://scrapingbot.io/api/v1/scrape with the page as url and render_js=true. The page loads in a real browser and the html field of the JSON response holds the rendered HTML, which you can pipe into any parser. Add wait_for with a CSS selector so the response waits for the content.
What does rendering cost?
A plain fetch is 1 credit; render_js=true is 5. Premium proxies are 10, stealth proxies 75, and AI extraction adds 5 to any of them. Failed requests and blocked attempts are refunded.
Should I use wait or wait_for?
Prefer wait_for with a selector for the element you need. It returns as soon as that element exists, while wait always pauses for the full number of milliseconds. All waiting counts toward the 45-second limit.
Do I need Selenium or Playwright to scrape a dynamic website in Python?
Not when the browser runs on the API side. Your Python code sends one HTTP request and gets the rendered HTML back, so there is no browser to install, update or keep running on your machine.