A language model only knows what was in its training data and what you put in the prompt. For questions about prices, products, releases or anything that changed this year, the reliable fix is to search first and hand the model the results. This guide builds that step.
You will run a Google search that comes back as JSON, turn the results into a compact context block with numbered sources, optionally add the text of the top pages, and pass the whole thing to whatever model you use. Nothing here depends on a particular LLM vendor: the model is a function you plug in. The search response and output below are real, for the query best espresso machine.
Before you start
- A ScrapingBot API key. Create a free account for 100 credits, enough for 10 searches. No card needed.
- Python 3.9 or newer and
pip install requests. HTML is parsed with the standard library'shtml.parser. - Any LLM you can call from Python, or none yet: the script prints the finished prompt.
One search, as JSON
Searches are a POST to https://scrapingbot.io/api/v1/google with /search as the endpoint. gl picks the country and hl the language:
curl -X POST https://scrapingbot.io/api/v1/google \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"endpoint": "/search", "params": {"q": "best espresso machine", "gl": "us", "hl": "en"}}'
Unlike the social endpoints, the results are at the top level of the response, not under data:
{
"success": true,
"searchParameters": { "q": "best espresso machine", "gl": "us", "hl": "en", "type": "search", "engine": "google" },
"organic": [
{
"position": 1,
"title": "The 7 Greatest Espresso Machines in 2026 [No *BS Guide]",
"link": "https://coffeechronicler.com/gear/espresso-machines/",
"snippet": "That being said, Rancilio Silvia is the gold standard when it comes to home espresso makers, and it has been this way for more than two decades ...",
"date": "Jul 11, 2026"
},
{
"position": 6,
"title": "Top Rated Espresso Machines",
"link": "https://www.seattlecoffeegear.com/collections/espresso-machines?srsltid=AU7gw4VzkkFBIJqO7z…",
"snippet": "Our Top Espresso Machines · Diletta Alto Espresso Machine · Breville Bambino Plus Baratza Encore ESP Bundle · Breville Barista Express ..."
},
…
],
"peopleAlsoAsk": [
{ "question": "What is the highest rated espresso machine for home use?" },
…
],
"relatedSearches": [
{ "query": "Best espresso machine for home" },
…
],
"duration": "0.85",
"statusCode": 200,
"creditsUsed": 10
}
This search returned 9 organic results, 4 People also ask questions and 8 related searches. A few things are worth knowing before you build on it:
dateis optional and loosely formatted. It wasJul 11, 2026on one result,6 months agoon another and missing on three. Pass it to the model as text; do not parse it as a date.- Some links carry a tracking parameter. Four of the nine links ended in a long
srsltidquery parameter. It adds tokens and makes citations ugly, so the script removes it. - People also ask holds questions only in this response, with no answer text. They are still useful: they tell the model what people usually mean by the query.
- There are no ads, and at most 10 organic results per page. Use
pagefor more.
From results to a cited context block
Save this as search_context.py. answer() runs the search, optionally fetches the top pages, builds the context block, wraps it in a prompt and, if you pass a model, asks it.
import sys
import time
from html.parser import HTMLParser
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
import requests
HEADERS = {"x-api-key": "YOUR_API_KEY"}
GOOGLE = "https://scrapingbot.io/api/v1/google"
SCRAPE = "https://scrapingbot.io/api/v1/scrape"
credits_used = 0
def google_search(q, gl="us", hl="en", tbs=None, page=1, tries=3):
"""One Google results page as JSON: 10 credits. Failed calls are refunded."""
global credits_used
params = {"q": q, "gl": gl, "hl": hl, "page": page}
if tbs:
params["tbs"] = tbs # qdr:h, qdr:d, qdr:w, qdr:m or qdr:y
for attempt in range(tries):
res = requests.post(GOOGLE, headers=HEADERS, timeout=60,
json={"endpoint": "/search", "params": params})
body = res.json()
if body.get("success"):
credits_used += body.get("creditsUsed", 0)
return body
if res.status_code in (400, 401, 402): # bad input, bad key, no credits
raise RuntimeError(body.get("error"))
time.sleep(2 ** attempt) # 408, 429 or 5xx: wait and try again
raise RuntimeError(f"search kept failing: {body.get('error')}")
class _Text(HTMLParser):
"""Collect visible text, skipping scripts, styles and other non-content tags."""
SKIP = {"script", "style", "noscript", "svg", "template", "nav", "footer"}
def __init__(self):
super().__init__()
self.parts, self.depth = [], 0
def handle_starttag(self, tag, attrs):
if tag in self.SKIP:
self.depth += 1
def handle_endtag(self, tag):
if tag in self.SKIP and self.depth:
self.depth -= 1
def handle_data(self, data):
if not self.depth and data.strip():
self.parts.append(data.strip())
def page_text(url, max_chars=3000):
"""Fetch a result page through the website API (1 credit) and keep its text."""
global credits_used
res = requests.get(SCRAPE, headers=HEADERS, params={"url": url}, timeout=60)
body = res.json()
credits_used += body.get("credits_used", 0)
if not body.get("success"):
return None # blocked or broken page: skip it, the snippet still counts
parser = _Text()
parser.feed(body["html"])
return " ".join(parser.parts)[:max_chars]
def clean_link(url):
"""Drop Google's srsltid tracking parameter so citations stay readable."""
parts = urlsplit(url)
query = [(k, v) for k, v in parse_qsl(parts.query) if k != "srsltid"]
return urlunsplit(parts._replace(query=urlencode(query)))
def build_context(results, pages=None, max_sources=8):
"""A compact, numbered context block plus the list of sources it cites."""
sp = results["searchParameters"]
lines = [f'Google results for "{sp["q"]}" ({sp.get("gl", "")}/{sp.get("hl", "")}):']
sources = []
for r in results.get("organic", [])[:max_sources]:
n = len(sources) + 1
link = clean_link(r["link"])
date = f" ({r['date']})" if r.get("date") else ""
lines.append(f"[{n}] {r['title']}{date}\n {link}\n {r.get('snippet', '')}")
if pages and pages.get(r["link"]):
lines.append(f" Page text: {pages[r['link']]}")
sources.append({"n": n, "title": r["title"], "url": link})
questions = [p["question"] for p in results.get("peopleAlsoAsk", [])]
if questions:
lines.append("People also ask:\n" + "\n".join(f"- {q}" for q in questions))
return "\n".join(lines), sources
PROMPT = """Answer the question using only the numbered sources below.
Cite every claim with its source number, like [2]. If the sources do not
answer the question, say so instead of guessing.
Question: {question}
{context}"""
def answer(question, llm=None, fetch_pages=0, **search_options):
"""llm is any function that takes a prompt string and returns text."""
results = google_search(question, **search_options)
pages = {}
for r in results.get("organic", [])[:fetch_pages]:
text = page_text(r["link"])
if text:
pages[r["link"]] = text
context, sources = build_context(results, pages)
prompt = PROMPT.format(question=question, context=context)
related = [s["query"] for s in results.get("relatedSearches", [])]
reply = llm(prompt) if llm else None
return {"prompt": prompt, "answer": reply, "sources": sources, "related": related}
if __name__ == "__main__":
question = sys.argv[1] if len(sys.argv) > 1 else "best espresso machine"
fetch_pages = int(sys.argv[2]) if len(sys.argv) > 2 else 0
out = answer(question, fetch_pages=fetch_pages) # pass llm=your_model to get an answer
print(out["prompt"])
print("\nfollow-up searches:", "; ".join(out["related"][:4]))
print(f"{len(out['sources'])} sources, {len(out['prompt'])} characters, "
f"credits used: {credits_used}")
Design choices that matter for an LLM:
- Numbered sources, with the instruction to cite them. Each result becomes
[n] title (date), its URL and its snippet. The prompt tells the model to cite by number and to say so when the sources do not answer. The returnedsourceslist maps numbers back to URLs, so you can turn[2]in the reply into a link. - Eight sources, not ten. Every extra result costs tokens, and the lowest-ranked ones tend to add the least.
max_sourcesis easy to change. - Related searches stay out of the prompt. They are returned separately, for an agent that wants to decide on a second, narrower search.
- Credits are counted from the responses. Search responses report
creditsUsedand website responsescredits_used; the script adds both.
Run it
python search_context.py "best espresso machine"
This is the prompt your model receives, built from the real response above:
Answer the question using only the numbered sources below.
Cite every claim with its source number, like [2]. If the sources do not
answer the question, say so instead of guessing.
Question: best espresso machine
Google results for "best espresso machine" (us/en):
[1] The 7 Greatest Espresso Machines in 2026 [No *BS Guide] (Jul 11, 2026)
https://coffeechronicler.com/gear/espresso-machines/
That being said, Rancilio Silvia is the gold standard when it comes to home espresso makers, and it has been this way for more than two decades ...
[2] The Best Espresso Machine (Apr 7, 2024)
https://coffeegeek.com/opinions/state-of-coffee/the-best-espresso-machine/
The Breville Bambino Plus: The Best Espresso Machine of All Time · PID stable temperature controls (200F, non changeable) at the grouphead
…
[6] Top Rated Espresso Machines
https://www.seattlecoffeegear.com/collections/espresso-machines
Our Top Espresso Machines · Diletta Alto Espresso Machine · Breville Bambino Plus Baratza Encore ESP Bundle · Breville Barista Express ...
[7] The Best Home Espresso Machine (Mar 17, 2026)
https://www.nytimes.com/wirecutter/reviews/best-espresso-machine-grinder-and-accessories-for-beginners/
After making (and tasting) dozens of espressos and lattes, we think the Profitec Go is the best machine for both new and skilled espresso ...
[8] Professional Espresso Machines
https://www.chriscoffee.com/collections/espresso-machines
The La Spaziale S11 Brio is a compact, professional-grade espresso machine designed for home baristas seeking café-quality results. With advanced PID ...
People also ask:
- What is the highest rated espresso machine for home use?
- What is the best espresso machine for home use?
- What is the best espresso coffee machine for home use?
- What is the best espresso machine right now?
follow-up searches: Best espresso machine for home; Best espresso machine professional; Professional espresso machine for home; Best espresso machine for home automatic
8 sources, 2792 characters, credits used: 10
2,792 characters is roughly 700 tokens by the usual estimate of four characters per token: small enough to add to any request. Notice what the snippets already give the model. Different sources name different machines (Rancilio Silvia, Breville Bambino Plus, Gaggia Classic Pro, Profitec Go), each with a date, so a model told to cite can say who recommends what and how recent the claim is, instead of naming one machine from memory.
Plug in your model
llm is any function that takes the prompt as a string and returns text. Wrap whichever client you use: a hosted API, a local model or a gateway. Here it is with a placeholder you would replace:
def my_model(prompt):
# Replace this body with a call to your model's API or SDK,
# and return its reply as a string.
return f"(model reply to a {len(prompt)}-character prompt)"
out = answer("best espresso machine", llm=my_model)
print(out["answer"])
for s in out["sources"][:3]:
print(f"[{s['n']}] {s['title']} - {s['url']}")
(model reply to a 2792-character prompt)
[1] The 7 Greatest Espresso Machines in 2026 [No *BS Guide] - https://coffeechronicler.com/gear/espresso-machines/
[2] The Best Espresso Machine - https://coffeegeek.com/opinions/state-of-coffee/the-best-espresso-machine/
[3] Best espresso machine for home use under moderate ... - https://www.reddit.com/r/JamesHoffmann/comments/1rsgpd9/best_espresso_machine_for_home_use_under_moderate/
If your model takes a system prompt, move the first paragraph of PROMPT there and send the question and context as the user message. Either way, keep the context after the question and the instruction to cite before it.
Add the page text when snippets are not enough
Snippets are one or two sentences, cut off mid-thought. For questions that need detail, pass a number as the second argument and the script fetches that many top results through the website API, strips them to text and adds the text under each source:
python search_context.py "best espresso machine" 3
Each page is a plain GET https://scrapingbot.io/api/v1/scrape?url=… at 1 credit, so that run costs 13 credits: 10 for the search and 3 for the pages. page_text() keeps only visible text, drops scripts, styles, navigation and footers, and caps each page at 3,000 characters so one long article cannot crowd out the rest. A page that fails to load returns None and the source keeps just its snippet.
When a page comes back empty
Some sites build their content with JavaScript, so the plain HTML has little text. Add "render_js": "true" to the params in page_text() for those sites; it costs 5 credits instead of 1. Scrape a JavaScript-rendered page covers how to tell. Do not send Google results URLs to the website API; it rejects them and points you to /search.
Search options worth using
| Parameter | What it does | Example |
|---|---|---|
gl | Country of the results. | us, gb, de |
hl | Language of the interface and snippets. | en, es, fr |
tbs | Only recent results: past hour, day, week, month or year. | qdr:d |
page | The next 10 results, 10 credits per page. | 2 |
q operators | site:, quotes and - work as on Google. | site:reddit.com espresso |
For news-like questions, tbs="qdr:w" is the simplest way to keep stale articles out of the context. answer() passes extra keyword arguments to the search, so answer(question, llm=my_model, tbs="qdr:w") is all it takes.
Let an agent search on its own
If your agent runs in an MCP client such as Claude Code or Cursor, you may not need this code at all. ScrapingBot's hosted MCP server exposes a googleSearch tool with q, gl, hl, num, page and tbs, plus a scrapeWebsite tool that returns pages as markdown. The agent decides when to search and what to read next. Add a web scraping MCP server to Claude Code and Cursor sets it up in one command.
Use the Python route when you want control: a fixed prompt format, a cap on sources and spend, or a pipeline that runs without a chat client.
What it costs
Each search is 10 credits; each plain page fetch is 1. On the Starter plan ($49.99 for 275,000 credits) that is about $1.82 per 1,000 searches, or about $2.36 per 1,000 if each one also reads three pages. Failed calls are refunded, and the free 100 credits cover 10 searches.
Where to go from here
- The Google API reference lists every parameter, plus the images, news, shopping and places endpoints.
- The Google Search API page has pricing and examples, and how to scrape Google results without getting blocked explains why an API beats scraping the results page yourself.
Common questions
How do I give an LLM live Google results?
Call the Google /search endpoint, which returns the results page as JSON, then turn the organic results and People also ask questions into a short numbered text block and put it in the prompt with an instruction to cite source numbers. The model never needs to browse.
Is a SERP API enough for RAG, or do I need the pages too?
Snippets are often enough for "what is" and "which one" questions and cost nothing extra. When the answer needs detail, fetch the top two or three result pages through the website API (1 credit each for a plain request) and add their text under each source.
How do I get only recent results?
Pass tbs: qdr:h for the past hour, qdr:d for the past day, qdr:w week, qdr:m month or qdr:y year. Use gl and hl to pick the country and language of the results.
How many results does one search return?
Up to 10 organic results per page; Google returns no more than that. Ask for page: 2 and so on for more, at 10 credits per page.
Can an agent call this without my Python code?
Yes. ScrapingBot's hosted MCP server has a googleSearch tool with the same parameters, so Claude Code, Cursor and other MCP clients can search Google directly. Each tool call is billed like the API call.