Hidden JSON APIs#
The page you want to scrape has already done the work for you. Find the JSON endpoint its own JavaScript calls, and skip the HTML entirely.
โฑ ~8 min read ยท ~15 min hands-on ๐ needs: HTTP clients ยท browser DevTools (Network tab)
Most modern sites load a nearly-empty HTML shell, then fetch their real data as JSON from a backend API. If you find that request, you get clean, structured data โ no brittle CSS selectors, no headless browser, often 100ร faster. This is the first thing to try on any dynamic site, before you reach for Playwright.
Try it in 5 minutes#
We’ll use quotes.toscrape.com/scroll โ a sandbox built for exactly this. It loads quotes as you scroll, so the data clearly isn’t in the first HTML response.
- Open https://quotes.toscrape.com/scroll in Chrome.
- Open DevTools (
F12orCtrl/Cmd+Shift+I) โ Network tab. - Click Fetch/XHR to filter out images, CSS, and fonts โ leaving only data requests.
- Reload the page and scroll down. Watch the rows appear.
- Click the request to
api/quotes?page=1. Open the Response tab โ that’s your data as JSON. - Right-click the request โ Copy โ Copy as cURL, and paste it into a terminal.
curl 'https://quotes.toscrape.com/api/quotes?page=1'โ You just got the same data the page shows โ as structured JSON, with no browser and no HTML parsing.
Why this works#
flowchart LR
B["Your browser"] -->|"1 ยท GET /scroll"| S["Server"]
S -->|"2 ยท empty HTML + JS"| B
B -->|"3 ยท JS calls GET /api/quotes?page=1"| A["Hidden JSON API"]
A -->|"4 ยท clean JSON"| B
B -->|"5 ยท JS paints the DOM"| D["What you see"]The HTML at step 2 is a shell โ the quotes aren’t in it. The browser runs JavaScript (step 3) that calls the real data source. Scrapers that only read step 2’s HTML find nothing; that’s why people wrongly conclude “this site needs a browser.”
The shortcut: call step 3 yourself and stop. You never need steps 1, 2, or 5.
Here’s the full extraction as a self-contained script โ no project setup, just uv run:
# /// script
# requires-python = ">=3.12"
# dependencies = ["httpx>=0.28", "polars>=1.0"]
# ///
"""Pull every quote from quotes.toscrape.com's hidden JSON API into Parquet.
Run: uv run hidden_api.py
"""
import httpx
import polars as pl
API = "https://quotes.toscrape.com/api/quotes"
def fetch_all() -> list[dict]:
rows, page = [], 1
# A Client reuses the TCP connection across pages โ faster and politer.
with httpx.Client(timeout=10, headers={"User-Agent": "tds-course-demo"}) as client:
while True:
# httpx's raise_for_status() returns the response, so we can chain .json()
data = client.get(API, params={"page": page}).raise_for_status().json()
for q in data["quotes"]:
rows.append(
{
"text": q["text"],
"author": q["author"]["name"],
"tags": ", ".join(q["tags"]),
}
)
if not data["has_next"]: # the API tells you when to stop
break
page += 1
return rows
if __name__ == "__main__":
rows = fetch_all()
pl.DataFrame(rows).write_parquet("quotes.parquet")
print(f"Saved {len(rows)} quotes to quotes.parquet")Notice what the API handed you for free: a has_next flag so you know when to stop, and a stable page structure. You wrote a loop, not a fragile HTML parser.
โ๏ธ Before you replay a request against a real site, check its Terms of Service and
robots.txt, and keep your rate low. An internal API being reachable is not the same as being allowed. See Legal & Ethical Scraping.
When it fails#
Copying the URL alone often isn’t enough โ the browser sent headers or cookies the server checks. Replay the whole request (that’s why “Copy as cURL” copies the headers too), then strip pieces away until you find the minimum that still works.
| Symptom on replay | Likely cause | Fix |
|---|---|---|
403 Forbidden | Server checks Referer / User-Agent / X-Requested-With | Send the same headers the browser did |
401 Unauthorized | Endpoint needs a token | Copy the Authorization header or cookie from the request |
| Empty result, no error | Needs a session cookie set by the HTML page | Hit the page first with an httpx.Client, reuse its cookies |
| Works once, then blocks you | Rate limit or expiring token | Slow down; refresh the token โ Rate Limits, Retries & Caching |
| No JSON anywhere in Fetch/XHR | Data is server-rendered into the HTML | This trick won’t help โ parse the HTML or use Playwright |
Pro tip: in the Network tab, use Search (Ctrl/Cmd+F) and type a value you can see on the page (an author’s name, a price). It jumps straight to the request that contains it โ faster than reading every row. Also look for URLs with api, graphql, /v1/, .json, or query in them.
Your turn (โ15 min)#
- Run the script above with
uv run hidden_api.pyand confirm you getquotes.parquet(~100 rows). - Pick a different endpoint on the same sandbox: open
quotes.toscrape.com/api/quotes?page=1and add filtering โ collect only quotes taggedlove. (Hint: the JSON also carries atagslist per quote.) - Query your Parquet without loading it into Python โ this is a one-liner with DuckDB:
duckdb -c "SELECT author, count(*) n FROM 'quotes.parquet' GROUP BY author ORDER BY n DESC LIMIT 5" - Stretch: open the Network tab on a real site you use (a news site, a store) and just find its data API. Don’t hammer it โ one look is the exercise.
Checklist#
- I can filter the Network tab to Fetch/XHR and find the request that returns the data.
- I can use Network Search to jump to the request containing a value I see on the page.
- I can turn a browser request into a replayable one with Copy as cURL.
- I can replay it with
httpxand know which headers decide success or a 403. - I can recognise when a site has no hidden API and I need a browser instead.
Go deeper#
- How I found the easiest way to scrape this site โ John Watson Rooney โ the network-tab technique, start to finish.
- How to Scrape Hidden APIs โ Scrapfly โ headers, pagination, and auth patterns in depth.
- Finding Undocumented APIs โ Inspect Element โ a friendly walk-through with real examples.
- httpx Clients โ official docs โ sessions, cookies, and connection reuse.
