services / page-extract
weboperational · 1721 ms
Structured Data Extraction (AI)
Extract structured JSON from a web page or text with fields you define: send url (or text) and fields like {"price": "price in EUR as number", "in_stock": {"type": "boolean"}}, get back exactly those fields, typed and validated against a JSON schema. Types: string, number, integer, boolean, string[], number[]. Fields the content does not state come back as null instead of being guessed. Up to 25 fields; optional instructions for extra rules. For product data, prices, contact and company details, job postings, event data, specs. Uses Claude Haiku 5.5 with structured outputs; page content is treated as data, instructions inside it are ignored. Not charged when the page cannot be read.
Run free trial ↗3 free calls per day with the example input. Paid: $0.007 USDC, no limit.
Call it
Input
| Field | Type | Description |
|---|---|---|
| url | string | Page to extract from (or send text) |
| text | string | Raw text instead of a URL, max 60,000 chars |
| fields * | object | Field name -> description string, or {type, description}. Types: string, number, integer, boolean, string[], number[] |
| instructions | string | Optional extra rules, max 500 chars |
Output
| Field | Type | Description |
|---|---|---|
| source | string | |
| title | string | |
| data | object | |
| fields_filled | integer | |
| fields_total | integer | |
| input_truncated | boolean | |
| model | string | |
| extracted_at | string |
Example response (data)
{
"source": "https://www.zalando.de/zalando-impressum/",
"title": "Zalando wohnen",
"data": {
"company": "Zalando SE",
"city": "Berlin",
"register_number": "HRB 158855 B",
"vat_id": "DE 260543043"
},
"fields_filled": 4,
"fields_total": 4,
"input_truncated": false,
"model": "claude-haiku-5-5",
"extracted_at": "2026-10-10T09:58:24.270Z"
}