Extract
Public HTML as clean text, when publisher and network policy allow it.
Live contract — August 11, 2026: Extract accepts public HTML
URLs only, up to 2,048 characters. It checks the index first. A live fetch on a miss is
gated: anonymous callers and free browser JWTs stay index-only; Growth plan
JWTs and API keys may live-fetch. Robots rules, noindex, nosnippet,
and private-network protections can refuse a URL.
What it does
One tool: extract_url(url). It returns title, text, and passages (paragraph blocks), plus a source field that tells you where the content came from:
"index"— we already crawled this page. The text is byte-identical to whatsearch_webquotes and carries its crawl date, so the page your agent cited is the page it can pull verbatim. No network call is made."live"— we hadn't crawled it, so it was fetched on the spot. Nothing is stored. Live fetch requires Growth (browser) or an API key — not the free/anon tier.
What it honors
We respect the publisher's own directives on both paths: a noindex page is refused outright, and a nosnippet page returns its title and link with empty text. That's a feature — a page we're not allowed to serve comes back as an explicit refusal, never as quietly-degraded content.
Why the limit is tighter
Extract runs at 10 calls per minute — tighter than search on purpose. An index hit is one local lookup, but a miss makes SeaWeb fetch someone else's page on your say-so.