# ThinSpider — A multi-tenant cloud crawl engine for engineers

> ThinSpider is a multi-tenant cloud crawl engine by ThinkAlike. Static pages, list-detail sites, SPA rendering and API endpoints share one task DSL; submit it through the step-by-step console wizard or a Firecrawl-style REST API, run it on tenant-isolated queues and an elastic worker pool, and take results out as JSON, CSV or webhook callbacks. One unified login spans the product suite, with per-instance licensing for on-premises delivery.

- Product: ThinSpider / ThinSpider
- Brand: ThinkAlike (Beijing ThinkAlike Technology Co., Ltd.)
- Trial: https://thinspider.thinkalike.com.cn
- Page: https://thinkalike.com.cn/en/thinspider.html

## Core capabilities
- **Four collection shapes** — Static and list-detail pages are parsed directly; SPA pages go through a rendering browser pool before content is taken; pure data endpoints are collected field by field.
- **One task DSL** — A single JSON declares entry URLs, traversal scope, depth and target fields, built in the console wizard or hand-written.
- **Firecrawl-style REST** — /v1/scrape, /v1/crawl, /v1/map and /v1/tasks/{id}; naming aligned with mainstream crawl services for drop-in integration.
- **Cleaning &amp; extraction** — URLs are normalized and deduped by content hash; extraction supports explicit rules with model assistance, degrading to CSS selectors when needed.
- **Isolated &amp; elastic** — Per-tenant queues with round-robin fairness so a slow job never starves others; the worker pool scales horizontally and business data stays separate.
- **Unified ID &amp; on-prem** — Auth code + PKCE with offline token verification for cross-product sign-in; the same core ships as an air-gapped, per-instance licensed package.

## Highlights
### From an entry URL to clean structured fields
Once submitted, the engine discovers links, normalizes and dedups, renders only when needed, and extracts body and custom fields into structured data ready for ThinkData/ThinKM.

### Console wizard, or a set of /v1 endpoints
The same task DSL has two doors: business users build it in the console, engineers POST /v1/crawl from their own systems — with progress, item counts and failure reasons coming from one source.

## Use cases
Content aggregation · Industry news · Competitor &amp; public data · SPA site collection · List-detail traversal · Knowledge-base corpus

ThinSpider targets publicly accessible data and does not promise unconditional breakthrough of any site; respect the destination's access rules and applicable law.

## FAQ
**What is ThinSpider?**

A <b>multi-tenant cloud crawl engine</b> by ThinkAlike: static pages, list-detail sites, SPA rendering and API endpoints declared in one <b>task DSL</b>, submitted via console or REST, delivered as JSON / CSV / webhook.

**How does the API relate to Firecrawl?**

It follows <b>Firecrawl-style</b> endpoint naming and request shapes (/v1/scrape, /v1/crawl, /v1/map, /v1/tasks/{id}, /v1/webhooks) for drop-in integration; the product's /v1 docs define exact fields.

**Does it support scheduled crawls?**

This release is <b>submission and queue driven</b>: the schedule field is recorded only and does not auto-trigger yet. Call /v1/crawl from your own scheduler for recurring collection.

**Can it crawl any website?**

It <b>does not promise unconditional breakthrough</b>. ThinSpider targets publicly accessible data, must respect the destination site's rules and applicable law, and runs within per-tenant quotas.

**How do results reach a knowledge base?**

Export <b>JSON / CSV</b> or receive <b>webhook callbacks</b> into your own system, then deposit into <b>ThinkData</b> and <b>ThinKM</b> for retrieval and Q&amp;A.

**How is it deployed on-premises?**

SaaS and the private package run the <b>same core</b>. Instances are activated with an <b>SN license file</b> (Ed25519 signature + AES-256-GCM protected content); unactivated instances block sensitive writes.
