Home / Products / ThinSpider
Multi-tenant crawl engine

ThinSpiderAny platform, any data — one task DSL takes it all

Static pages, list-detail sites, SPA rendering and pure API sources share one task model. Submit it through the step-by-step console wizard or a Firecrawl-style REST API, run it on tenant-isolated queues and an elastic worker pool, and take results out as JSON, CSV or webhook callbacks — ready to deposit into ThinkData and ThinKM. One unified login spans the product suite, with per-instance licensing for on-premises delivery.

thinspider.thinkalike.com.cn / crawl console
4source shapes
2submit paths
3result outlets

One engine, four source shapes

Declare, submit, extract and deliver — end to end

Four collection shapes

Static and list-detail pages are parsed directly; SPA pages go through a rendering browser pool before content is taken; pure data endpoints are collected field by field.

One task DSL

A single JSON declares entry URLs, traversal scope, depth and target fields. Business users build it in the console wizard; engineers hand-write it and submit as-is.

Firecrawl-style REST

/v1/scrape for a page, /v1/crawl for a site, /v1/map for link discovery, /v1/tasks/{id} for progress — naming aligned with mainstream crawl services for drop-in integration.

Cleaning & extraction

URLs are normalized and deduped by content hash so one page never enters twice; extraction supports explicit rules with model assistance, degrading to CSS selectors when needed.

Isolated & elastic

Each tenant gets its own queue with round-robin fairness, so a slow job never starves others; the worker pool scales horizontally and business data stays in separate stores.

Unified ID & on-prem

Sign in once at the unified identity center (auth code + PKCE, offline token verification) and move across products; the same core ships as an air-gapped package with per-instance licensing.

From an entry URL to clean structured fields

Once a task is submitted the engine discovers links, normalizes and dedups, renders only when needed, and turns pages into searchable, reusable structured data.

  • Link discovery: traverse within the declared scope, deduping as it goes.
  • On-demand rendering: static pages parsed directly, SPAs rendered first.
  • Field landing: title, body, time and custom fields feed ThinkData/ThinKM.
Task execution · queue / worker pool
discover
extract

Console wizard, or a set of /v1 endpoints

The same task DSL has two doors: business users click it together in the console, engineers simply POST /v1/crawl from their own systems — no private protocol to learn.

  • Standard API: /v1/scrape · /v1/crawl · /v1/map · /v1/tasks/{id} · /v1/webhooks.
  • Same state: progress, item counts and failure reasons come from one source for both doors.
  • Direct output: download JSON / CSV, or let webhooks push results downstream.
POST /v1/crawl · submit a task
task DSL
result
4
source shapes
2
submit paths
3
result outlets
1
login cross-product

Four steps from login to results

Sign in, declare the task, let the engine run, take the data

1

Sign in

Redirect to the unified identity center; one login grants the token.

2

Declare

Build the task in the wizard or POST /v1/crawl with entry, scope and fields.

3

Execute

The task enters its tenant queue; isolated workers render and crawl as needed.

4

Deliver

JSON / CSV export or webhook callbacks; deposit into ThinkData/ThinKM.

Use cases

Content aggregationIndustry newsCompetitor & public dataSPA site collectionList-detail traversalKnowledge-base corpus

ThinSpider targets publicly accessible data and does not promise unconditional breakthrough of any site; respect the destination's access rules and applicable law.

Interface at a glance

Turn any platform's data into usable input

Try ThinSpider's four collection shapes, task DSL and tenant-isolated engine.

Try ThinSpider now

Frequently asked about ThinSpider

What is ThinSpider?

A multi-tenant cloud crawl engine by ThinkAlike: static pages, list-detail sites, SPA rendering and API endpoints declared in one task DSL, submitted via console or REST, delivered as JSON / CSV / webhook.

How does the API relate to Firecrawl?

It follows Firecrawl-style endpoint naming and request shapes (/v1/scrape, /v1/crawl, /v1/map, /v1/tasks/{id}, /v1/webhooks) for drop-in integration; the product's /v1 docs define exact fields.

Does it support scheduled crawls?

This release is submission and queue driven: the schedule field is recorded only and does not auto-trigger yet. Call /v1/crawl from your own scheduler for recurring collection.

Can it crawl any website?

It does not promise unconditional breakthrough. ThinSpider targets publicly accessible data, must respect the destination site's rules and applicable law, and runs within per-tenant quotas.

How do results reach a knowledge base?

Export JSON / CSV or receive webhook callbacks into your own system, then deposit into ThinkData and ThinKM for retrieval and Q&A.

How is it deployed on-premises?

SaaS and the private package run the same core. Instances are activated with an SN license file (Ed25519 signature + AES-256-GCM protected content); unactivated instances block sensitive writes.