Home / Products / ThinkData
Data → knowledge assets

ThinkDataFrom multi-source data to knowledge assets

Unify and govern data scattered across systems, curate it with quality checks, build a searchable and answerable knowledge base, mine it deeply with entity recognition and knowledge graphs, and register it as reusable, grantable knowledge assets. FastAPI + SQLModel + SQLite core, DuckDB curation to Parquet, Vue3 + Element Plus UI.

ThinkData / knowledge base & assets
36sources
53kb docs
1K+entities

One pipeline from raw data to knowledge assets

Collect, govern, curate, build KB, graph, assetize, serve and license — data keeps gaining value as it flows

Multi-source collection

Local-directory, URL-fetch and DB-snapshot collectors with scheduled dispatch and job mutex locks build a reliable data foundation.

Cloud storage governance

Object-storage intranet archiving with sha256 dedup, lifecycle tiering and signed downloads; capacity and cost are visible, staging overflow degrades gracefully.

Curation & quality

In-process DuckDB curation produces Parquet with encoding normalization, row dedup and type inference; a built-in quality rule engine grades results pass/warn/fail.

Document knowledge base

Parses docx/xlsx/pptx/pdf into a knowledge base with FTS5 Chinese full-text and vector hybrid retrieval (RRF fusion) and knowledge-base Q&A.

NER & knowledge graph

Dual-mode NER (rule dictionary + LLM) extracts entities and relations, exports JSON/CSV/Cypher, and renders a force-directed graph you can reason over.

Asset catalog & lineage

Datasets, knowledge bases and services are cataloged together with auto-registered upstream/downstream lineage and cycle detection, under controlled sensitivity and permissions.

Open API service

API keys (SHA256, scopes, rate limits, revocation) expose search, Q&A and dataset-snapshot endpoints for upstream AI apps and digital employees.

SN licensing & online update

One-machine-one-code SN activation, a 90-day free trial and a read-only demo mode when unlicensed; one-click online update with automatic rollback on failure.

Store data — and mine it into knowledge

Built-in document parsing, hybrid retrieval, entity/relation extraction and graph pipelines mine unstructured documents and structured tables alike, forming a searchable, answerable, reasoning-capable knowledge asset.

  • Entities & relations: dual-mode NER (rule dictionary + LLM) extracts and links to grow a living knowledge graph.
  • Hybrid retrieval: FTS5 Chinese full-text and vector (numpy, optional pgvector backend) fused via RRF, with knowledge-base Q&A.
  • Graceful degradation: when embeddings or the LLM are unreachable it falls back to full-text and marks it honestly, never breaking ingestion.
Mining pipeline · entity/graph
NER
graph

Make data a reusable, value-appreciating asset

Processed results are registered as data assets — searchable, grantable, reusable — and, via the open API, linked to the suite's knowledge bases and digital employees so data keeps appreciating in use.

  • Asset catalog: datasets / KBs / services unified, with sensitivity tiers and controlled permissions.
  • Lineage: auto-registers the collect → dataset → KB → service chain and detects cycles.
  • Open service: API keys expose search, Q&A and snapshots for digital employees to call.
Asset catalog · lineage
catalog
lineage
3
collector types
6
doc formats parsed
1click
hybrid search & Q&A
1/1
SN one-machine-one-code

The journey from data to assets

Standardized, traceable, reusable

1

Collect

Local / URL / DB-snapshot ingestion with scheduled dispatch and mutex locks.

2

Govern

Object-storage intranet archiving, dedup and lifecycle tiering with cost visibility.

3

Curate & build KB

DuckDB curation to Parquet with quality checks; parse documents into a searchable, answerable KB.

4

Mine & assetize

NER and graphs mine deeply; register assets and open APIs — searchable, grantable, linked.

Use cases

Enterprise data governanceIndustry knowledge graphDocument assetizationQ&A data supplyPrivate knowledge base

Interface at a glance

Screenshots taken from ThinkData's production UI with real data

Turn data into knowledge assets

ThinkData offers an SN-licensed 90-day free trial and offline on-premises deployment — a single 4GB server is enough. Contact us to book a demo or get a deployment plan.

Book a demo / inquiry

Frequently asked about ThinkData

What is ThinkData?

ThinkData is a data & knowledge platform by ThinkAlike spanning multi-source collection, cloud storage governance, curation & quality, document knowledge base, entity recognition & knowledge graph, asset catalog & lineage, open API service and SN licensing & online update.

Is there a free trial? How is it licensed?

Yes. ThinkData uses SN licensing: after installation you can claim a 90-day free trial online and activate with a one-machine-one-code SN before it expires; unlicensed it runs in a read-only demo mode (readable, not writable) — no data loss.

How does it upgrade safely?

One-click online update from the SN platform: the install script backs up code and database, runs a health check after replacement, and automatically rolls back to the previous version on failure.

What stack does it use?

Backend FastAPI + SQLModel with SQLite for metadata/runtime; DuckDB for curation producing Parquet; retrieval is FTS5 Chinese full-text + vector (numpy, optional pgvector backend) fused via RRF; front end Vue3 + Element Plus.

How does it relate to the knowledge base?

ThinkData processes data into reusable knowledge assets that, through the open API, link with ThinKM and the digital employees for retrieval.

Does it support private deployment?

Yes. ThinkData ships as an offline, on-premises intranet deployment; data never leaves your domain and a single 4GB server is enough. The hosted demo is coming online — for a demo or deployment plan contact contact@thinkalike.com.cn.