A structured knowledge pipeline for scientific discovery

Crystal

Just like the universe, available information keeps expanding. Yet identifying meaningful insights only becomes harder without the proper navigation tools. Crystal transforms unstructured text into a queryable knowledge graph, surfacing non-obvious relationships, patterns, and potential knowledge gaps to guide hypothesis generation across fields that may otherwise remain disconnected or overlooked by individual researchers.

Same method, different domains. Questions fit for each. Fig. 1 - Illustrative; not live output
SCIENCE DOMAIN: NEUROBIOLOGY OF VOCAL LEARNING FOXP2-LINKED VOCAL CIRCUIT HVC RA Area X SRPX2 0 PubMed results Potential gap identified FICTION DOMAIN: CHARACTERS IN FINAL FANTASY PROTECTIVE LOYALTY Vivi Garnet Auron Zidane A protector becomes the protected
archetype entity gap: PubMed-verified absence structurally notable: glow + spark

the problem

Search finds papers.
It doesn't build understanding.

Current tools for navigating text sources surface connections like keyword matches and cross-document references, or provide unstable, opaquely-generated summaries. What's missing is a way to navigate the map of concepts that emerges from the text itself, in a traceable, reproducible manner.

01
Keyword search — Finds documents that contain a term. Returns a pile, not a picture.
PubMed, Google Scholar
02
Citation networks — Shows which papers reference each other. Still says nothing about what they claim, and is shaped by author behavior.
connected papers
03
AI chat summaries — Read for you, fluently, but in a non-transparent manner that can be hard to repeat.
no persistent structure
04
Crystal — Transforms entities and relationships into a graph where every claim carries its source, surfacing non-obvious structural insights fully transparently.
structured + traceable + reproducible

how it works

A shared core. Different questions, per domain.

The pipeline doesn't know or care whether it's reading a review article or an official encyclopedia of a complex, fictional world — the same extraction, normalization, and clustering machinery builds a conceptual topology either way. Only in the final step does that shared structure become a domain-specific question asked against a map you can already understand. The method works on domains as different as natural science and human storytelling, reflecting shared principles in the representation of complex knowledge that can be explored as a navigable conceptual topology.

0
Entity discovery & extraction
The model analyzes the raw text to identify individual entities that carry consistent meaning — genes, places, characters, whatever the text continually references — then pulls every relevant sentence for each one, copied exactly as written (no paraphrasing). A follow-up step catches cases where two sources call the same thing by different names and merges them.
1
Atomic semantic extraction
Each entity's collected sentences are segmented into individual, checkable claims, everything from "What traits characterize it?" to "How does it relate to other entities?" Each claim traces back to the exact source sentence it came from.
2
Structural compression
Everything known about an entity gets compressed into a compact profile related to the entity's position within the extracted network. A follow-up step then reconciles vocabulary across different sources, so two write-ups describing the same thing in different words end up sharing one term instead of splitting into two.
3
Archetype formation
The model looks across every entity's profile at once and groups the ones that share real structural patterns, while also flagging exactly how each member departs from those same patterns. Crystal calls these groups archetypes — recurring structural patterns, not statistical clusters, that are analogous to how a crystal's geometry determines its form but not other details like size or composition.
4
Cross-source confirmation
A fully deterministic check, not a model call: how many independent sources support each pattern, which entities bridge more than one, and which patterns are too thin or single-sourced to trust. This is what makes "established across 2+ sources" a real, checked filter rather than a claim.
5
Surface structural insights
A domain-specific question run against that shared topology. What's a potential knowledge gap across fields? Which entity breaks the pattern everything else follows? Tailor the question for the insight you want to surface.

pass 5 — fit for domain

The same discipline across domains: keep the deterministic legwork separate from the LLM steps, and constrain whichever calls are LLM-driven to work only from verified, structured inputs — never room to invent a new claim.

Science — PubMed-verified gaps

1. Topology query (deterministic) — filter the archetype graph by a research question, e.g. "gene targets not yet studied in songbirds for research in vocal communication."

2. Gap verification (deterministic) — targeted PubMed searches; a null result is kept as a citable fact, not an assumption.

3. Report synthesis (LLM) — one call per candidate, given only the topology position, evidence strings, and PubMed result. It packages verified claims — it doesn't generate new ones.

Fiction — structural inversion

1. Candidate filter (deterministic) — archetypes confirmed across 2+ sources, with a member whose profile diverges from the shared pattern.

2. Inversion scoring (LLM, temp 0.1 — minimal randomness) — judges genuine inversion vs. minor variant on one specific shared axis, grounded only in the evidence given.

3. Readable explanation (LLM) — translates an already-correct technical finding into plain language. Translation, not judgment.

the discipline

The model explains. It doesn't decide.

In Pass 5, every claim is checked before the model ever sees it — PubMed searches run, evidence traced, thresholds applied. The one LLM call per finding comes last, and its only job is turning an already-verified result into plain language.

Verified inputs only

Every research report carries the same notice: gap claims come from PubMed searches run at generation time; topology claims trace to specific papers via evidence reference IDs. The one LLM call in the whole pass receives only that verified, structured material — it packages claims, it doesn't generate them.

No forced comparisons

If two entities were never actually assessed on the same trait, Crystal doesn't compare them on it by matching similar-sounding labels. That comparison is simply left out — not silently faked to look complete.

real output, trimmed
SRPX2 AND songbird  →  0 results (PubMed, searched 2026-07-14)
SRPX2 AND "zebra finch"  →  0 results (PubMed, searched 2026-07-14)
evidence_ref: Konopka, G., & Roberts, T. F. (2016). Animal Models of Speech and Vocal Communication Deficits Associated With Psychiatric Disorders. Biological Psychiatry.
Its relevance to vocal learning stems directly from its position as a transcriptional target of FOXP2. Searches of PubMed for "SRPX2 AND vocal learning," "SRPX2 AND songbird," and "SRPX2 AND zebra finch" each return zero results as of July 14, 2026, confirming that this gene has not been studied in any avian vocal learner.

results

Real Pass 5 output, from two wildly different domains.

From scientific literature

Query: "genes, proteins, or other discrete molecular targets mechanistically linked to vocal learning circuits, not yet studied in songbirds." Below are the candidates where PubMed verification actually came back thin or empty; not every candidate the pass surfaced is shown here.

SRPX2

A FOXP2 transcriptional target that promotes synapse formation; rodent knockdown reduces both synapse density and ultrasonic vocalization. PubMed: zero results for "vocal learning," "songbird," or "zebra finch."

Candidate signal, not a validated result — flagged for follow-up, not a conclusion.
SLIT1

FOXP2-regulated axon guidance ligand; human FOXP2 drives stronger SLIT1 upregulation than chimp FOXP2, and vocal-learner species show convergent SLIT1 down-regulation in motor regions. PubMed: 1–2 results across all three searches — no functional songbird study published.

Candidate signal, not a validated result — flagged for follow-up, not a conclusion.
LHX9

A regional-identity transcription factor specifically downregulated in the developmental tissue that becomes RA, the song system's premotor output nucleus. PubMed: zero results for "vocal learning," 2 each for "songbird" and "zebra finch."

Candidate signal, not a validated result — flagged for follow-up, not a conclusion.

From complex, fictional worlds

Structurally notable characters.

Same approach; a different question: which entities read as the outlier in two or more independently-confirmed archetypes at once.

Zidane (Final Fantasy IX)

Genuine inversion in 2 archetypes. The protector he'd been for most of the story briefly becomes the protected once he learns his origin — and that revelation produces collapse before resilience, rather than the self-grounding such revelations usually bring.

Structural pattern, not a claim about authorial intent — flagged by deviation from the shared pattern.
Tidus (Final Fantasy X)

Genuine inversion in 3 archetypes at once. His existence is contingent on the perpetuation of an external dream, so completing his protective mission, resolving his conflict with his father, and any "chosen" sacrifice are all the same predetermined event — not an earned outcome.

Structural pattern, not a claim about authorial intent — flagged by deviation from the shared pattern.

see it run

The interactive graphs, live.

View a fully explorable graph — zoom into a cluster, click a node for its evidence trail, see all extracted characteristics and relationships. Two domains, same pipeline.