Skip to main content

Tier Pipeline

Differens parses files at six levels of abstraction, T0 through T5. Each tier is a parser that produces a Node tree. Each tier can fall back to the tier below it. The tier number is the abstraction ladder: T0 is bytes, T5 is code.

The tiers

T0: binary (hash-only)

Byte-level hash comparison. Used for images, fonts, archives, and anything else that fails the text sniff. Output is a single Update when the hash changes, with the byte-size delta as context. T0 never attempts to parse. Format plugins can upgrade T0 per file type: register a BinaryDiffPlugin for an extension and it answers or declines, with the byte-delta report as the default. Image perceptual diff, EXIF diff, ELF symbol diff are all plugin-shaped, none built in.

T1: raw (line diff)

The safety net: a line-level diff over the raw text. It is the fallback for every tier above it. Small inputs get an exact longest-common-subsequence diff; large inputs get a linear-space Myers diff, so a 100,000-line file with a one-line edit reports exactly that one edit. There is no size cap. A file that shares no lines with the other side at all reports as a single rewrite.

T2: prose (word-level)

Word-level diff for plain text and logs. Paragraphs and sentences become tree nodes, so a moved paragraph is a Move instead of a delete-plus-insert. Punctuation and whitespace changes update leaf nodes rather than rewriting whole paragraphs.

T3: markup (lenient HTML/XML)

A lenient HTML/XML tokenizer that tolerates malformed documents: unclosed tags, misnested elements, unescaped entities. Element hierarchy becomes the node tree. Attribute order changes are reorders, not updates.

T4: data (JSON / YAML / TOML)

Value trees for JSON, the YAML subset Differens supports, and the TOML subset. The diff is key-path based: nested objects become nested nodes, so a key moved between two objects reports as a Move with the key path in the label. Array reordering is a Reorder.

T5: code (tree-sitter)

A tree-sitter CST with per-language semantic extractors. The extractors map raw tree-sitter node types to canonical concepts such as Function, Class, Method, and Import. The narration can then say “function” whether the source said fn, def, or func. All sixteen registered languages have extractors: TypeScript/JavaScript, Python, Rust, Go, C, C++, Java, Ruby, PHP, Swift, Kotlin, C#, Scala, Lua, and shell. Languages without an extractor still work at the generic level: structural diff with raw tree-sitter node type labels. The extractor only adds human-readable concept names.

Graceful degradation

The tiers form a chain. Each tier can hand off to the tier below it:
  1. T5 parse fails (syntax error, unsupported construct) → tree-sitter produces ERROR nodes and the structural diff still runs.
  2. No grammar registered (extension has no grammar, or the grammar could not load) → T1 line diff.
The worst case for text is a line diff. It is less informative than a structural diff, but it still runs, at any size.

Classification rules

The content router classifies by extension, then by content sniff: Extension-based classification is the first cut; the content sniff is the backstop. A .txt file that is actually minified JSON still gets a T2 word-level diff. That is acceptable, because the tier below can take over.

Tier summary

Adding a tier adapter: create the adapter in packages/tiers/src/, produce a Node tree (the interface lives in packages/core/src/index.ts), register it in the diffWithTier() switch in packages/tiers/src/index.ts, and add extension rules to classifyFile(). The tier number should reflect its position on the abstraction ladder.