Skip to main content

Cross-File Correlation

A line diff cannot tell you that a function was deleted from utils.ts and recreated in helpers.ts. It shows a delete here and an insert there. The cross-file correlator runs after the per-pair diff, over the whole set of changed files. Its job is to determine whether code moved between files.

Stage 1: structure hash buckets

Every node from every old file and every new file is bucketed by structure hash. Two nodes in the same bucket have identical shape: same kind, same children, regardless of labels or values. Identical subtrees that appear in different files are exact moves. A function body moved to another file with no edits is found here.

Stage 2: content hash exact matches

Structure buckets say shapes match. Content hashes say code matches. Within a structure bucket, nodes are compared by content hash: kind, label, value, and children. A function moved unchanged to another file matches on both hashes. The correlator emits a Move with source and destination file paths instead of a Delete plus an Insert. Jaccard similarity catches renames at a distance (stage 3).

Stage 3: Jaccard similarity for modified moves

The hard case is a function that moved and changed: new parameter, rewritten body, renamed. No hash matches. The correlator falls back to token-level Jaccard similarity between the old and new node’s token sets:
The pair with the highest similarity above threshold is reported as the move’s origin, provided no better match exists elsewhere. This is deliberately conservative. A low-similarity pairing is not reported, because a wrong move attribution is worse than none.
Jaccard is only applied within the same structure family. A function candidate is never compared against a class or an import. The structure bucket from stage 1 constrains the search space before similarity scoring runs.

The edit action

Correlated moves surface as a single Move action carrying both paths: The narration engine renders that as “moved computeTotal from src/utils.ts to src/helpers.ts: one sentence instead of two unrelated hunks.

Current limitations

Leaf-level renames still report as a Delete plus an Insert. A single renamed identifier has no subtree to match on, so the correlator cannot pair it with confidence. The narration convention handles the cosmetic case: when a delete and an insert are the only changes in their respective parents, the narration renders “renamed X to Y” even though the underlying actions are separate. Exact-match granularity below the node level is on the roadmap.