Skip to content

ADR 0003 — Parser backend

  • Status: Accepted (amended 2026-09-26, see below)
  • Date: 2026-09-26
  • Decides: D2 in roadmap.md

Context

Every language needs a parser that produces a concrete syntax tree (CST) with every token and the exact byte range of all trivia (whitespace and comments). The parser layer also has to support:

  • many languages, including HTML, CSS, SCSS, JavaScript, TypeScript/TSX, Svelte, Python, and user-supplied DSLs;
  • embedded languages (<script>, <style>, Svelte blocks, Markdown fences);
  • a way for a new language to be added without writing a parser by hand;
  • use from Python (see ADR 0002).

Decision

  1. tree-sitter is the default parser backend, used through the official tree-sitter Python bindings. Grammars are consumed as the published per-language wheels (tree-sitter-python, tree-sitter-javascript, …) where available, or as compiled grammars shipped by a language pack.
  2. The core depends on a Parser protocol, not on tree-sitter directly. The protocol produces the common CST node interface (type, children, byte range, point range, field names, is_error/is_missing). The tree-sitter adapter is one implementation; a language pack may provide another (e.g. a hand-written parser for a very small DSL, or a PEG parser) as long as it satisfies the same interface.
  3. Trivia is reconstructed from byte ranges. tree-sitter does not keep whitespace in the tree; the adapter derives each trivia run from the gaps between token byte ranges, so the original source is always exactly recoverable from the CST.
  4. Injections use tree-sitter's query language (injections.scm), following the convention used by editors (Neovim, Helix, Zed), so existing injection queries can be reused.
  5. Grammar versions are pinned per language pack. A grammar upgrade is a pack release that must pass the pack's full fixture suite, since node types and field names can change between grammar versions.

Consequences

  • Maintained grammars exist for every Phase 1–4 language; parsing is fast (C) and incremental, which also benefits the LSP server.
  • tree-sitter is error-tolerant: files with syntax errors still parse, with ERROR/MISSING nodes. The core must check for these and apply the on_error policy (default: leave the file untouched).
  • Grammars are written for highlighting and editing, not formatting. Some produce trees that are coarser or shaped differently than a formatter wants (e.g. unparsed text nodes inside template languages). Language packs may need small grammar forks or post-parse tree adjustments; the Parser protocol (decision 2) is the escape hatch when a grammar is unsuitable.
  • Grammar wheels add compiled dependencies per language; language packs declare them, so users install only the grammars they need.
  • New DSLs can be supported by writing a tree-sitter grammar (a JavaScript grammar DSL, compiled to C) — no rainbow-specific parser framework is needed.

Alternatives considered

  • Per-language native parsers (Python ast/tokenize, Babel/SWC, PostCSS, …). Each is the most accurate parser for its language, but they span several host languages, expose incompatible tree shapes, and many discard trivia. Supporting them all would mean one adapter per language and would make new DSLs expensive.
  • A rainbow-specific parser framework (e.g. PEG-based). Full control over tree shape, but it would mean writing and maintaining grammars for every mainstream language ourselves.
  • Mixed by default (native parsers for mainstream languages, tree-sitter for the rest). Rejected as the default because it forfeits a single CST model; it remains possible per language through the Parser protocol.

Amendment 1 — 2026-09-26 (implementation, TASKS.md T7)

Decided while implementing the adapter (rainbow_fmt.core.treesitter):

  1. is_extra is part of the CstNode protocol. It marks trivia that may appear anywhere, such as comments, so comment handling can be written once for all languages. Comments are leaves; only whitespace lives in the gaps between leaves.
  2. Error nodes are never is_extra. tree-sitter reports ERROR nodes as extras; the adapter reports is_extra = False for them, so rules can rely on is_extra meaning trivia.
  3. A leading UTF-8 byte order mark is the pipeline's job. tree-sitter skips it, so it is not part of any node. The format pipeline strips it before parsing and restores it after printing. A characterization test (test_leading_utf8_bom_is_skipped_by_tree_sitter) catches a change in grammar behaviour.
  4. Generic tree helpers iter_nodes, iter_leaves and has_errors live next to the protocols (rainbow_fmt.core.parser), work for any Parser implementation, and are iterative.
  5. Wrappers are created lazily (children are wrapped on first access and cached) and hold a reference to their tree, so a node stays valid for as long as it is referenced.