ADR 0003 — Parser backend¶
- Status: Accepted (amended 2026-09-26, see below)
- Date: 2026-09-26
- Decides: D2 in
roadmap.md
Context¶
Every language needs a parser that produces a concrete syntax tree (CST) with every token and the exact byte range of all trivia (whitespace and comments). The parser layer also has to support:
- many languages, including HTML, CSS, SCSS, JavaScript, TypeScript/TSX, Svelte, Python, and user-supplied DSLs;
- embedded languages (
<script>,<style>, Svelte blocks, Markdown fences); - a way for a new language to be added without writing a parser by hand;
- use from Python (see ADR 0002).
Decision¶
- tree-sitter is the default parser backend, used through the official
tree-sitterPython bindings. Grammars are consumed as the published per-language wheels (tree-sitter-python,tree-sitter-javascript, …) where available, or as compiled grammars shipped by a language pack. - The core depends on a
Parserprotocol, not on tree-sitter directly. The protocol produces the common CST node interface (type, children, byte range, point range, field names,is_error/is_missing). The tree-sitter adapter is one implementation; a language pack may provide another (e.g. a hand-written parser for a very small DSL, or a PEG parser) as long as it satisfies the same interface. - Trivia is reconstructed from byte ranges. tree-sitter does not keep whitespace in the tree; the adapter derives each trivia run from the gaps between token byte ranges, so the original source is always exactly recoverable from the CST.
- Injections use tree-sitter's query language (
injections.scm), following the convention used by editors (Neovim, Helix, Zed), so existing injection queries can be reused. - Grammar versions are pinned per language pack. A grammar upgrade is a pack release that must pass the pack's full fixture suite, since node types and field names can change between grammar versions.
Consequences¶
- Maintained grammars exist for every Phase 1–4 language; parsing is fast (C) and incremental, which also benefits the LSP server.
- tree-sitter is error-tolerant: files with syntax errors still parse, with
ERROR/MISSINGnodes. The core must check for these and apply theon_errorpolicy (default: leave the file untouched). - Grammars are written for highlighting and editing, not formatting. Some
produce trees that are coarser or shaped differently than a formatter
wants (e.g. unparsed text nodes inside template languages). Language packs
may need small grammar forks or post-parse tree adjustments; the
Parserprotocol (decision 2) is the escape hatch when a grammar is unsuitable. - Grammar wheels add compiled dependencies per language; language packs declare them, so users install only the grammars they need.
- New DSLs can be supported by writing a tree-sitter grammar (a JavaScript grammar DSL, compiled to C) — no rainbow-specific parser framework is needed.
Alternatives considered¶
- Per-language native parsers (Python
ast/tokenize, Babel/SWC, PostCSS, …). Each is the most accurate parser for its language, but they span several host languages, expose incompatible tree shapes, and many discard trivia. Supporting them all would mean one adapter per language and would make new DSLs expensive. - A rainbow-specific parser framework (e.g. PEG-based). Full control over tree shape, but it would mean writing and maintaining grammars for every mainstream language ourselves.
- Mixed by default (native parsers for mainstream languages, tree-sitter
for the rest). Rejected as the default because it forfeits a single CST
model; it remains possible per language through the
Parserprotocol.
Amendment 1 — 2026-09-26 (implementation, TASKS.md T7)¶
Decided while implementing the adapter (rainbow_fmt.core.treesitter):
is_extrais part of theCstNodeprotocol. It marks trivia that may appear anywhere, such as comments, so comment handling can be written once for all languages. Comments are leaves; only whitespace lives in the gaps between leaves.- Error nodes are never
is_extra. tree-sitter reportsERRORnodes as extras; the adapter reportsis_extra = Falsefor them, so rules can rely onis_extrameaning trivia. - A leading UTF-8 byte order mark is the pipeline's job. tree-sitter
skips it, so it is not part of any node. The format pipeline strips it
before parsing and restores it after printing. A characterization test
(
test_leading_utf8_bom_is_skipped_by_tree_sitter) catches a change in grammar behaviour. - Generic tree helpers
iter_nodes,iter_leavesandhas_errorslive next to the protocols (rainbow_fmt.core.parser), work for anyParserimplementation, and are iterative. - Wrappers are created lazily (children are wrapped on first access and cached) and hold a reference to their tree, so a node stays valid for as long as it is referenced.