Architecture¶
This document describes the major components of rainbow-fmt and the data
flowing between them. It is a design target, not a description of existing
code.
Pipeline¶
source text
│
▼
┌──────────────┐ ┌───────────────┐
│ 1. Parsing │◀──│ Language pack │ grammar, injections, rules, options
└──────┬───────┘ └───────┬───────┘
│ CST + trivia │
▼ │
┌──────────────┐ │
│ 2. Injection │ split embedded regions, recurse into (1) per language
└──────┬───────┘ │
│ tree of CSTs │
▼ ▼
┌──────────────┐ ┌───────────────┐
│ 3. Rule │◀──│ 5. Config │ resolved options for this file/node
│ engine │ │ resolver │
└──────┬───────┘ └───────────────┘
│ Doc IR
▼
┌──────────────┐
│ 4. Printer │ language-agnostic layout (width, indent, line breaks)
└──────┬───────┘
│ text
▼
┌──────────────┐
│ 6. Verifier │ re-parse, compare trees, check idempotency
└──────┬───────┘
▼
formatted text ──▶ 7. Front-ends: CLI · library API · LSP · pre-commit
1. Parsing layer¶
Responsibility: turn source text into a concrete syntax tree that keeps every token and the exact byte range of every piece of trivia (whitespace, comments).
- Implemented: the
Parser/CstNodeprotocols and tree helpers inrainbow_fmt.core.parser, the tree-sitter adapter inrainbow_fmt.core.treesitter(ADR 0003). - Primary backend: tree-sitter. Mature grammars exist for every initial target (HTML, CSS, SCSS, JavaScript, TypeScript/TSX, Svelte, Python, JSON, TOML, YAML, Markdown), and new grammars can be generated from a JavaScript grammar DSL. Its query language also gives us injections and pattern matching for free.
- Pluggable backends. A
Parserinterface allows a language pack to use a different parser (e.g. Python'sast/tokenize, a hand-written parser for a tiny DSL, or a PEG grammar) as long as it produces the common CST node interface. - Trivia attachment. Comments are attached to the nearest node (leading,
trailing, or dangling) using a deterministic algorithm, so rules can move
nodes without losing comments. This is historically the hardest part of
every formatter and gets its own test suite. Implemented for the
members of a node (
rainbow_fmt.core.comments.attach_comments, TASKS.md T13): a comment starting on the row where a member or its separator ends trails it, any other comment leads the next member or dangles after the last; blank lines before each member are recorded. It uses rows only, so it works for anyCstNodeimplementation. The verifier checks that no comment is lost or reordered. - Error tolerance. Files with syntax errors are either left untouched
(default) or formatted outside the error region (
on_error = "partial").
2. Injection layer (embedded languages)¶
Responsibility: find regions of one language inside another and format them with the correct language pack.
- Injections are declared in the language pack (tree-sitter
injections.scmqueries, or an equivalent declarative form):<script lang="ts">→ TS,<style lang="scss">→ SCSS, Markdown fences → by info string, Python strings tagged# rainbow: lang=sql→ SQL. - The inner region is formatted independently, with its base indentation and available width supplied by the host, then spliced back as an opaque Doc fragment.
- Template interpolations (Svelte
{expr},{#if}blocks, Jinja{{ }}) are handled by placeholder substitution: the host masks interpolations, formats, then restores them — each interpolation is itself formatted as an expression in the inner language.
3. Rule engine¶
Responsibility: translate a CST into Doc IR by applying the language pack's rules, parameterized by the resolved options.
- A rule = a selector (today a node type; later optionally a
tree-sitter query or a CSS-like path such as
call_expression > arguments) + a function producing Doc IR from the node, its children and the resolved options. - Language packs write rules as Python functions (
(node, ctx) -> Doc); users override a single rule with a[[rule]]expression template that is checked on load and cannot run code (ADR 0005,extending.md). - Specificity and overriding. A user
[[rule]]replaces the pack's rule for the same selector; with richer selectors, more specific rules will win, as in CSS. preservesupport. Every option-driven decision has a "look at the source" branch: e.g.blank_lines_between_members = preserveconsults the trivia recorded in the CST;line_breaks = preserveemits a hard line wherever the source had one.- Default rule. A node with no rule is printed verbatim (source slice), with only indentation re-based. This guarantees that an incomplete language pack is safe: unknown constructs are never mangled.
- Deep nesting. Rules recurse once per nesting level, so a pack runs
them with
rainbow_fmt.core.stack.call_with_large_stack(a worker thread with a 64 MiB stack and a recursion limit of 100,000; TASKS.md T12). JSON formats more than 10,000 levels; beyond the limit,formatraisesNestingTooDeepError. The parser, the printer and the verifier are iterative.
4. Printer (layout engine)¶
Responsibility: turn Doc IR into text that fits the configured width.
- Based on the Wadler/Leijen "prettier printer" algebra, as used by Prettier
and many others:
text,line,softline,hardline,group,indent,align,fill,if_break,line_suffix, plus rainbow additions: table— align cells across consecutive lines (assignment alignment, dict/object value alignment, CSS property alignment), a common request that opinionated formatters refuse.verbatim— emit source text untouched (forpreserveandrainbow: offregions).- Completely language-agnostic; it knows nothing about syntax.
- Configurable:
max_width,indent_style(spaces/tabs),indent_size,tab_width(for width computation),newline, Unicode width handling. - Implemented in
rainbow_fmt.core.docandrainbow_fmt.core.printer; reference:doc-ir.md.
5. Configuration system¶
Responsibility: determine the effective value of every option for a given file and, where relevant, a given node.
Detailed in configuration.md. In brief:
- Typed option schemas (declared by core and by each language pack), with documentation and defaults generated from the schema.
- Format-independent model; each file format is a thin source that parses
into it and records positions. TOML is primary; YAML and
package.jsonare supported (ADR 0001). - Cascading sources: built-in defaults → preset(s) →
.editorconfig→ project config (rainbow.toml/rainbow.yaml/[tool.rainbow]inpyproject.toml/"rainbow"inpackage.json) → path-glob overrides → inline directives. - Every resolved value records its provenance (which file and line set
it), which powers
rainbow-fmt explain.
6. Verifier¶
Responsibility: guarantee that formatting never changes meaning and is stable.
- Implemented (
rainbow_fmt.verify, TASKS.md T11):verify(language, source, output, options)checks that the output parses, that the trees without comments are equal (node types, shape, token text, and the own text of nodes with children, such as CSS100%), that the comments are the same and in the same order, and that a second pass changes nothing.Language.optional_tokens(CSS:;) are ignored, and so areLanguage.optional_trailingtokens directly before a closing bracket (Python: trailing commas),Language.optional_lastandoptional_children(JavaScript: semicolons and empty statements). Token texts are compared with line endings normalized; comments also without whitespace at the start and end of their lines (a pack may re-indent the lines of a block comment). The CLI verifies every file before writing it (--no-verify,[files] verify = false). - Equivalence check: re-parse the output; the CST with trivia stripped must equal the input's. Planned: language packs declare normalizations that are allowed to differ (quote style, optional trailing commas, optional parentheses, semicolons in JS); JSON needs none.
- Idempotency check:
format(format(x)) == format(x). - On by default; skipped on request for speed (editors).
7. Front-ends¶
- CLI —
rainbow-fmt format,check,diff,explain,infer,languages,options. Parallel file processing and a content-hash cache. - Library API —
rainbow_fmt.format_text(text, language=..., config=...). - LSP server — document and range formatting for any editor.
- Integrations — pre-commit hook, GitHub Action, VS Code extension (thin LSP client).
8. Style inference (rainbow-fmt infer)¶
Responsibility: read an existing codebase and emit a config that reproduces its dominant style, so adopting rainbow produces a minimal diff.
- For each option, format a sample of files under each candidate value and
pick the value that minimizes the diff against the originals (or report
that the codebase is inconsistent and suggest
preserve). - This is a direct payoff of the "every decision is an option" design.
Package layout (proposed)¶
rainbow-fmt/
├── src/rainbow_fmt/
│ ├── core/ # CST node interface, trivia, Doc IR, printer
│ ├── rules/ # rule types, CST helpers, override templates
│ ├── config/ # schema, sources, cascade, provenance
│ ├── inject/ # injection detection and splicing
│ ├── verify/ # equivalence + idempotency
│ ├── cli/ # command-line interface
│ ├── lsp/ # language server
│ └── languages/ # built-in language packs (one sub-package each)
│ ├── json/
│ ├── css/ scss/
│ ├── python/
│ ├── javascript/ typescript/
│ └── html/ svelte/
├── tests/
│ ├── core/ config/ rules/ ...
│ └── fixtures/<language>/<case>/{input, options.toml, expected}
└── docs/
Third-party language packs are separate distributions discovered via the
rainbow_fmt.languages entry-point group.