Skip to content

Architecture

This document describes the major components of rainbow-fmt and the data flowing between them. It is a design target, not a description of existing code.

Pipeline

 source text
     │
     ▼
┌──────────────┐   ┌───────────────┐
│ 1. Parsing   │◀──│ Language pack │  grammar, injections, rules, options
└──────┬───────┘   └───────┬───────┘
       │ CST + trivia      │
       ▼                   │
┌──────────────┐           │
│ 2. Injection │  split embedded regions, recurse into (1) per language
└──────┬───────┘           │
       │ tree of CSTs      │
       ▼                   ▼
┌──────────────┐   ┌───────────────┐
│ 3. Rule      │◀──│ 5. Config     │  resolved options for this file/node
│    engine    │   │    resolver   │
└──────┬───────┘   └───────────────┘
       │ Doc IR
       ▼
┌──────────────┐
│ 4. Printer   │  language-agnostic layout (width, indent, line breaks)
└──────┬───────┘
       │ text
       ▼
┌──────────────┐
│ 6. Verifier  │  re-parse, compare trees, check idempotency
└──────┬───────┘
       ▼
 formatted text  ──▶  7. Front-ends: CLI · library API · LSP · pre-commit

1. Parsing layer

Responsibility: turn source text into a concrete syntax tree that keeps every token and the exact byte range of every piece of trivia (whitespace, comments).

  • Implemented: the Parser/CstNode protocols and tree helpers in rainbow_fmt.core.parser, the tree-sitter adapter in rainbow_fmt.core.treesitter (ADR 0003).
  • Primary backend: tree-sitter. Mature grammars exist for every initial target (HTML, CSS, SCSS, JavaScript, TypeScript/TSX, Svelte, Python, JSON, TOML, YAML, Markdown), and new grammars can be generated from a JavaScript grammar DSL. Its query language also gives us injections and pattern matching for free.
  • Pluggable backends. A Parser interface allows a language pack to use a different parser (e.g. Python's ast/tokenize, a hand-written parser for a tiny DSL, or a PEG grammar) as long as it produces the common CST node interface.
  • Trivia attachment. Comments are attached to the nearest node (leading, trailing, or dangling) using a deterministic algorithm, so rules can move nodes without losing comments. This is historically the hardest part of every formatter and gets its own test suite. Implemented for the members of a node (rainbow_fmt.core.comments.attach_comments, TASKS.md T13): a comment starting on the row where a member or its separator ends trails it, any other comment leads the next member or dangles after the last; blank lines before each member are recorded. It uses rows only, so it works for any CstNode implementation. The verifier checks that no comment is lost or reordered.
  • Error tolerance. Files with syntax errors are either left untouched (default) or formatted outside the error region (on_error = "partial").

2. Injection layer (embedded languages)

Responsibility: find regions of one language inside another and format them with the correct language pack.

  • Injections are declared in the language pack (tree-sitter injections.scm queries, or an equivalent declarative form): <script lang="ts"> → TS, <style lang="scss"> → SCSS, Markdown fences → by info string, Python strings tagged # rainbow: lang=sql → SQL.
  • The inner region is formatted independently, with its base indentation and available width supplied by the host, then spliced back as an opaque Doc fragment.
  • Template interpolations (Svelte {expr}, {#if} blocks, Jinja {{ }}) are handled by placeholder substitution: the host masks interpolations, formats, then restores them — each interpolation is itself formatted as an expression in the inner language.

3. Rule engine

Responsibility: translate a CST into Doc IR by applying the language pack's rules, parameterized by the resolved options.

  • A rule = a selector (today a node type; later optionally a tree-sitter query or a CSS-like path such as call_expression > arguments) + a function producing Doc IR from the node, its children and the resolved options.
  • Language packs write rules as Python functions ((node, ctx) -> Doc); users override a single rule with a [[rule]] expression template that is checked on load and cannot run code (ADR 0005, extending.md).
  • Specificity and overriding. A user [[rule]] replaces the pack's rule for the same selector; with richer selectors, more specific rules will win, as in CSS.
  • preserve support. Every option-driven decision has a "look at the source" branch: e.g. blank_lines_between_members = preserve consults the trivia recorded in the CST; line_breaks = preserve emits a hard line wherever the source had one.
  • Default rule. A node with no rule is printed verbatim (source slice), with only indentation re-based. This guarantees that an incomplete language pack is safe: unknown constructs are never mangled.
  • Deep nesting. Rules recurse once per nesting level, so a pack runs them with rainbow_fmt.core.stack.call_with_large_stack (a worker thread with a 64 MiB stack and a recursion limit of 100,000; TASKS.md T12). JSON formats more than 10,000 levels; beyond the limit, format raises NestingTooDeepError. The parser, the printer and the verifier are iterative.

4. Printer (layout engine)

Responsibility: turn Doc IR into text that fits the configured width.

  • Based on the Wadler/Leijen "prettier printer" algebra, as used by Prettier and many others: text, line, softline, hardline, group, indent, align, fill, if_break, line_suffix, plus rainbow additions:
  • table — align cells across consecutive lines (assignment alignment, dict/object value alignment, CSS property alignment), a common request that opinionated formatters refuse.
  • verbatim — emit source text untouched (for preserve and rainbow: off regions).
  • Completely language-agnostic; it knows nothing about syntax.
  • Configurable: max_width, indent_style (spaces/tabs), indent_size, tab_width (for width computation), newline, Unicode width handling.
  • Implemented in rainbow_fmt.core.doc and rainbow_fmt.core.printer; reference: doc-ir.md.

5. Configuration system

Responsibility: determine the effective value of every option for a given file and, where relevant, a given node.

Detailed in configuration.md. In brief:

  • Typed option schemas (declared by core and by each language pack), with documentation and defaults generated from the schema.
  • Format-independent model; each file format is a thin source that parses into it and records positions. TOML is primary; YAML and package.json are supported (ADR 0001).
  • Cascading sources: built-in defaults → preset(s) → .editorconfig → project config (rainbow.toml / rainbow.yaml / [tool.rainbow] in pyproject.toml / "rainbow" in package.json) → path-glob overrides → inline directives.
  • Every resolved value records its provenance (which file and line set it), which powers rainbow-fmt explain.

6. Verifier

Responsibility: guarantee that formatting never changes meaning and is stable.

  • Implemented (rainbow_fmt.verify, TASKS.md T11): verify(language, source, output, options) checks that the output parses, that the trees without comments are equal (node types, shape, token text, and the own text of nodes with children, such as CSS 100%), that the comments are the same and in the same order, and that a second pass changes nothing. Language.optional_tokens (CSS: ;) are ignored, and so are Language.optional_trailing tokens directly before a closing bracket (Python: trailing commas), Language.optional_last and optional_children (JavaScript: semicolons and empty statements). Token texts are compared with line endings normalized; comments also without whitespace at the start and end of their lines (a pack may re-indent the lines of a block comment). The CLI verifies every file before writing it (--no-verify, [files] verify = false).
  • Equivalence check: re-parse the output; the CST with trivia stripped must equal the input's. Planned: language packs declare normalizations that are allowed to differ (quote style, optional trailing commas, optional parentheses, semicolons in JS); JSON needs none.
  • Idempotency check: format(format(x)) == format(x).
  • On by default; skipped on request for speed (editors).

7. Front-ends

  • CLI — rainbow-fmt format, check, diff, explain, infer, languages, options. Parallel file processing and a content-hash cache.
  • Library API — rainbow_fmt.format_text(text, language=..., config=...).
  • LSP server — document and range formatting for any editor.
  • Integrations — pre-commit hook, GitHub Action, VS Code extension (thin LSP client).

8. Style inference (rainbow-fmt infer)

Responsibility: read an existing codebase and emit a config that reproduces its dominant style, so adopting rainbow produces a minimal diff.

  • For each option, format a sample of files under each candidate value and pick the value that minimizes the diff against the originals (or report that the codebase is inconsistent and suggest preserve).
  • This is a direct payoff of the "every decision is an option" design.

Package layout (proposed)

rainbow-fmt/
├── src/rainbow_fmt/
│   ├── core/          # CST node interface, trivia, Doc IR, printer
│   ├── rules/         # rule types, CST helpers, override templates
│   ├── config/        # schema, sources, cascade, provenance
│   ├── inject/        # injection detection and splicing
│   ├── verify/        # equivalence + idempotency
│   ├── cli/           # command-line interface
│   ├── lsp/           # language server
│   └── languages/     # built-in language packs (one sub-package each)
│       ├── json/
│       ├── css/  scss/
│       ├── python/
│       ├── javascript/  typescript/
│       └── html/  svelte/
├── tests/
│   ├── core/  config/  rules/ ...
│   └── fixtures/<language>/<case>/{input, options.toml, expected}
└── docs/

Third-party language packs are separate distributions discovered via the rainbow_fmt.languages entry-point group.