Skip to content

Extending rainbow-fmt: new languages and DSLs

Adding a language should be a matter of hours for a simple DSL and days for a mainstream language — not the months it takes to write a printer from scratch.

Anatomy of a language pack

rainbow-lang-foo/
├── pyproject.toml          # entry point: rainbow_fmt.languages = foo = "rainbow_lang_foo"
└── rainbow_lang_foo/
    ├── language.toml       # metadata: name, file extensions, comment syntax, parser
    ├── options.py          # option declarations (see Options below)
    ├── format.py           # rules: node type → function returning a Doc
    ├── injections.scm      # optional: embedded-language queries
    ├── normalize.toml      # optional: allowed differences for the verifier
    └── fixtures/           # golden tests: input + options → expected

Levels of support

A pack can start small and grow; each level is useful on its own.

Level Provides Result
0 — Registration language.toml + grammar Re-indentation, whitespace/EOL normalization, rainbow: off; everything else verbatim
1 — Structure Rules for blocks / statements Consistent indentation and blank-line handling
2 — Layout Rules for expressions, lists, calls Width-aware line breaking
3 — Style Language options (quotes, commas, braces …) Full configurability
4 — Embedding injections.scm Formats nested languages

Because unmatched nodes are printed verbatim, a Level 0 pack is already safe to run on real code.

Rules

Rules are Python functions (ADR 0005). A rule receives a CST node and a context, and returns a Doc (doc-ir.md); the pack maps node types to rules. The JSON pack (rainbow_fmt/languages/json/format.py) is the reference:

from rainbow_fmt.core.doc import Doc, concat
from rainbow_fmt.core.parser import CstNode
from rainbow_fmt.rules import Rule, field_node, has_comments, source_text


def pair(node: CstNode, ctx: Context) -> Doc:
    if has_comments(node):  # comment between key and value: keep as written
        return source_text(node, ctx)
    return concat(ctx.doc(field_node(node, "key")), ": ", ctx.doc(field_node(node, "value")))


RULES: dict[str, Rule[Context]] = {"pair": pair, "object": object_, ...}
  • ctx.doc(child) formats a child with its own rule; ctx.text(node) is the source text; the pack's context also carries its resolved options.
  • Nodes without a rule are printed verbatim (source_text), so a pack is safe before it is complete.
  • Rules recurse once per nesting level. Run the root rule with rainbow_fmt.core.stack.call_with_large_stack(ctx.doc, root), and turn a RecursionError into rainbow_fmt.languages.NestingTooDeepError, as format_json does; otherwise input nested a few hundred levels deep exceeds Python's default recursion limit.
  • Shared helpers (field_node, has_comments, source_text) live in rainbow_fmt.rules.
  • Comments are children with is_extra set. For a node with members (a list, a block), rainbow_fmt.core.comments.attach_comments(children, separators=(",",)) returns each member with its leading and trailing comments and blank_lines_before, plus the dangling comments after the last member; the JSON pack's container shows how to print them (leading comments on their own lines, trailing ones with line_suffix, a container with comments always broken).

Options

A pack declares its options on its Language; the resolver (rainbow_fmt.options.resolve_options) validates them, applies the configuration levels and records provenance. The JSON pack:

JSON_OPTIONS = (
    Option("object_wrap", Choice("preserve", "fit", "always"), "preserve",
           "When objects break: as written, only when too long, or always."),
    Option("align_values", Boolean(), False,
           "Align the values of a broken object whose values are all scalars."),
)
LANGUAGE = Language("json", (".json", ".jsonc"), format_json, JSON_OPTIONS, PARSER)

format(source, options, strict=False) receives a ResolvedOptions mapping every core, shared and pack option to its value (and the user's [[rule]] templates in options.rules). Pack option names must be unique and must not repeat a core or shared option, except to narrow it: a Choice with a subset of its values (the JSON pack narrows trailing_comma to "never"). A narrowed option belongs to the pack; [core] and [shared] values do not apply to it. defaults changes only the default of core or shared options for the pack (the Python pack: {"trailing_comma": "multiline"}); [core] and [shared] values still apply. Unknown names, narrowed names and invalid values are rejected when the Language is created.

parser (optional) is the pack's Parser; the verifier (rainbow_fmt.verify) re-parses the output with it to check that the tree and comments are unchanged. Without a parser, only stability (format(format(x)) == format(x)) is verified. optional_tokens names token types the formatter may add or remove, which the verifier ignores (the CSS pack: ;, for last_semicolon). optional_trailing maps a node type to a token type that may be added or removed only directly before the node's closing bracket (the Python pack: , in argument_list, list and others, but not subscript, where a[1,] differs from a[1]), and only after a named node ([,] has one element). optional_last maps a node type to a token type that may be added or removed as its last child, and optional_children a node type to child types that may be added or removed anywhere among its children (the JavaScript pack: a statement's final ;, and empty statements in a block).

rainbow_fmt.languages.pack.format_with_rules(language, source, options, rules, helpers, make_context, strict=...) does what every tree-sitter pack needs: it resolves the options, adds the user's [[rule]] templates, keeps the byte order mark, handles syntax errors and deep nesting, and prints with the core options; the JSON, CSS, Python and JavaScript packs are three-line wrappers around it.

User overrides

A user replaces one rule from their configuration without forking the pack:

[[rule]]
language = "json"
select   = "pair"                                  # node type
template = "[field('key'), ' : ', field('value')]"

A template is a restricted Python expression: names, string and number literals, lists (concatenated), calls of named functions without keyword arguments, and a if condition else b. It can use:

Name Meaning
concat, group, indent, join, if_break, text Doc builders
line, softline, hardline line breaks
children the node's children, formatted
has_comments whether a comment is a direct child
field(name) the child in grammar field name, formatted
source() the node's source text, unchanged
pack helpers exported by the pack; JSON: container(open, close)

Templates are checked when the configuration is loaded: attribute access, subscripts, operators, comprehensions and lambdas are rejected with the rule's position, e.g. rule[0] (json/pair): unsupported syntax 'a.b' at column 1. A pack documents the helpers it exports as part of its API.

Embedding other languages

A pack whose files contain code of other languages (HTML with <style> and <script>, Svelte with {expressions}) formats such a region by splicing the embedded pack's Doc into its own, so the one printer lays out the whole file: widths, indentation and verbatim lines inside the region come out right without re-indenting text (rainbow_fmt.inject, TASKS.md T24).

  • The pack names the languages it embeds: Language(..., embeds={"css": CSS, "javascript": JAVASCRIPT}). Their options are resolved for the host file and configuration (options.embedded["css"]: [language.css], [core], .editorconfig, --set apply as they would to a CSS file at that path) and are part of the cache key.
  • A rule gets the region's Doc with embedded_doc(ctx.options, CSS, region_bytes) and prints it, for example as indent(hardline, doc); None means the region does not parse, and the rule prints it as written. wrap=("(", ")") parses the region between the two strings and gives the Doc of the node inside (a Svelte {expr} parsed as a JavaScript expression, so that { a } is an object, not a block).
  • Language.injection(node, source) returns an Injection(child, "css", wrap) for a node holding a region, so that the verifier compares the region by the CSS pack's rules (tree and comments; with whitespace collapsed when it does not parse) rather than as one token.
  • Language.collapsed names token types the verifier compares with whitespace runs collapsed (HTML text, which the pack re-wraps).

Every built-in pack provides Language.doc(source, options, wrap=...) (rainbow_fmt.languages.pack.doc_with_rules), so any of them can be embedded.

Registering a pack

A pack is a Python distribution that names its Language in the rainbow_fmt.languages entry-point group:

# pyproject.toml of rainbow-lang-scheme
[project.entry-points."rainbow_fmt.languages"]
scheme = "rainbow_lang_scheme:LANGUAGE"

Once installed (pip install -e .), rainbow-fmt formats the pack's extensions, [language.scheme] is valid configuration, and the pack's options show up in rainbow-fmt options. Built-in packs are consulted first, so a plugin cannot take over .py or .css. A pack that fails to import, or whose entry point is not a Language, stops rainbow-fmt with an error naming the entry point. A worked example, with a hand-written parser, is in howto/define-language-module.

Bringing a grammar

  • tree-sitter grammar exists: reference it (compiled wheel such as tree-sitter-python, or a path to a built grammar).
  • No grammar: write one with tree-sitter's grammar DSL, or implement the Parser interface directly for very small DSLs.

Testing a pack

rainbow-fmt test-pack path/to/pack runs:

  1. Golden fixtures — every fixtures/<case>/input.* formatted with options.toml must equal expected.*.
  2. Safety — the verifier's equivalence and idempotency checks on every fixture and on an optional corpus of real-world files.
  3. Option matrix — each option is exercised across all its values on the corpus, checking safety (not output) for every combination sampled.