Extending rainbow-fmt: new languages and DSLs¶
Adding a language should be a matter of hours for a simple DSL and days for a mainstream language — not the months it takes to write a printer from scratch.
Anatomy of a language pack¶
rainbow-lang-foo/
├── pyproject.toml # entry point: rainbow_fmt.languages = foo = "rainbow_lang_foo"
└── rainbow_lang_foo/
├── language.toml # metadata: name, file extensions, comment syntax, parser
├── options.py # option declarations (see Options below)
├── format.py # rules: node type → function returning a Doc
├── injections.scm # optional: embedded-language queries
├── normalize.toml # optional: allowed differences for the verifier
└── fixtures/ # golden tests: input + options → expected
Levels of support¶
A pack can start small and grow; each level is useful on its own.
| Level | Provides | Result |
|---|---|---|
| 0 — Registration | language.toml + grammar |
Re-indentation, whitespace/EOL normalization, rainbow: off; everything else verbatim |
| 1 — Structure | Rules for blocks / statements | Consistent indentation and blank-line handling |
| 2 — Layout | Rules for expressions, lists, calls | Width-aware line breaking |
| 3 — Style | Language options (quotes, commas, braces …) | Full configurability |
| 4 — Embedding | injections.scm |
Formats nested languages |
Because unmatched nodes are printed verbatim, a Level 0 pack is already safe to run on real code.
Rules¶
Rules are Python functions (ADR 0005). A rule
receives a CST node and a context, and returns a Doc
(doc-ir.md); the pack maps node types to rules. The JSON
pack (rainbow_fmt/languages/json/format.py) is the reference:
from rainbow_fmt.core.doc import Doc, concat
from rainbow_fmt.core.parser import CstNode
from rainbow_fmt.rules import Rule, field_node, has_comments, source_text
def pair(node: CstNode, ctx: Context) -> Doc:
if has_comments(node): # comment between key and value: keep as written
return source_text(node, ctx)
return concat(ctx.doc(field_node(node, "key")), ": ", ctx.doc(field_node(node, "value")))
RULES: dict[str, Rule[Context]] = {"pair": pair, "object": object_, ...}
ctx.doc(child)formats a child with its own rule;ctx.text(node)is the source text; the pack's context also carries its resolved options.- Nodes without a rule are printed verbatim (
source_text), so a pack is safe before it is complete. - Rules recurse once per nesting level. Run the root rule with
rainbow_fmt.core.stack.call_with_large_stack(ctx.doc, root), and turn aRecursionErrorintorainbow_fmt.languages.NestingTooDeepError, asformat_jsondoes; otherwise input nested a few hundred levels deep exceeds Python's default recursion limit. - Shared helpers (
field_node,has_comments,source_text) live inrainbow_fmt.rules. - Comments are children with
is_extraset. For a node with members (a list, a block),rainbow_fmt.core.comments.attach_comments(children, separators=(",",))returns each member with itsleadingandtrailingcomments andblank_lines_before, plus thedanglingcomments after the last member; the JSON pack'scontainershows how to print them (leading comments on their own lines, trailing ones withline_suffix, a container with comments always broken).
Options¶
A pack declares its options on its Language; the resolver
(rainbow_fmt.options.resolve_options) validates them, applies the
configuration levels and records provenance. The JSON pack:
JSON_OPTIONS = (
Option("object_wrap", Choice("preserve", "fit", "always"), "preserve",
"When objects break: as written, only when too long, or always."),
Option("align_values", Boolean(), False,
"Align the values of a broken object whose values are all scalars."),
)
LANGUAGE = Language("json", (".json", ".jsonc"), format_json, JSON_OPTIONS, PARSER)
format(source, options, strict=False) receives a ResolvedOptions
mapping every core, shared and pack option to its value (and the user's
[[rule]] templates in options.rules). Pack option names must be unique
and must not repeat a core or shared option, except to narrow it: a
Choice with a subset of its values (the JSON pack narrows
trailing_comma to "never"). A narrowed option belongs to the pack;
[core] and [shared] values do not apply to it.
defaults changes only the default of core or shared options for the
pack (the Python pack: {"trailing_comma": "multiline"}); [core] and
[shared] values still apply. Unknown names, narrowed names and invalid
values are rejected when the Language is created.
parser (optional) is the pack's Parser; the verifier
(rainbow_fmt.verify) re-parses the output with it to check that the tree
and comments are unchanged. Without a parser, only stability
(format(format(x)) == format(x)) is verified. optional_tokens names
token types the formatter may add or remove, which the verifier ignores
(the CSS pack: ;, for last_semicolon). optional_trailing maps a node
type to a token type that may be added or removed only directly before
the node's closing bracket (the Python pack: , in argument_list,
list and others, but not subscript, where a[1,] differs from
a[1]), and only after a named node ([,] has one element). optional_last
maps a node type to a token type that may be added or removed as its last
child, and optional_children a node type to child types that may be
added or removed anywhere among its children (the JavaScript pack: a
statement's final ;, and empty statements in a block).
rainbow_fmt.languages.pack.format_with_rules(language, source, options,
rules, helpers, make_context, strict=...) does what every tree-sitter pack
needs: it resolves the options, adds the user's [[rule]] templates, keeps
the byte order mark, handles syntax errors and deep nesting, and prints
with the core options; the JSON, CSS, Python and JavaScript packs are three-line wrappers
around it.
User overrides¶
A user replaces one rule from their configuration without forking the pack:
[[rule]]
language = "json"
select = "pair" # node type
template = "[field('key'), ' : ', field('value')]"
A template is a restricted Python expression: names, string and number
literals, lists (concatenated), calls of named functions without keyword
arguments, and a if condition else b. It can use:
| Name | Meaning |
|---|---|
concat, group, indent, join, if_break, text |
Doc builders |
line, softline, hardline |
line breaks |
children |
the node's children, formatted |
has_comments |
whether a comment is a direct child |
field(name) |
the child in grammar field name, formatted |
source() |
the node's source text, unchanged |
| pack helpers | exported by the pack; JSON: container(open, close) |
Templates are checked when the configuration is loaded: attribute access,
subscripts, operators, comprehensions and lambdas are rejected with the
rule's position, e.g. rule[0] (json/pair): unsupported syntax 'a.b' at
column 1. A pack documents the helpers it exports as part of its API.
Embedding other languages¶
A pack whose files contain code of other languages (HTML with <style>
and <script>, Svelte with {expressions}) formats such a region by
splicing the embedded pack's Doc into its own, so the one printer lays
out the whole file: widths, indentation and verbatim lines inside the
region come out right without re-indenting text
(rainbow_fmt.inject, TASKS.md T24).
- The pack names the languages it embeds:
Language(..., embeds={"css": CSS, "javascript": JAVASCRIPT}). Their options are resolved for the host file and configuration (options.embedded["css"]:[language.css],[core],.editorconfig,--setapply as they would to a CSS file at that path) and are part of the cache key. - A rule gets the region's Doc with
embedded_doc(ctx.options, CSS, region_bytes)and prints it, for example asindent(hardline, doc);Nonemeans the region does not parse, and the rule prints it as written.wrap=("(", ")")parses the region between the two strings and gives the Doc of the node inside (a Svelte{expr}parsed as a JavaScript expression, so that{ a }is an object, not a block). Language.injection(node, source)returns anInjection(child, "css", wrap)for a node holding a region, so that the verifier compares the region by the CSS pack's rules (tree and comments; with whitespace collapsed when it does not parse) rather than as one token.Language.collapsednames token types the verifier compares with whitespace runs collapsed (HTMLtext, which the pack re-wraps).
Every built-in pack provides Language.doc(source, options, wrap=...)
(rainbow_fmt.languages.pack.doc_with_rules), so any of them can be
embedded.
Registering a pack¶
A pack is a Python distribution that names its Language in the
rainbow_fmt.languages entry-point group:
# pyproject.toml of rainbow-lang-scheme
[project.entry-points."rainbow_fmt.languages"]
scheme = "rainbow_lang_scheme:LANGUAGE"
Once installed (pip install -e .), rainbow-fmt formats the pack's
extensions, [language.scheme] is valid configuration, and the pack's
options show up in rainbow-fmt options. Built-in packs are consulted
first, so a plugin cannot take over .py or .css. A pack that fails to
import, or whose entry point is not a Language, stops rainbow-fmt
with an error naming the entry point. A worked example, with a
hand-written parser, is in
howto/define-language-module.
Bringing a grammar¶
- tree-sitter grammar exists: reference it (compiled wheel such as
tree-sitter-python, or a path to a built grammar). - No grammar: write one with tree-sitter's grammar DSL, or implement the
Parserinterface directly for very small DSLs.
Testing a pack¶
rainbow-fmt test-pack path/to/pack runs:
- Golden fixtures — every
fixtures/<case>/input.*formatted withoptions.tomlmust equalexpected.*. - Safety — the verifier's equivalence and idempotency checks on every fixture and on an optional corpus of real-world files.
- Option matrix — each option is exercised across all its values on the corpus, checking safety (not output) for every combination sampled.