Skip to content

Benchmark

python -m rainbow_fmt.benchmark formats and verifies a generated corpus and prints the best time of several runs per case, in seconds (ADR 0002, decision 3):

$ python -m rainbow_fmt.benchmark
case           size   format   verify
large        143104    0.238    0.921
comments      20799    0.026    0.044
deep           2001    0.352    0.389
css           95537    0.322    1.099
python       107625    0.578    1.403
javascript   110504    0.871    1.743
typescript    95368    0.560    1.206
html         101468    0.195    0.553
svelte        99045    0.841    2.752
calibration 0.019 s
Case Contents
large 139 KiB of JSON, 1,000 objects with 6,000 key/value pairs
comments JSONC, 1,000 members each with a trailing comment
deep JSON nested 1,000 levels deep (indent_size = 0)
css 93 KiB of CSS: 1,250 rules with selector lists, 25 @media blocks, comments
python 105 KiB of Python: 120 classes with 480 methods, 120 functions, comprehensions, comments
javascript 108 KiB of JavaScript without semicolons: 110 classes with fields, 330 async methods, arrow functions, member chains, template literals, comments
typescript 93 KiB of TypeScript: 110 generic interfaces, union type aliases, enums, classes with parameter properties and overloads, generic arrow functions
html 100 KiB of HTML: 240 sections with headings, filled text, links, entities, lists and inputs, a <style> and a <script>, comments
svelte 97 KiB of Svelte: a script, 300 {#if} blocks with {#each}/{:else} inside, expressions in content and attributes, components, a <style>

The svelte case is slow because tree-sitter-svelte 1.0.2 parses in time quadratic in the file size (0.5 s for 100 KiB, 0.04 s for 25 KiB), and a file is parsed four times (format, verification of input and output, the stability pass); components of ordinary size are not affected (2,386 real components: 18 s with four workers).

The corpus is generated deterministically (rainbow_fmt.benchmark.CORPUS), so no large files are checked in. Each Case names the language that formats and verifies it (Case(name, generate, options, language); JSON by default).

Compared with Black and Prettier

The same corpus files, written to disk and checked from the command line (check, best of three runs, process start-up included; Windows, four CPUs; Black 26.5 compiled, Prettier 3.9 on Node, rainbow-fmt 0.4.0.dev0 on CPython 3.12). Each tool does a different amount of work: rainbow-fmt verifies its output by default (a re-parse and a second formatting pass, --no-verify skips it), Black re-parses to compare ASTs, Prettier does not check its output.

Case Size rainbow-fmt rainbow-fmt --no-verify Black Prettier
large (JSON) 139 KiB 1.28 s 0.48 s 0.24 s
css 93 KiB 1.44 s 0.51 s 0.41 s
python 105 KiB 1.73 s 0.71 s 1.91 s
javascript 107 KiB 2.29 s 0.90 s 0.34 s
typescript 93 KiB 1.83 s 0.73 s 0.38 s
html 99 KiB 0.97 s 0.48 s 0.30 s
svelte 96 KiB 3.88 s 1.15 s 1.00 s

Start-up alone (--version): rainbow-fmt 0.26 s, Black 0.18 s, Prettier 0.10 s. On directories (one run each): a slice of the standard library (86 files) takes rainbow-fmt 2.1 s with four workers (7.8 s with one) and Black 2.3 s; the 1,143-page HTML corpus takes rainbow-fmt 71 s and Prettier 41 s.

So, as of Phase 4: formatting alone is on a par with Black and about twice as slow as Prettier per file; verification doubles the time again. The verifier is the place to look first (TASKS.md T23): it re-parses the input that formatting already parsed, and the stability pass formats the whole file a second time.

Options

Option Meaning
--repeat N runs per case; the best is reported (default 5)
--output FILE write the results as JSON
--baseline FILE compare with a results file; exit 1 on a regression

Comparing with a baseline

Machines differ in speed, so every time is divided by the calibration: the time of a fixed pure-Python workload measured in the same run. A case is a regression when its normalized format or verify time is more than twice the baseline's:

large: format 2.10x the baseline (tolerance 2.00x)

CI runs python -m rainbow_fmt.benchmark --baseline benchmarks/baseline.json --output benchmark.json (job benchmark) and keeps benchmark.json as an artifact. After an intended change in speed, replace benchmarks/baseline.json with a new results file, preferably the artifact of a CI run.