Benchmark¶
python -m rainbow_fmt.benchmark formats and verifies a generated corpus
and prints the best time of several runs per case, in seconds
(ADR 0002, decision 3):
$ python -m rainbow_fmt.benchmark
case size format verify
large 143104 0.238 0.921
comments 20799 0.026 0.044
deep 2001 0.352 0.389
css 95537 0.322 1.099
python 107625 0.578 1.403
javascript 110504 0.871 1.743
typescript 95368 0.560 1.206
html 101468 0.195 0.553
svelte 99045 0.841 2.752
calibration 0.019 s
| Case | Contents |
|---|---|
large |
139 KiB of JSON, 1,000 objects with 6,000 key/value pairs |
comments |
JSONC, 1,000 members each with a trailing comment |
deep |
JSON nested 1,000 levels deep (indent_size = 0) |
css |
93 KiB of CSS: 1,250 rules with selector lists, 25 @media blocks, comments |
python |
105 KiB of Python: 120 classes with 480 methods, 120 functions, comprehensions, comments |
javascript |
108 KiB of JavaScript without semicolons: 110 classes with fields, 330 async methods, arrow functions, member chains, template literals, comments |
typescript |
93 KiB of TypeScript: 110 generic interfaces, union type aliases, enums, classes with parameter properties and overloads, generic arrow functions |
html |
100 KiB of HTML: 240 sections with headings, filled text, links, entities, lists and inputs, a <style> and a <script>, comments |
svelte |
97 KiB of Svelte: a script, 300 {#if} blocks with {#each}/{:else} inside, expressions in content and attributes, components, a <style> |
The svelte case is slow because tree-sitter-svelte 1.0.2 parses in time
quadratic in the file size (0.5 s for 100 KiB, 0.04 s for 25 KiB), and a
file is parsed four times (format, verification of input and output, the
stability pass); components of ordinary size are not affected (2,386 real
components: 18 s with four workers).
The corpus is generated deterministically (rainbow_fmt.benchmark.CORPUS),
so no large files are checked in. Each Case names the language that
formats and verifies it (Case(name, generate, options, language); JSON by
default).
Compared with Black and Prettier¶
The same corpus files, written to disk and checked from the command line
(check, best of three runs, process start-up included; Windows, four
CPUs; Black 26.5 compiled, Prettier 3.9 on Node, rainbow-fmt 0.4.0.dev0
on CPython 3.12). Each tool does a different amount of work: rainbow-fmt
verifies its output by default (a re-parse and a second formatting pass,
--no-verify skips it), Black re-parses to compare ASTs, Prettier does
not check its output.
| Case | Size | rainbow-fmt | rainbow-fmt --no-verify |
Black | Prettier |
|---|---|---|---|---|---|
large (JSON) |
139 KiB | 1.28 s | 0.48 s | 0.24 s | |
css |
93 KiB | 1.44 s | 0.51 s | 0.41 s | |
python |
105 KiB | 1.73 s | 0.71 s | 1.91 s | |
javascript |
107 KiB | 2.29 s | 0.90 s | 0.34 s | |
typescript |
93 KiB | 1.83 s | 0.73 s | 0.38 s | |
html |
99 KiB | 0.97 s | 0.48 s | 0.30 s | |
svelte |
96 KiB | 3.88 s | 1.15 s | 1.00 s |
Start-up alone (--version): rainbow-fmt 0.26 s, Black 0.18 s, Prettier
0.10 s. On directories (one run each): a slice of the standard library
(86 files) takes rainbow-fmt 2.1 s with four workers (7.8 s with one) and
Black 2.3 s; the 1,143-page HTML corpus takes rainbow-fmt 71 s and
Prettier 41 s.
So, as of Phase 4: formatting alone is on a par with Black and about twice as slow as Prettier per file; verification doubles the time again. The verifier is the place to look first (TASKS.md T23): it re-parses the input that formatting already parsed, and the stability pass formats the whole file a second time.
Options¶
| Option | Meaning |
|---|---|
--repeat N |
runs per case; the best is reported (default 5) |
--output FILE |
write the results as JSON |
--baseline FILE |
compare with a results file; exit 1 on a regression |
Comparing with a baseline¶
Machines differ in speed, so every time is divided by the calibration: the time of a fixed pure-Python workload measured in the same run. A case is a regression when its normalized format or verify time is more than twice the baseline's:
large: format 2.10x the baseline (tolerance 2.00x)
CI runs python -m rainbow_fmt.benchmark --baseline
benchmarks/baseline.json --output benchmark.json (job benchmark) and
keeps benchmark.json as an artifact. After an intended change in speed,
replace benchmarks/baseline.json with a new results file, preferably the
artifact of a CI run.