Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Badness

Badness is a language server, formatter, and linter for LaTeX.

It parses LaTeX into a lossless concrete syntax tree and builds three tools on top of it:

  • a formatter (badness format) that lays out source deterministically,
  • a linter (badness lint) that reports diagnostics, and
  • a language server (badness lsp) that brings both to your editor, as well as many other features.

The architecture follows rust-analyzer: a generic, error-tolerant, hand-written parser produces a lossless tree, semantics are layered on top as a separate concern, and recomputation is incremental.

The Architecture

Badness treats input as generic TeX surface syntax. It never requires resolving macros or catcodes to succeed—doing that in full generality is equivalent to running a TeX engine, so anything it cannot statically recognize degrades to generic nodes rather than a crash. Two properties are guaranteed by construction and enforced as test oracles:

  • Losslessness: the parsed tree reconstructs the input byte-for-byte.
  • Idempotence: formatting an already-formatted file changes nothing.

Where to Go Next

Funding

I am grateful for the financial support offered by the TeX Users Group through the TeX Development Fund.

Installation

Badness is distributed as a single binary, badness. The current version is 0.27.0. It is available from several sources:

  • crates.io: cargo install badness
  • Homebrew: brew install jolars/tap/badness
  • npm: npm install -g badness (bundles a prebuilt binary)
  • PyPI: uv tool install badness/pipx install badness
  • AUR (Arch Linux): yay -S badness-bin (prebuilt binary)
  • Prebuilt binaries: from the releases page
  • VS Code/Open VSX: the Badness extension (also on Open VSX; works in Positron and Cursor)

The editor extension bundles a platform-specific badness binary and starts the language server automatically, so no separate CLI install is required. See Editor Setup for configuration.

From Source

Badness is written in Rust. With a Rust toolchain installed, build from a checkout:

git clone https://github.com/jolars/badness
cd badness
cargo build --release

The binary lands at target/release/badness. Copy it onto your PATH, or run it in place.

To install it into Cargo’s bin directory instead:

cargo install --path .

Verifying the Install

badness --version

This should print badness 0.27.0.

Getting Started

Badness’s main subcommands are format, lint, and lsp (with parse and init as helpers). This page walks through formatting and linting from the command line. For editor integration, see Editor Setup.

Formatting a File

Format a file in place:

badness format paper.tex

Pass several paths to format them all:

badness format intro.tex methods.tex results.tex

Pass - to read from standard input and write the formatted result to standard output—handy for piping or editor integrations:

cat paper.tex | badness format -

A piped standard input is also read when you pass no paths at all, so the shorter cat paper.tex | badness format works too. At an interactive prompt, though, where there is nothing to pipe, badness format with no paths reports a usage error rather than silently waiting on the terminal.

Checking Without Writing

In CI you usually want to verify that files are already formatted rather than rewrite them. The --check flag prints a diff of what would change and exits non-zero if any file is not already formatted:

badness format --check paper.tex
Diff in paper.tex:12:
 \section{Introduction}
-Some    text with   odd spacing.
+Some text with odd spacing.
1 of 1 file(s) would be reformatted

Since --check writes nothing, that report is the only account of what would change, which is why it is shown by default. Pass --quiet for just the file list and the summary—useful when a first run over an unformatted project would otherwise flood a CI log:

badness format --check --quiet .

The report goes to stdout (only errors use stderr) and is colorized when writing to a terminal; --color always|never overrides that, and NO_COLOR is honored.

Linting

lint parses each file and reports any diagnostics found, rendered with source snippets. It exits non-zero when there is at least one diagnostic:

badness lint paper.tex

Like format, it reads standard input when given - (or when piped with no paths):

cat paper.tex | badness lint -

The snippets go to stderr and are colorized when it is a terminal, under the same --color always|never and NO_COLOR rules as the --check diff. The --output concise and --output json forms are meant for other programs to read, so they stay plain whatever --color says.

Adjusting Layout

The formatter takes a few style options on the command line:

badness format --line-width 100 --indent-width 4 --wrap preserve paper.tex

See the CLI Reference for every flag and the Configuration reference for what --wrap controls.

Formatting

badness format lays out LaTeX source deterministically. Output is decided solely by the formatter’s rules and its layout engine—there are no per-construct special cases to memorize.

In Place, stdin, or check

badness format paper.tex          # rewrite the file in place
cat paper.tex | badness format    # stdin → stdout
badness format --check paper.tex  # diff, don't write; non-zero if unformatted

--check prints a diff of the pending change for each file, then a summary; add --quiet to reduce that to the file list and the summary. See Checking without writing.

Style Options

The style flags—including --line-width, --indent-width, --item-indent, and --wrap—mirror the [format] section of badness.toml and override it for a single run. Each option’s default and meaning is listed in the Configuration reference.

For persistent settings, badness reads a badness.toml discovered from the working directory upward; pass --config <PATH> to point at a specific file or --no-config to ignore any discovered one. Run badness init to write a starter badness.toml.

Under reflow, textual optional arguments wrap at existing top-level comma-space boundaries and fill each line to the configured width. Continuation lines are indented, and brackets remain attached wherever adding a space would change the argument. Known key-value arguments instead expand to one entry per line when they do not fit. Opaque braced environment arguments keep values such as {section in head/foot} together instead of wrapping their words. Existing comments retain their binding; the formatter does not insert % markers to create new break opportunities.

For a literal top-level \documentclass{cas-sc} or \documentclass{cas-dc}, badness treats the braced affiliation fields and the trailing address options of \affiliation as key-value lists. Short lists stay inline; longer lists expand to one entry per line. Other classes retain ordinary argument formatting, and definitions in the document or a loaded local package override the class signature.

In align and other math alignment grids, including nested aligned environments, a leading \label{key} sits on its own line. Formula rows align without counting the label toward column widths. Trailing comments stay attached to their labels, and labels within formula rows stay in place. This layout applies independently of math-wrap.

Under math-wrap = "break" (the default unless wrap = "preserve"), multiline sibling environments in display math form separate blocks, even when a joined line would fit. Intervening expressions sit on their own lines, and punctuation stays attached to the preceding block. A prefix such as A = can still introduce the first environment on the same line. Explicit preserve and single-line math modes retain their existing line-break policies.

Inside algorithm and algorithm2e environments (including their starred forms), badness normalizes text in \KwIn, \KwOut, \KwData, and \KwResult. It indents the braced bodies of \For, \ForEach, \ForAll, \While, \If, \ElseIf, \Else, \eIf, and \Repeat, placing each statement ending in \; on its own source line. Math spacing commands remain inside their formulas, and trailing comments stay attached to their statements. Control-flow commands need their complete braced arguments to receive this layout; custom commands and forms with parenthesized side comments use the ordinary fallback. Top-level \; statements require a recognized control-flow call in the environment’s direct body, since the separate algorithm package uses the same float name and can use \; for ordinary spacing.

Turning the formatter off

Sometimes a block is laid out by hand and should stay that way—a tikzpicture aligned by eye, a table whose columns line up in the source. Comment directives turn the formatter off over exactly as much as you point at, and content inside is reproduced byte for byte.

Skip the next construct:

% badness-format skip: hand-aligned by eye
\begin{tikzpicture}
  \foreach \p/\pos in {A/left, B/left, C/right, D/right}%
  \node[\pos] at (\p) {$\p$};%
\end{tikzpicture}

Skip a region:

% badness-format off
\begin{tabular}{ll}
  a   &   b \\
  ccc &   d \\
\end{tabular}
% badness-format on

Skip a whole file, wherever in it the directive sits:

% badness-format skip-file: generated, do not edit

An off with no matching on runs to the end of the file. The : <reason> is optional everywhere and is never interpreted—it is there for the next person to read.

Each directive has a bare counterpart that turns off both the formatter and every lint rule over the same span: % badness skip, % badness off / % badness on, and % badness skip-file. Use the -format spelling when you want the linter to keep reporting.

To exclude whole files by path instead, use exclude/extend-exclude in badness.toml; see the Configuration reference. That is the better tool when you control the config, since it keeps the directive out of the document.

The mirror of these is % badness-lint, which suppresses diagnostics without touching layout and takes an optional rule name; see Linting.

One note on where directives are read: a directive must be its own % comment. In a .dtx documentation line the leading % is a documentation margin rather than a comment, so a directive written there is inert (inside a macrocode chunk it works normally).

Guarantees

The formatter is built around a small set of invariants that double as test oracles:

  • Idempotence: format(format(x)) == format(x).
  • Losslessness: the parsed tree reconstructs the input byte-for-byte, so the formatter never loses or corrupts content.
  • Protected regions: verbatim-like content (verbatim, lstlisting, \verb, comments) is never altered. An environment badness cannot tell is verbatim—one built by machinery no scan follows—can be named in [environments], which is also how you teach it a \bea/\eea pair defined in a sibling .sty.
  • Whitespace-only: formatting changes whitespace, line breaks, and comment placement, and nothing else. It never inserts, deletes, or rewrites a token of real content.

Content normalizations—rewriting x^{2} to x^2, or $$…$$ to \[…\]—are therefore lint fixes, not formatting. Run badness lint --fix for those.

Linting

badness lint parses each file and reports diagnostics, rendered with source snippets pointing at the offending range. It exits non-zero when there is at least one diagnostic, which makes it usable as a CI gate.

badness lint paper.tex
cat paper.tex | badness lint   # stdin

Parse diagnostics

Alongside the rules, the linter surfaces parse diagnostics: places where the parser recovered from malformed input. Because the parser is error-tolerant, a single problem never aborts the parse—badness anchors recovery on clean LaTeX boundaries (\end{…}, \begin, a blank line, }, $, &, \\) and keeps going, so one file can report several independent diagnostics in one run. Parse diagnostics carry the rule id parse and are never silenced by select/ignore.

Rules

Beyond parse recovery, badness ships a growing set of built-in rules (deprecated-command, dollar-display-math, undefined-ref, and more). Each has a stable id used in diagnostics, config, and suppression comments. See the Linter Rules reference for the full catalogue, or print a single rule’s description and examples from the terminal:

badness lint --explain deprecated-command

Every rule is on by default. Narrow the active set through the [lint] table in badness.toml or the matching --select/--ignore CLI flags; see the Configuration reference.

Suppress a rule at one site with a comment directive:

% badness-lint skip deprecated-command: legacy code
{\bf here}

The verb carries the scope, and there are three:

ScopeDirective
The next construct% badness-lint skip <rule>: <reason>
A region% badness-lint off <rule> … % badness-lint on <rule>
The whole file% badness-lint skip-file <rule>: <reason>

Naming the <rule> is optional—leave it out and the directive covers every rule over that same span. An off with no matching on runs to the end of the file. The : <reason> tail is optional everywhere and is never interpreted.

Each has a bare counterpart that turns off the formatter at the same time: % badness skip, % badness off / % badness on, and % badness skip-file. For layout only, use the % badness-format spellings described in Formatting.

In .bib files the same grammar rides an @comment entry, since BibTeX has no line-comment token:

@comment{badness-lint skip missing-required-field: publisher long gone}
@book{oldbook, title = {An Orphaned Book}}

The inert-suppression rule warns when a directive cannot act—for example, a dangling skip, an unmatched on, an unclosed off, a directive written as typeset prose on a .dtx documentation line, or a format-only directive in a .bib file.

Some rules ship an auto-fix. badness lint --fix applies the meaning-preserving (Safe) ones; --unsafe-fixes also applies fixes that may change output, such as missing-nonbreaking-space (inserting a tie changes line breaking), abbreviation-spacing (inserting \ or \@ changes sentence spacing), or space-before-command (deleting a space before \footnote changes spacing).

Machine-readable output

badness lint --output json emits the findings as a JSON array on stdout (the human-readable pretty and concise modes write to stderr). A clean run emits [], so consumers always receive valid JSON; the exit code still signals whether findings exist. This is the contract external tools consume, e.g. panache when linting latex code blocks in Markdown documents.

[
  {
    "rule": "ellipsis",
    "severity": "warning",
    "path": "paper.tex",
    "start": 5,
    "end": 8,
    "message": "literal `...` ellipsis; use `\\dots`",
    "fix": {
      "edits": [{ "content": "\\dots", "start": 5, "end": 8 }],
      "applicability": "safe",
      "description": "Replace `...` with `\\dots`"
    },
    "related": []
  }
]

Ranges are 0-indexed byte offsets into the named file (no line/column resolution). severity is one of error, warning, info, or hint; applicability is safe or unsafe (the --fix/--unsafe-fixes split). The fix key is omitted when a finding has no auto-fix. An edit carries a path key only when it targets a different file than the diagnostic (a cross-file fix); related lists secondary “see also” locations.

Compared to the sibling tools arity and fatou, the schema differs in two ways: offsets are flat start/end keys rather than a range object, and message is a plain string rather than a structured object.

Editor Setup

Badness ships a language server. Start it with:

badness lsp

The server speaks the Language Server Protocol over stdio. Point your editor’s LSP client at the badness binary with the lsp argument and associate it with LaTeX (.tex) and BibTeX (.bib) files.

Settings can be supplied as initializationOptions at startup or through workspace/didChangeConfiguration, either as a bare object or namespaced under a badness key.

Formatter widths: lineWidth and indentWidth. They act as a fallback: a discovered badness.toml always wins outright, and absent one, your editor’s tab size (sent with each formatting request) overrides the indent width.

The language server is also the sole consumer of the [build] section of badness.toml, which locates the compile’s .aux artifacts; see the Configuration reference.

Command completion

Command suggestions include short signatures, such as \section[]{} and \vspace{}, before you select an item. Clients that support completion label details can display the argument suffix beside the command name. Other clients receive the full signature in the completion item’s detail field. Full documentation loads when the client resolves the selected item.

Signatures use the document’s definitions, loaded local packages, and Badness’s built-in data. They display the known brace and bracket argument slots; they do not describe every TeX argument protocol. The signature display does not insert arguments.

Known math and text symbols and logos, such as \omega, \hbar, \copyright, and \LaTeX, have the completion kind Constant. Known argument-free control, spacing, and declaration commands, such as \newpage, \par, \quad, and \bfseries, have kind Keyword. Argument-taking commands such as \vspace retain Function. Editors can use these distinctions for icons and automatic brackets; for example, blink.cmp can insert braces after \vspace while leaving \omega and \newpage bare. Recognized definitions in the document or loaded local packages, and explicit project declarations, override the built-in classification. Commands without a curated classification keep their existing completion kinds; an empty signature alone does not establish zero arguments.

LaTeX3 completion

Badness completes expl3 functions, variables, and constants inside \ExplSyntaxOn regions, after \ProvidesExplPackage, \ProvidesExplClass, or \ProvidesExplFile, and inside recognized expl3 macrocode regions in .dtx files. For example, \tl_ offers \tl_set:Nn, and \l_tmpa_ offers scratch variables. The built-in catalog ships with Badness and needs no TeX installation.

Completion also includes literal definitions in the current file and loaded local packages and classes, including \cs_new:Npn functions, variable and constant declarations, conditional forms, and generated variants. Badness does not expand macros to discover computed names. It skips incomplete definitions and definitions stored as token-list data. These names support completion; expl3 definition navigation and argument signature help are not yet provided.

Renaming source files

Invoke your editor’s LSP rename command inside a literal source path, such as \input{chapters/introduction}. Badness renames the file and updates its references across the discovered workspace. This also works with \include, \subfile, \subfileinclude, \import, \subimport, \loadglsentries, and the parent path in \documentclass[...]{subfiles}. Literal \includeonly lists are updated along with their targets.

The new name uses the same base directory as the original argument. For example, renaming chapters/introduction to appendix moves the file to appendix.tex beside the referring document. Use chapters/appendix to keep it in the same directory. Import commands use their directory argument as the base. Badness preserves the file extension when you omit it and keeps each reference’s extension spelling when it still resolves correctly. Renaming an imported file preserves the import directory argument unless that directory itself moves.

Moves stay within the same workspace root and never overwrite existing files or destination buffers. New parent directories are allowed when the editor creates them while applying the file operation; Neovim supports this. Without workspace folders, Badness uses the initiating document’s directory as the boundary.

File explorers can also rename source files and folders through workspace/willRenameFiles and workspace/didRenameFiles. Your explorer must send these requests and notifications. When files move, Badness adjusts their recognized references, including references to assets inside moved folders. Cursor rename includes its reference edits even when explorer hooks are enabled. Unsaved editor buffers take precedence over disk contents. Each referring file uses its own project’s declarations and exclusions, including nested projects and other workspace folders.

Badness declines moves across directories when a moved source contains relative file arguments. Their resolution can depend on the compilation directory or an import context, which cannot be inferred from the source’s location alone. It also checks the compilation and import directories inherited through literal source loads. If a reference requires different edits in those contexts, Badness declines the rename, even when the referring file stays in place.

File rename requires a client that supports LSP resource-rename operations. Dynamic paths, braceless inputs, \graphicspath, and symlink aliases are not resolved for rename. Source and destination paths cannot pass through symlinks inside the workspace, and new names cannot contain quotation marks or TeX delimiters. Badness also declines names containing spaces when a reference uses \usepackage, \RequirePackage, or \bibliography, which strip those spaces. Unresolved references and files excluded from discovery remain unchanged. Installed TEXMF files and navigation-only .dtx fallbacks are never renamed. Renaming directly from bibliography, graphics, package, or class arguments is not supported.

TEXMF discovery

How the language server discovers the installed TeX tree for package resolution: document links, package hover, go-to-definition, and installed-set completion. Where a TeX installation lives is a fact about the machine, not the project, so these settings come from the editor rather than badness.toml, and they never affect badness format or badness lint, whose output stays a pure function of the input regardless of what is installed.

A texmf object with three keys, all optional:

  • enabled (boolean, default true): whether to scan the TEXMF tree at all. When false, package resolution stays local to the document’s directory.
  • roots (array of paths, default []): extra TEXMF root directories to index in addition to (and ahead of) the discovered ones. Useful for a non-standard install that kpsewhich can’t see.
  • useKpsewhich (boolean, default true): whether to shell out to kpsewhich to discover the TEXMF tree roots. When false, discovery falls back to default-path heuristics only.
{ "texmf": { "enabled": true, "roots": ["/opt/texmf"], "useKpsewhich": true } }

Jump between a source line and the matching place in the compiled PDF.

Badness never typesets, and it never reads a .synctex.gz. Forward search works out three things — the file your cursor is in, the root document’s PDF, and the line number — and hands them to a viewer you configure. Every SyncTeX-aware viewer (zathura, Okular, SumatraPDF, Skim) links libsynctex and does the mapping itself, which is why they all want a file and a line rather than a coordinate. Inverse search runs in the other direction and is started by the viewer.

You need a PDF compiled with SyncTeX enabled — latexmk -pdf -synctex=1, or -synctex=1 passed to pdflatex/lualatex directly. Badness will not run that for you; use your existing build setup, or an extension like LaTeX Workshop.

Configuring the viewer

Which viewer is installed on your machine, and under what name, is a fact about the machine rather than the project — so these settings come from the editor, like TEXMF discovery, and not from badness.toml. Where the PDF lives is project data and belongs to the [build] section instead.

A forwardSearch object:

  • executable (string): the viewer program. Spawned directly, not through a shell, so it is a program name and never a command line — putting flags here ("zathura --synctex-forward") silently fails to launch. This is the most common misconfiguration.
  • args (array of strings): the viewer’s arguments. Required — there is no useful default, since every viewer spells forward search differently. Without it, forward search reports itself unconfigured.
  • ipcDir (path, optional): where inverse-search servers advertise themselves. An escape hatch for containers and sandboxes; see below.

Each argument may carry:

PlaceholderExpands to
%fthe .tex file the cursor is in
%pthe root document’s PDF
%lthe line number, counting from 1
%%fa literal %f

An argument wrapped entirely in " is passed through with the quotes stripped and nothing substituted — the escape hatch when a viewer needs a literal %.

Recipes, matching texlab’s, so an existing configuration ports unchanged:

Viewerexecutableargs
zathurazathura["--synctex-forward", "%l:1:%f", "%p"]
Okularokular["--unique", "file:%p#src:%l%f"]
SumatraPDFSumatraPDF["-reuse-instance", "%p", "-forward-search", "%f", "%l"]
Skimdisplayline["%l", "%p", "%f"]
Evinceevince-synctex["-f", "%l", "%p", "\"code -g %f:%l\""]
qpdfviewqpdfview["--unique", "%p#src:%f:%l:1"]
{
  "forwardSearch": {
    "executable": "zathura",
    "args": ["--synctex-forward", "%l:1:%f", "%p"]
  }
}

The server handles textDocument/forwardSearch, a custom request taking the standard { textDocument, position } params — the same method name and shape texlab uses, so a client written for texlab works unchanged. It never fails the request; it answers with a status:

StatusMeaning
0the viewer was launched
1the viewer would not start
2no PDF on disk, or the buffer has no path — build the document first
3no viewer configured

The capability is advertised as experimental.textDocumentForwardSearch.

If forward search opens the wrong PDF, or reports status 2 on a project that has been built, the root document is probably not being found — see root in the [build] reference.

Configure your viewer to run:

badness inverse-search --input "%f" --line "%l"

substituting the viewer’s own placeholders. For zathura that is:

zathura --synctex-editor-command "badness inverse-search --input %{input} --line %{line}"

Use --line0 instead if your viewer counts lines from zero. (--line1 is accepted as a synonym for --line, so a texlab configuration ports directly.)

The command finds the language server whose workspace contains the file and asks it to reveal the position, so an editor must already have that project open, and its LSP client must support window/showDocument. Servers whose client does not support it never register, which is why inverse search silently does nothing in an editor lacking it — the command says so when nothing is listening.

With several editor windows open, the server whose workspace root contains the file wins; the longest matching root is preferred, so nested projects resolve deterministically.

Servers advertise themselves in $BADNESS_IPC_DIR, else a per-user directory under your runtime directory ($XDG_RUNTIME_DIR), else the temporary directory. The forwardSearch.ipcDir setting overrides all of these — useful when the viewer and the server see different filesystems, as in a container or a remote development setup. Keep it short: a Unix socket path cannot exceed about 100 bytes, and badness says so explicitly in its log if yours does. On a system with no $XDG_RUNTIME_DIR and a /tmp shared between users, that last fallback is worth knowing about: the directory is created 0700, the advertisements 0600, and badness ignores any advertisement it does not own, so another user can neither read nor impersonate one.

One caveat inherent to SyncTeX: it maps the source as it was compiled. With unsaved edits, buffer line numbers and PDF line numbers drift apart until you rebuild.

Table refactoring

With the cursor inside a statically understood tabular, tabular*, or array environment, the Add column at end code action appends a centered c column to the preamble and an empty trailing cell to every row. The action is withheld when the preamble uses unknown column types, a row has an ambiguous width, or the environment has been redefined, so it never applies a partial table rewrite.

Neovim

With the built-in vim.lsp client (Neovim 0.11+):

vim.lsp.config.badness = {
  cmd = { "badness", "lsp" },
  filetypes = { "tex", "latex", "plaintex", "bib" },
  root_markers = { "badness.toml", ".git" },
  init_options = { lineWidth = 80, indentWidth = 2 },
}
vim.lsp.enable("badness")

The init_options block is optional; omit it to use the defaults or a badness.toml.

VS Code

Install the Badness extension from the VS Code Marketplace or the Open VSX extension. It bundles a platform-specific badness binary and starts the language server automatically when you open a .tex file, so no separate CLI install is required.

The extension is configured through badness.* settings. By default it uses the bundled binary (badness.executableStrategy: "bundled"); set the strategy to environment to use a badness on your PATH, or path with badness.executablePath to point at a specific binary. See the extension’s README for the full list of settings.

Using only some features

The formatter, linter, and language features share one server but can be turned off independently, so you can adopt just the parts you want:

  • badness.formatting.enable — use Badness as a formatter.
  • badness.diagnostics.enable — show Badness diagnostics (the linter).
  • badness.languageFeatures.enable — hover, completion, navigation, symbols, rename, code actions, and the rest.

All three default to true. They are client-side gates, so the server keeps running and the toggles take effect without a reinstall. For a formatter-only setup, turn off the other two:

{
  "badness.diagnostics.enable": false,
  "badness.languageFeatures.enable": false
}

Turning off badness.diagnostics.enable this way suppresses every diagnostic, including the syntax/parse errors that a badness.toml [lint] selection cannot silence. The badness.toml route stays the right tool when you want to keep parse errors but mute specific lint rules across every editor and the CLI.

Using with LaTeX Workshop

Badness works alongside LaTeX Workshop rather than replacing it. The two divide cleanly: LaTeX Workshop handles building, PDF preview, and SyncTeX, while badness handles formatting, linting, and navigation. Run both, and let each own its half.

Formatting. The badness extension registers itself as the default formatter for LaTeX files. LaTeX Workshop’s own formatter integration is disabled by default (latex-workshop.formatting.latex is "none"); leave it that way so there is a single formatting authority. For BibTeX files, LaTeX Workshop ships a built-in formatter, so pick badness explicitly:

{
  "[bibtex]": {
    "editor.defaultFormatter": "jolars.badness"
  }
}

Linting. LaTeX Workshop’s ChkTeX and lacheck integrations are disabled by default (latex-workshop.linting.chktex.enabled and latex-workshop.linting.lacheck.enabled). Leave them off; enabling them alongside badness produces overlapping diagnostics for many common issues.

Completion. Both extensions contribute completion items, so you may see duplicate suggestions for commands, environments, or citations. This is harmless, but if it bothers you, the latex-workshop.intellisense.* settings let you turn off the overlapping parts on the LaTeX Workshop side.

Other Editors

Any LSP-capable editor can run badness: configure a server whose command is badness lsp, communicating over stdio, for LaTeX documents. Consult your editor’s LSP client documentation for the exact configuration shape.

Pre-commit

Badness ships a pre-commit hook through the badness-pre-commit mirror repository. The hook installs the prebuilt badness wheel from PyPI, so no Rust toolchain is needed. Add this to your .pre-commit-config.yaml:

repos:
  - repo: https://github.com/jolars/badness-pre-commit
    # badness version
    rev: v0.9.0
    hooks:
      # Lint .tex, .sty, .cls, .dtx, .ins, and .bib files
      - id: badness-lint
      # Format the same files in place
      - id: badness-format

Tags mirror badness releases: rev: v0.9.0 runs badness 0.9.0.

To apply safe lint autofixes before formatting (the fix-then-format pipeline), pass --fix:

- id: badness-lint
  args: [--fix]
- id: badness-format

To check formatting without rewriting files:

- id: badness-format
  args: [--check]

The hook then writes nothing, so pre-commit has no modified file to show you; the diff --check prints is the whole report. Add --quiet alongside it if you would rather see only the list of files that would be reformatted.

Configuration

Badness is configured through a badness.toml file. All keys are optional and spelled in kebab-case; an unknown key or section is a hard error, not a silent no-op. Run badness init to write a commented starter file with defaults and examples.

# extend = "../shared/badness.toml"

# Gitignore-style patterns to skip during directory discovery.
# exclude = [".git/"]
# extend-exclude = []

[format]
# line-width = 80  # 0 disables width-based wrapping
# indent-width = 2
# item-indent = "hang"  # hang | indent | none
# wrap = "reflow"  # reflow | stable | sentence | semantic | preserve
# line-ending = "auto"  # auto | lf | crlf | native

[lint]
# select = ["..."]  # if set, only these rules run
# ignore = []       # rules to disable

Discovery

For each input, Badness walks from the file’s directory upward and uses the first badness.toml it finds. The walk stops at a directory containing a .git entry (the repository root), so a config file outside your repository is never picked up.

If no project file is found, Badness next checks the BADNESS_CONFIG environment variable. When set (and non-empty), it names a config file to use instead of the global user config below—handy for keeping one config on a synced drive and pointing every machine at it. A set BADNESS_CONFIG shadows the global config entirely.

If BADNESS_CONFIG is unset, badness falls back to a global user config: the first existing file among

  1. $XDG_CONFIG_HOME/badness/config.toml
  2. ~/.config/badness/config.toml
  3. the platform config directory (%APPDATA%\badness\config.toml on Windows, ~/Library/Application Support/badness/config.toml on macOS)

The BADNESS_CONFIG and global files use the same schema as a project badness.toml and are whole-file fallbacks, never merged with a project config. Relative exclude patterns in them resolve against the working directory (CLI) or the document’s directory (language server) rather than the config’s own directory. The language server uses the same resolution, so both are easy ways to set editor-wide defaults such as wrap = "preserve" (an edit is picked up when the server restarts). If none of these files is found, built-in defaults apply.

Two global CLI flags override discovery:

  • --config <PATH> uses that file instead of discovering one.
  • --no-config ignores any project, BADNESS_CONFIG, or global file and uses built-in defaults.

CLI flags for individual options (--line-width, --wrap, --select, etc.) override the corresponding config values for a single run.

Editor support

Badness publishes a JSON Schema for badness.toml so editors with TOML support can provide key and value completion, hover documentation, and validation.

Schema URL: https://badness.dev/badness.schema.json

The schema is generated from the configuration types and checked against them in the test suite. The version at the URL tracks the latest released version of Badness.

VS Code (Even Better TOML)

With the Even Better TOML extension installed, add this association to your user or workspace settings.json:

{
  "evenBetterToml.schema.associations": {
    "^(.*/)?badness\\.toml$": "https://badness.dev/badness.schema.json"
  }
}

Inline #:schema directive

TOML tooling that supports inline schema directives can select it from the configuration file itself:

#:schema https://badness.dev/badness.schema.json

This is also useful for the global user configuration, whose generic config.toml file name should not be associated with Badness automatically.

Other editors

Any editor or language server that consumes JSON Schemas—including Helix, Neovim with taplo-lsp, Zed, and IntelliJ—can use the same URL through its TOML schema settings.

Top level

extend

Load another configuration file as a base. A relative path starts at the directory containing the file that declares extend; absolute paths and paths starting with ~/ also work. An extended file may itself use extend. A missing file or a cycle is an error.

Badness merges tables by key, so a child can override one [format] setting while inheriting the others. Values in the child replace inherited values, including arrays. extend-exclude is the exception: its patterns append to the inherited patterns, in base-to-child order. Named [commands] and [environments] declarations merge by name. Relative paths in inherited settings use the selected config’s path root. For project configs, that root is the selected config’s directory.

Default value: unset

Type: string

Example:

extend = "../shared/badness.toml"

[format]
line-width = 100

exclude

Gitignore-style patterns to exclude from directory discovery, resolved relative to the directory containing the badness.toml. Excludes apply to both format and lint, which share one file walk, so this is a top-level key rather than a [format] option.

When set, this replaces the built-in default set ([".git/"]); use extend-exclude to add patterns without restating the defaults. Patterns given with the --exclude CLI flag are always added on top.

Default value: [".git/"]

Type: array of strings

Example:

exclude = ["vendor/", "old-drafts/"]

extend-exclude

Gitignore-style patterns added in addition to the base set selected by exclude (the built-in defaults when exclude is unset). Use this to skip a few extra paths without replacing the defaults.

Default value: []

Type: array of strings

Example:

extend-exclude = ["build/"]

[format]

Options for badness format. Each mirrors a CLI flag of the same name, which takes precedence for a single run.

line-width

Maximum line width before the formatter breaks a line. Must be between 0 and 1000. Set line-width = 0 to disable width-based wrapping. Sentence breaks, authored breaks retained by the selected wrap mode, and structural line breaks still apply. This setting also controls width-based layout in display math, command arguments, and BibTeX values. Protected content and indivisible atoms may exceed a positive width.

Default value: 80

Type: integer

Example:

[format]
line-width = 100

indent-width

Spaces per indent step. Must be between 1 and 1000.

Default value: 2

Type: integer

Example:

[format]
indent-width = 4

item-indent

How continuation lines in list items are indented relative to the \item command.

ModeBehavior
hangAlign under the body following a bare \item (the default).
indentAdd one indent-width step from the \item column.
noneAlign with the \item command.

Labels and Beamer overlays do not widen the hang offset, so items retain one continuation edge regardless of marker width.

Default value: "hang"

Type: string

Example:

[format]
item-indent = "indent"

wrap

How the formatter lays out line breaks inside a paragraph. It does not affect structure, only where soft line breaks fall.

ModeBehavior
reflowGreedy fill: pack words up to line-width, breaking only where the next word would overflow.
stablePreserve acceptable authored breaks and rebalance only text that no longer fits (keeps revision diffs small).
preserveLeave the authored line breaks untouched.
sentenceOne sentence per line. Line width is ignored—a long sentence stays on one line.
semanticSemantic line breaks: keep authored breaks, add sentence breaks, and wrap overlong lines to line-width.

Both sentence and semantic split a paragraph at sentence boundaries. Boundary detection is a small per-language rule engine over the words: a ., !, or ? ends a sentence unless the word is a known abbreviation (e.g., Fig., Dr., etc.) an ellipsis (..., …), or a contextual abbreviation whose following word signals that the sentence continues (U.S. Government stays together, U.S. However splits). The abbreviation profile is chosen by lang and extended by no-break-abbreviations.

In sentence mode, citation commands follow their grammatical role. Parenthetical and other postpositive forms, such as \parencite, \citep, and \autocite, stay with the preceding sentence—even when an authored line break has stranded one on the next line. Textual forms, such as \textcite and \citet, may begin the next sentence. Plain \cite is package- and style-dependent, so the formatter follows the source: it attaches on the same line but preserves an authored line break.

semantic additionally preserves the author’s own line breaks on top of the sentence breaks (the sembr convention). It does not detect clause boundaries itself—a break after a comma or and survives only where the author placed a newline. Long sentences wrap at line-width, so one sentence can span several lines. Each sentence starts on a new line even if it would fit beside the preceding one.

Earlier versions ignored width in semantic mode. To keep unlimited semantic prose, use:

[format]
wrap = "semantic"
line-width = 0

Zero also disables width-based wrapping outside prose. Authored breaks remain preserved, including breaks inserted by an earlier formatting pass: increasing the width does not automatically join them.

stable also preserves authored line breaks, but treats them as preferred anchors rather than hard boundaries. It is aimed at keeping revision diffs small: a small prose edit perturbs the smallest possible region. Each prose run is solved as one global layout problem. Candidate layouts are compared lexicographically by total overflow, underflow below a soft target (line-width - 15), changed authored breaks, displacement from the nearest authored break, raggedness around that target, and line count. This makes the hard width non-negotiable before minimizing source churn, while a short final line remains unpenalized. Blank lines and command-only lines bound each independently optimized run, and code-like statement bodies retain ordinary greedy fill. (The soft target is not currently configurable.) With line-width = 0, stable retains authored breaks without balancing line lengths.

When omitted, every file kind reflows—.tex, .bib, .sty, .cls, .dtx, and .ins alike. A file’s extension is not a layout input.

That is safe because reflow is never the thing that decides whether content may move. The formatter declines to reflow anything it cannot lay out without changing meaning, in every wrap mode and regardless of what you configure: verbatim bodies and \verb, comments, .dtx documentation margins and docstrip guards (which must stay at column 0), and any documentation block whose rewrapping would push a % off column 0. Asking for wrap = "reflow" on a .dtx cannot corrupt it; asking for wrap = "preserve" on a .tex is a stylistic choice, not a safety one.

Beamer overlay bodies in \only, \uncover, \visible, \invisible, \onslide, \action, \alt, and \temporal retain their authored line breaks in every wrap mode. Multiline bodies stay multiline, and inline bodies stay inline even when they exceed line-width. Indentation still normalizes.

Code, in practice, has little to reflow: expl3 regions (\ExplSyntaxOn…\ExplSyntaxOff) are laid out by their own rules whatever wrap says, and a source line consisting only of commands keeps its own line. So a package or class body formats much as it did before, and preserve remains available if you want authored breaks kept verbatim.

Default value: unset (reflow, for every file kind)

Type: "reflow" | "stable" | "sentence" | "semantic" | "preserve"

Example:

[format]
wrap = "stable"

math-wrap

How the formatter lays out line breaks inside display math: \[…\], $$…$$, and single-formula math environments such as equation. Alignment-grid environments (align, gather, matrices) and inline $…$ math are not affected.

ModeBehavior
autoDerive from the effective wrap: preserve keeps authored math breaks, every other mode breaks (amsmath).
preserveKeep the authored line breaks inside the body. Spacing within each line is still normalized.
single-lineNever insert breaks: the body stays on one line, overflowing line-width if too long (like inline math).
breakBreak a too-long body before its top-level relations and binary operators, aligning a relation chain (amsmath style).

In break mode, multiline sibling environments also form separate blocks at the display body’s indentation. This structural layout applies regardless of line-width; punctuation remains attached, and intervening expressions get their own lines.

Default value: "auto"

Type: "auto" | "preserve" | "single-line" | "break"

Example:

[format]
math-wrap = "preserve"

line-ending

How the line breaks in formatted output are spelled. The layout engine always decides where breaks go; this decides only the bytes they render as, and it applies to the whole document — including inside verbatim-style protected regions, which would otherwise keep their authored endings and leave the file mixed.

ModeBehavior
autoKeep the endings the file was written with: CRLF if its first line break is one, LF otherwise.
lfAlways \n.
crlfAlways \r\n.
nativeThe platform’s convention: \r\n on Windows, \n elsewhere.

The default is auto, so formatting never rewrites a repository’s line endings on its own — set lf (or add a .gitattributes rule) if you want them normalized.

Default value: "auto"

Type: "auto" | "lf" | "crlf" | "native"

Example:

[format]
line-ending = "lf"

lang

Document language as a BCP-47-style code (en, de, pt-BR, …), used by the sentence and semantic wrap modes to pick the sentence-boundary abbreviation profile. Built-in profiles cover English (default), Czech, German, Spanish, and French; the region subtag is folded away, and an unknown or unset language falls back to English. (Automatic detection from babel/polyglossia is not yet implemented.)

Default value: unset (English)

Type: string

Example:

[format]
lang = "de"

no-break-abbreviations

User-supplied no-break abbreviations for the sentence and semantic wrap modes, keyed by language code or the literal default bucket (applied to every document). An abbreviation listed here never ends a sentence, so no line break is inserted after it. Merged on top of the built-in per-language lists.

Default value: {}

Type: table of string arrays, keyed by language code or default

Example:

[format.no-break-abbreviations]
default = ["ibid."]         # applied to every document
de = ["bzw.", "Abb."]       # applied only when lang resolves to German

[lint]

Rule selection for badness lint, shared by the LaTeX and BibTeX rule sets. Most rules are on by default; each rule’s reference entry states whether it is default-enabled. An unknown rule id is reported at lint time, not rejected at config-parse time.

select

Explicit allowlist of rule ids. When set, only these rules run.

Default value: unset (all default-enabled rules run)

Type: array of strings

Example:

[lint]
select = ["deprecated-command", "dollar-display-math"]

To enable the opt-in dash check:

[lint]
select = ["dash-length"]

ignore

Rule ids to disable, applied on top of either select or the default rule set.

Default value: []

Type: array of strings

Example:

[lint]
ignore = ["missing-nonbreaking-space"]

[build]

Where the TeX compiler leaves its artifacts, and which file it was run on. Read by the language server only — it pulls resolved label and section numbers from the .aux files for hover and document symbols, and locates the compiled PDF for forward search. Never read by the formatter or linter.

aux-dir

Directory holding the build’s .aux files (latexmk’s -auxdir/-outdir), resolved relative to the root document’s directory when not absolute. When unset, each document’s .aux is expected next to it, as in plain latex/pdflatex runs.

Default value: unset (sibling .aux files)

Type: path

Example:

[build]
aux-dir = "out"

pdf-dir

Directory holding the build’s PDF output (latexmk’s -outdir), resolved relative to the root document’s directory when not absolute. When unset, the PDF is expected next to the root document.

Default value: unset (the root document’s own directory)

Type: path

Example:

[build]
pdf-dir = "out"

pdf-filename

The compiled PDF’s file name, when the build does not name it after the root document (latexmk’s -jobname). A bare file name, never a path — use pdf-dir for the directory — and .pdf is appended when it carries no extension, so "thesis" and "thesis.pdf" mean the same thing.

Default value: unset (<root document stem>.pdf)

Type: string

Example:

[build]
pdf-filename = "thesis.pdf"

root

The project’s root document — the file the compiler was run on — resolved relative to this badness.toml’s directory when not absolute.

Normally the root is found by scanning the project for a file carrying \documentclass or \begin{document}, and you do not need this key. But that scan only sees files the server has already loaded, and it loads them one directory at a time: editing chapters/ch1.tex in a project rooted at ../main.tex never loads main.tex, so the scan finds no root at all and forward search resolves the wrong PDF. Set root for that layout.

Default value: unset (scan the project for a document root)

Type: path

Example:

[build]
root = "main.tex"

[commands]

Declares project commands whose first braced argument contains label or citation keys. This covers wrappers that Badness cannot recognize without macro expansion, which remains deliberately out of scope.

Entries are keyed by the command name without its leading backslash:

# A comma-separated list of label keys.
[commands.eqrefs]
like = "cref"

# A comma-separated list of bibliography keys.
[commands.projectcite]
like = "parencite"

The like target must be a curated reference or citation command. It determines the key behavior: ref/eqref accept one label, cref and its list-valued siblings split on commas, citation commands split on commas, and nocite preserves the special * wildcard.

Command declarations affect linting, label/citation navigation, rename, and key completion. They do not expand the macro, declare its arity, change argument attachment, or lend formatter layout. Use the command whose observable key behavior matches the wrapper: an eqrefs command that accepts several labels is like = "cref", even if its implementation calls \eqref once per key.

Anything that would silently do nothing is a configuration error: an empty entry, an invalid control-word name, an unknown or non-ref/cite like target, or an attempt to reclassify a curated built-in command.

Command like

The curated reference or citation command whose key behavior this project command copies.

Type: string

Example:

[commands.eqrefs]
like = "cref"

[environments]

Declares environments Badness cannot recognize from the file alone: one that behaves like a built-in but has no built-in counterpart, one whose body is verbatim, and one reached through command spellings rather than \begin/\end. This is the only section that changes how your files are parsed, so it is read by format, lint, and the language server alike; editing it makes the server reparse the project.

Entries are keyed by the environment’s own name, whether or not Badness already knows it:

# \begin{myenv} … \end{myenv}, with no built-in counterpart
[environments.myenv]
like = "align"

# extra delimiter spellings for an environment Badness already knows
[environments.eqnarray]
begin = ['\bea']
end = ['\eea']

# one side alone: `\bsplit` expands to `\begin{split}`, so a written-out
# `\end{split}` closes it — there is no closing command to declare
[environments.split]
begin = ['\bsplit']

# both at once: an environment reached only through commands
[environments.mytheorem]
like = "theorem"
begin = ['\startmyenv']
end = ['\endmyenv']

Write control words as TOML literal strings (single quotes) so the backslash needs no escaping: '\bea', not "\\bea". Both spellings are accepted, and so is a name with no backslash at all — a control word can never contain one, so there is nothing to disambiguate.

A declaration names a spelling, never a pairing. Every structural rule still applies, so a declared \bea whose \eea is unreachable — stranded inside a brace group, or simply missing — stays an ordinary command, exactly as it would without the declaration. A wrong declaration therefore does nothing to your document; it cannot corrupt it.

What a declaration cannot do is invent behavior. It only ever points at an environment Badness already curates, so there is no way to spell out “this one is math, takes two arguments, and has a verbatim body” key by key. If nothing built in resembles yours, that is worth an issue rather than a workaround.

Anything a declaration cannot satisfy is an error at config load, reported against the key you wrote, rather than a block that parses and quietly does nothing:

  • an entry with no keys under it, which would declare nothing at all
  • like naming an environment Badness does not know
  • delimiter spellings for a verbatim environment (the closing command is never seen — the verbatim body has already swallowed it)
  • delimiter spellings for an environment that takes arguments (a bare command carries none)
  • delimiter spellings for an environment with no like and no built-in of that name, so its behavior is unknown
  • one spelling claimed by two entries, or listed twice by one
  • a spelling that is the delimiter itself ('\end{split}') rather than a command standing in for one — the written-out delimiter already pairs with a declared spelling, so the key can just be removed
  • a spelling that could never be a single control word ('\b ea', '\bea2')
  • a spelling that is already a LaTeX command Badness knows ('\emph'), which would change what that command means throughout the project

like

The built-in environment whose behavior this one copies: whether its body is math, whether it aligns on &, whether it is verbatim, and every such property at once. This is also how you name a verbatim environment defined by machinery no scan can follow — like = "lstlisting" protects its body from reflowing and from lint findings.

The target is looked up among the environments Badness curates by hand; a misspelled one is an error rather than a silent no-op.

Default value: unset

Type: string

Example:

[environments.mycode]
like = "lstlisting"

begin

Command spellings that stand in for this environment’s \begin{…}. Any of them opens it, and any spelling in end closes it — pairing is by side, not by position, so the two lists need not be the same length.

The written-out \end{…} closes it too, which is why end is optional. A command defined as \def\bsplit{\begin{split}} expands to \begin{split}, so \bsplit … \end{split} is a perfectly ordinary environment and there may be no closing command to name at all.

Use this when the definition is somewhere Badness cannot see: a sibling .sty, or one built by machinery no scan follows. A definition written with a plain \newcommand or \def in the same file — \newcommand{\bea}{\begin{eqnarray}}, or \def\bsplit{\begin{split}} on its own — is already recognized without any configuration.

A spelling must be a command of your own. Naming one Badness already knows ('\emph', '\section') is an error rather than a redefinition: the declaration would apply everywhere that command appears, which is never what a delimiter declaration means.

Default value: []

Type: array of strings (control words)

Example:

[environments.eqnarray]
begin = ['\bea', '\beqa']
end = ['\eea']

end

Command spellings that stand in for this environment’s \end{…}, the mirror of begin in every respect — including that it stands alone. A command defined as \def\eeq{\end{equation}} closes a written-out \begin{equation}, so an entry may name a closing spelling without naming an opening one.

Default value: []

Type: array of strings (control words)

Example:

[environments.eqnarray]
begin = ['\bea']
end = ['\eea']

Note: TEXMF-tree discovery (the former [texmf] section) is configured through your editor’s LSP settings, not badness.toml. Where a TeX installation lives is a fact about the machine, not the project, so it does not belong in a file shared across contributors. See Editor Setup.

Command-line reference

A formatter, linter, and language server for LaTeX

Usage: badness [OPTIONS] <COMMAND>

Options

--config <PATH>

Path to a badness.toml to use instead of discovering one. Applies to format and lint; ignored by parse, lsp, and init

--no-config

Ignore any badness.toml (project, $BADNESS_CONFIG, or global) and use built-in defaults

--color <WHEN>

When to use color in output

Default value: auto

Possible values:

  • auto: Colorize when writing to a terminal and NO_COLOR is unset (default)
  • always: Always colorize
  • never: Never colorize
-q, --quiet

Suppress non-essential output (errors are still shown). Under format --check this drops the per-file diff, leaving the list of files that would be reformatted and the summary

badness format

Format LaTeX source.

With paths, formats each file in place. Reads stdin (to stdout) when given -, or when paths are omitted and stdin is not a terminal.

Usage: badness format [OPTIONS] [PATHS]...

Arguments

<PATHS>...
Files or directories to format. Pass - for stdin, which is also read when paths are omitted and stdin is not a terminal

Options

--check

Report which files would change without writing them. Exits non-zero if any file is not already formatted. Requires path arguments: there is no file on disk to report on when reading stdin

--stdin-filepath <PATH>

Name the stdin buffer so its language is dispatched by extension (.bib → BibTeX, anything else → LaTeX). No file is read or written; only the extension is used. Ignored when paths are given

--line-width <LINE_WIDTH>

Maximum line width before the formatter breaks a line; 0 disables width-based wrapping

--indent-width <INDENT_WIDTH>

Number of spaces per indent step

--item-indent <ITEM_INDENT>

How to indent continuation lines in list items

Possible values:

  • hang: Align continuations under the body following a bare \item (default)
  • indent: Indent continuations by one indent-width step
  • none: Align continuations with the \item command
--wrap <WRAP>

How to lay out line breaks inside a paragraph

Possible values:

  • reflow: Greedy fill: wrap words to the line width (default)
  • stable: Preserve acceptable authored breaks and rebalance only nearby text (revision-stable wrapping)
  • sentence: One sentence per line (line width ignored)
  • semantic: Semantic line breaks (sembr.org): keep authored breaks, add sentence breaks, and wrap overlong lines to the line width
  • preserve: Leave authored line breaks untouched
--math-wrap <MATH_WRAP>

How to lay out line breaks inside display math

Possible values:

  • auto: Derive from the effective wrap mode: preserve → preserve, else break (default)
  • preserve: Keep authored line breaks inside display-math bodies
  • single-line: Never insert breaks; a long body overflows the line width
  • break: Break a too-long body before its top-level operators (amsmath style)
--line-ending <LINE_ENDING>

How to spell the line breaks in the formatted output

Possible values:

  • auto: Keep the endings the file was written with (default)
  • lf: Always LF (\n)
  • crlf: Always CRLF (\r\n)
  • native: The platform’s convention: CRLF on Windows, LF elsewhere
--exclude <PATTERN>

Gitignore-style pattern to skip during directory discovery (repeatable). Added on top of any exclude/extend-exclude from badness.toml

--force-exclude

Apply exclude patterns to files named explicitly on the command line too (they are normally always processed). For runners like pre-commit that pass staged files as arguments

badness lint

Lint LaTeX source, reporting parse diagnostics.

With paths, lints each file. Reads stdin when given -, or when paths are omitted and stdin is not a terminal. Exits non-zero if any diagnostics are reported.

Usage: badness lint [OPTIONS] [PATHS]...

Arguments

<PATHS>...
Files or directories to lint. Pass - for stdin, which is also read when paths are omitted and stdin is not a terminal

Options

--fix

Apply safe autofixes in place, then report what remains. Requires path arguments; has no effect on stdin (there is nothing to write)

--unsafe-fixes

Also apply fixes that may change typeset output (requires --fix)

--stdin-filepath <PATH>

Name the stdin buffer so its language is dispatched by extension (.bib → BibTeX, anything else → LaTeX). No file is read or written; only the extension is used. Ignored when paths are given

--exclude <PATTERN>

Gitignore-style pattern to skip during directory discovery (repeatable). Added on top of any exclude/extend-exclude from badness.toml

--force-exclude

Apply exclude patterns to files named explicitly on the command line too (they are normally always processed). For runners like pre-commit that pass staged files as arguments

--select <RULE>

Run only these rules (repeatable). Overrides [lint] select from badness.toml when given

--ignore <RULE>

Disable these rules (repeatable). Overrides [lint] ignore from badness.toml when given

--explain <RULE>

Print the description and examples for a rule id, then exit. Ignores paths, config, and fixes

--output <OUTPUT>

Output format for findings. The human modes write to stderr; json writes to stdout

Default value: pretty

Possible values:

  • pretty: Source-snippet output with caret spans, on stderr (default)
  • concise: One path:line:col: severity [rule] message line per finding, on stderr
  • json: A machine-readable JSON array of findings on stdout ([] when clean), with byte-offset ranges and fix data

badness parse

Parse LaTeX source and print its concrete syntax tree (CST).

A debugging aid: prints the lossless parse tree as an indented KIND@range listing, with token text, followed by any parse errors. With a path, parses that file. Reads stdin when given -, or when the path is omitted and stdin is not a terminal.

Usage: badness parse [PATH]

Arguments

<PATH>
File to parse. Pass - for stdin, which is also read when the path is omitted and stdin is not a terminal

badness lsp

Run the language server over stdio

Usage: badness lsp

Answer a PDF viewer’s inverse (backward) search.

Point your viewer’s inverse-search command here — for zathura, --synctex-editor-command "badness inverse-search --input %{input} --line %{line}". The position is handed to a running badness language server, which reveals it in your editor via window/showDocument, so the file must belong to a workspace some editor currently has open.

Usage: badness inverse-search [OPTIONS] --input <PATH>

Options

-i, --input <PATH>

The .tex file the viewer resolved

-l, --line <LINE>

Line number, counting from 1 — what SyncTeX-aware viewers emit.

Required unless --line0 is given. Deliberately not enforced by clap, whose message for that would name only --line and so send a --line0 user the wrong way.

--line0 <LINE>

Line number counting from 0, for a viewer that reports it that way

--character <COLUMN>

Column, counting from 0, when the viewer supplies one

Default value: 0

--ipc-dir <DIR>

Directory holding the servers’ IPC advertisements. Defaults to $BADNESS_IPC_DIR, then a per-user directory under the runtime (or temporary) directory

badness init

Write a commented starter badness.toml to the current directory

Usage: badness init [OPTIONS]

Options

--force
Overwrite an existing badness.toml

Linter Rules

badness lint runs a set of built-in rules over each file’s parse tree and reports a diagnostic for every finding. This page is the catalogue: one section per rule, keyed by its stable rule id. That id is what appears in a diagnostic, what [lint] select/ignore (and --select/--ignore) target, and what a % badness-lint skip <id> comment suppresses.

Most rules are on by default. Each rule’s section states its default; enable an opt-in rule with select, or narrow the default set with select/ignore in the [lint] table (see the Configuration reference). Where a rewrite is unambiguous a rule carries an auto-fix: a safe fix (shown below as “After applying the fix”) is applied by badness lint --fix; an unsafe fix, one that may change output such as inserting a line-breaking tie, is applied only with --unsafe-fixes or as an editor code action, so it has no “after” block here.

Each example below is linted live to produce its diagnostic and fixed output, so this page never drifts from the rules’ actual behavior.

This page covers the LaTeX linter. BibTeX files have a parallel set of rules (a separate BibRule registry under src/bib/linter/), selectable through the same [lint] config and catalogued in BibTeX Linter Rules.

abbreviation-spacing

Flag TeX’s sentence-vs-interword spacing going wrong around abbreviations and acronyms (ChkTeX 12/13). Outside \frenchspacing, TeX widens the space after ./?/! unless the punctuation follows an uppercase letter. Two shapes defeat that: a lowercase abbreviation (e.g., i.e., etc., et al.) gets a too-wide space, fixed with \ (e.g.\ foo); and an uppercase acronym ending a sentence (USA.) gets a too-narrow space, fixed with \@ (USA\@.). To stay conservative the first fires only before a lowercase word (the sentence clearly continues) and the second only for a run of two or more capitals before the period and before an uppercase word (a new sentence), so initials (J.), dotted forms (U.S.A.), and mid-sentence acronyms are left alone. Both fixes are unsafe – they change the typeset spacing – so --fix leaves them alone while --unsafe-fixes and the editor code action apply them. The rule is silent under \frenchspacing, and never touches comments, verbatim, or math.

This rule is enabled by default.

A lowercase abbreviation followed by more text takes an interword space:

We tried several methods, e.g. gradient descent.
warning: abbreviation-spacing
 --> example.tex:1:31
  |
1 | We tried several methods, e.g. gradient descent.
  |                               ^ `e.g.` is an abbreviation, not a sentence end; use an interword space `\ ` (`e.g.\ `) so TeX does not widen the gap

An acronym ending a sentence takes intersentence spacing:

The rover reached the USA. Then it stopped.
warning: abbreviation-spacing
 --> example.tex:1:26
  |
1 | The rover reached the USA. Then it stopped.
  |                          ^ capital before sentence-ending punctuation suppresses intersentence spacing; use `\@` (`Word\@.`) to restore it

blank-line-in-keyval

Flag a blank line at the top level of a key=value argument. A blank line is a \par token and a keyval processor walks its entries with macros that are not \long, so the call aborts – and the error TeX reports names the processor rather than the command the author wrote (\hypersetup yields “Paragraph ended before \kv@processor@default was complete”), which is what makes the finding worth more than the compiler’s own message. Scoped by measurement: a blank line nested inside a value’s brace group (\tikzset{aa/.style={draw,\n\nthick}}) compiles clean and is not flagged, an unclosed { is left to the parse error it already draws, and only the hand-curated signature tier is consulted. The autofix drops the blank line and keeps the following indentation; it is safe by construction, since it edits only whitespace and ContentKind::Keyval is exactly the claim that the processor strips spaces around entries.

This rule is enabled by default.

A blank line separating two keys, which aborts the call:

\hypersetup{colorlinks=true,

linkcolor=blue}
error: blank-line-in-keyval
 --> example.tex:1:29
  |
1 |   \hypersetup{colorlinks=true,
  |  _____________________________^
2 | |
3 | | linkcolor=blue}
  | |_^ blank line in `\hypersetup`'s key-value argument; the `\par` aborts the call

After applying the fix:

\hypersetup{colorlinks=true,
linkcolor=blue}

duplicate-label

Flag a label key defined more than once in the same label namespace – within one file, or across files that share a document when a project view is available. LaTeX itself only warns and silently keeps the last definition. Within a file, a warning requires a prior definition in the same conditional branch or an enclosing context. Separate conditional tests are treated as uncertain and do not trigger a warning. Recognizes \if...\else...\fi and common macros with complete braced arguments, including \ifthenelse, \iftoggle, and \IfFileExists. Predicates are not evaluated, and coverage across branches is not combined. Cross-file checks use label namespaces. No autofix: resolving a collision (rename vs delete) is the author’s call.

This rule is enabled by default.

The same key defined twice in one file:

\section{One}\label{sec:x}
\section{Two}\label{sec:x}
warning: duplicate-label
 --> example.tex:2:14
  |
1 | \section{One}\label{sec:x}
  |                     ----- first definition of `sec:x`
2 | \section{Two}\label{sec:x}
  |              ^^^^^^^^^^^^^ label `sec:x` is defined more than once

deprecated-command

Flag the obsolete two-letter font switches (\bf, \it, \rm, \sf, \tt, \sc, \sl) that LaTeX 2e superseded with the \...series/\...shape/\...family declarations. \em is not flagged; it is still the supported emphasis switch. A name the file redefines (\renewcommand{\sl}{…}, \def\rm{…}) is the user’s macro, not the switch, so it is not flagged anywhere. The autofix swaps just the control word (\bf -> \bfseries), leaving any following text untouched, so it is correct by construction; it is withheld where the switch is merely referenced (\let\x\rm, \ifx\rm\y).

This rule is enabled by default.

An obsolete two-letter font switch:

{\bf important}
warning: deprecated-command
 --> example.tex:1:2
  |
1 | {\bf important}
  |  ^^^ `\bf` is deprecated; use `\bfseries`

After applying the fix:

{\bfseries important}

deprecated-suppression-syntax

Flag the retired % badness-ignore <rule> and % badness-ignore-file [<rule>] suppression spellings, which remain accepted for compatibility but are no longer documented. The Safe autofix rewrites only the family and verb to % badness-lint skip <rule> or % badness-lint skip-file [<rule>]; the selector and reason remain byte-for-byte unchanged, and the edit stays entirely inside a comment. This meta diagnostic is not silenced by the retired directive it reports; use [lint].ignore to disable the rule deliberately.

This rule is enabled by default.

A retired suppression directive:

% badness-ignore deprecated-command: legacy source
{\bf text}
warning: deprecated-suppression-syntax
 --> example.tex:1:3
  |
1 | % badness-ignore deprecated-command: legacy source
  |   ^^^^^^^^^^^^^^ retired suppression syntax; use `% badness-lint skip` instead

After applying the fix:

% badness-lint skip deprecated-command: legacy source
{\bf text}

missing-nonbreaking-space

Flag a plain space where a TeX tie (~) belongs, before a command whose output a line break would orphan: a bare-number reference (Figure \ref{x}, \eqref, \pageref) or a bracketed citation (see \cite{a}, \parencite, \autocite). A tie keeps the reference on the same line. Self-describing references (\autoref, \cref) and textual citations (\textcite, \citet) are not flagged – they emit their own noun, so a break orphans nothing. Both a same-line space and a single source line break before the command are flagged (a blank line is not – that starts a new paragraph). For a same-line space the fix is unsafe – inserting a tie changes line breaking – so --fix leaves it alone; --unsafe-fixes and the editor code action apply it. A line break is report-only: rewriting the newline to ~ would join the two lines, a reflow the formatter owns.

This rule is enabled by default.

A plain space where a tie belongs before a cross-reference:

see Figure \ref{fig:plot}
warning: missing-nonbreaking-space
 --> example.tex:1:11
  |
1 | see Figure \ref{fig:plot}
  |           ^ missing non-breaking space before `\ref`; use a tie `~` so the reference stays on the same line

obsolete-environment

Flag math environments the community has superseded, naming the modern replacement in the message. The canonical case is eqnarray, which amsmath replaced with align decades ago (it mis-spaces relations and is a perennial l2tabu warning). The autofix renames the \begin/\end pair in place, leaving the body untouched, so it is correct by construction.

This rule is enabled by default.

The superseded eqnarray environment:

\begin{eqnarray}
  a &=& b
\end{eqnarray}
warning: obsolete-environment
 --> example.tex:1:7
  |
1 | \begin{eqnarray}
  |       ^^^^^^^^^^ `eqnarray` is obsolete; use `align`

After applying the fix:

\begin{align}
  a &=& b
\end{align}

primitive-command

Flag raw plain-TeX primitives discouraged in LaTeX source, naming the LaTeX construct that supersedes each one (ChkTeX 41, lacheck, l2tabu). A sibling of deprecated-command, which covers the obsolete font switches. Most primitives are reported only: their LaTeX replacement restructures arguments (a \over b becomes \frac{a}{b}, \centerline{x} becomes a \centering declaration or a center environment), so no single textual edit can rewrite them correctly by construction. A few carry a Safe autofix — a 1:1 control-word swap for a primitive whose LaTeX form is a single meaning-identical token (\sb/\sp become _/^); the swap replaces just the control word, so it stays lossless and meaning-preserving, and is withheld where the primitive is merely referenced (\let\x\sp, \ifx\sp\y). A name the file redefines (\renewcommand\sp{…}) is the user’s macro, not the primitive, so it is not flagged anywhere. The implicit braces \bgroup and \egroup are not flagged: replacing them with literal braces can change macro argument and definition boundaries.

This rule is enabled by default.

A plain-TeX fraction primitive (report-only; the LaTeX form restructures its operands):

$a \over b$
warning: primitive-command
 --> example.tex:1:4
  |
1 | $a \over b$
  |    ^^^^^ `\over` is a raw TeX primitive; use `\frac{...}{...}`

The plain-TeX subscript alias, carrying a safe swap to _:

$x\sb2$
warning: primitive-command
 --> example.tex:1:3
  |
1 | $x\sb2$
  |   ^^^ `\sb` is a raw TeX primitive; use `_`

After applying the fix:

$x_2$

dollar-display-math

Flag plain-TeX $$...$$ display math. $$ is a TeX primitive that bypasses amsmath spacing hooks and breaks fleqn/\everydisplay, so LaTeX steers users to \[...\]. The autofix swaps the delimiters in place and leaves the body untouched, so it parses and stays lossless; it is withheld when the display math is unclosed.

This rule is enabled by default.

Plain-TeX display math:

$$a + b = c$$
warning: dollar-display-math
 --> example.tex:1:1
  |
1 | $$a + b = c$$
  | ^^ `$$…$$` is plain-TeX display math; use `\[…\]`

After applying the fix:

\[a + b = c\]

ellipsis

Flag a literal run of three or more periods (...) where a real ellipsis command belongs. ... sets three tight full stops; LaTeX’s ellipsis commands set correctly spaced dots. In text the fix is a safe swap to \dots (a space is added before a following letter so the control word cannot glue onto the next word). In math \ldots (baseline, for comma lists) and \cdots (centered, for operator chains) are not interchangeable, so the fix is unsafe: it guesses from the neighboring atoms – an operator or relation picks \cdots, otherwise \ldots – and applies only under --unsafe-fixes or as an editor code action. Comments and verbatim are never touched.

This rule is enabled by default.

Literal dots in text:

See Chapter 2, 3, ... for details.
warning: ellipsis
 --> example.tex:1:19
  |
1 | See Chapter 2, 3, ... for details.
  |                   ^^^ literal `...` ellipsis; use `\dots`

After applying the fix:

See Chapter 2, 3, \dots for details.

Literal dots in a math sum (an operator neighbor picks \cdots):

$a_1 + ... + a_n$
warning: ellipsis
 --> example.tex:1:8
  |
1 | $a_1 + ... + a_n$
  |        ^^^ literal `...` ellipsis; use `\cdots` in math (`\ldots` for lists, `\cdots` for operator chains)

expl3-invalid-message-parameter

Flag #5 through #9 in either text argument of a literal expl3 msg_new, msg_set, or msg_gset definition. Messages accept only #1 through #4. Escaped hashes and parameters belonging to an enclosing function definition are distinguished from message parameters. Checks cover recognized executable calls and unexpanded function bodies; stored token lists, expanded text arguments, and unresolved calls stay silent. Report-only: the intended message argument is unknown.

This rule is enabled by default.

A message refers to a fifth parameter:

\ExplSyntaxOn
\msg_new:nnn { demo } { bad-value } { Invalid~value:~#5 }
\ExplSyntaxOff
warning: expl3-invalid-message-parameter
 --> example.tex:2:54
  |
2 | \msg_new:nnn { demo } { bad-value } { Invalid~value:~#5 }
  |                                                      ^^ invalid expl3 message parameter `#5`; messages accept only `#1` through `#4`

expl3-protected-predicate

Flag a protected expl3 conditional definition whose literal condition list requests a p predicate. Predicates must be expandable, which protection prevents. The new, set, and gset families are checked in recognized executable code, including unexpanded function bodies. Computed condition lists and unresolved calls stay silent. Report-only: choosing between protection and the predicate changes the function’s API or meaning.

This rule is enabled by default.

A protected conditional requests a predicate:

\ExplSyntaxOn
\prg_new_protected_conditional:Nnn \demo_ready: { p, TF }
  { \prg_return_true: }
\ExplSyntaxOff
warning: expl3-protected-predicate
 --> example.tex:2:51
  |
2 | \prg_new_protected_conditional:Nnn \demo_ready: { p, TF }
  |                                                   ^ a protected expl3 conditional cannot define an expandable `p` predicate

expl3-variant-type

Flag incompatible or deprecated argument-type conversions in literal expl3 variant-generation calls. A shorter variant inherits the original suffix. Unchanged letters are valid, N may become c, and n may become o, V, v, f, e, or x. Conversions between these two families are deprecated; other changes are incompatible. Checks cover recognized executable calls and unexpanded function bodies, not stored token lists or unresolved expansion. Report-only: the intended signature is the author’s decision.

This rule is enabled by default.

A variant cannot add arguments:

\ExplSyntaxOn
\cs_generate_variant:Nn \demo_use:n { nn }
\ExplSyntaxOff
warning: expl3-variant-type
 --> example.tex:2:39
  |
2 | \cs_generate_variant:Nn \demo_use:n { nn }
  |                                       ^^ incompatible expl3 variant conversion from `n` to `nn`

Converting a single-token argument to a token-list argument is deprecated:

\ExplSyntaxOn
\cs_generate_variant:Nn \demo_use:Nn { nn }
\ExplSyntaxOff
warning: expl3-variant-type
 --> example.tex:2:40
  |
2 | \cs_generate_variant:Nn \demo_use:Nn { nn }
  |                                        ^^ deprecated expl3 variant conversion from `Nn` to `nn`

extra-alignment-tab

Flags a row in a built-in tabular, tabular*, or array environment that consumes more columns than its column preamble declares. LaTeX cannot place the overflowing cell and reports an extra alignment tab. Short rows are valid and are not flagged. Custom column types and dynamic \multicolumn spans are left alone when their width cannot be established statically. No autofix is offered because either the row or the preamble may be wrong.

This rule is enabled by default.

A row that exceeds the declared table width:

\begin{tabular}{ll}
  a & b & c \\
\end{tabular}
error: extra-alignment-tab
 --> example.tex:2:9
  |
1 | \begin{tabular}{ll}
  |                 -- table preamble declares 2 columns
2 |   a & b & c \\
  |         ^ row uses at least 3 columns, but the table preamble declares 2

extra-math-linebreak

Flag a plain \\ at the end of an align, alignat, flalign, gather, or multline environment, including their starred forms, or immediately after \intertext{...} or \shortintertext{...} in an environment that supports intertext. These breaks add an empty row, increasing vertical space and potentially adding an equation number. Breaks before intertext, starred breaks, explicit spacing arguments, subsidiary environments such as aligned, and locally redefined environments or intertext commands are left alone. The fix deletes only the offending \\, preserving comments and surrounding whitespace. It is unsafe because it changes typeset spacing and potentially numbering; use --fix --unsafe-fixes or an explicit editor action.

This rule is enabled by default.

A final linebreak adds an empty equation row:

\begin{align}
  a &= b \\
\end{align}
warning: extra-math-linebreak
 --> example.tex:2:10
  |
2 |   a &= b \\
  |          ^^ final linebreak adds an empty math row

Intertext already separates the surrounding equation rows:

\begin{align*}
  a &= b \\
  \intertext{and therefore}\\
  c &= d
\end{align*}
warning: extra-math-linebreak
 --> example.tex:3:28
  |
3 |   \intertext{and therefore}\\
  |                            ^^ linebreak after `\intertext` adds an empty math row

hard-coded-reference

Flag a literal cross-reference written in prose – Figure 3, Table~1, Section 2 – instead of \ref/\cref to a \label (textidote sh:hcfig/hctab/hcsec). Hard-coding the number defeats LaTeX’s automatic numbering: renumbering a float or reordering sections silently breaks the reference and drops the hyperlink. The rule is report-only – the correct rewrite needs the label the number refers to, which is not in the text, so no autofix is offered. To stay conservative it fires only for a capitalized reference word (Figure, Table, Section, Eq., …) matched as a whole word and directly followed, across one space or a tie ~, by an arabic number; plurals, lowercase, Figure~\ref{x}, and Figure three are left alone. It also skips a citation locator (\cite[Section~8.1]{...}, a reference into external work), an environment title (\begin{thm}[Conway's Theorem 0], a proper name), and an \item[label] description-list caption (\item[Part 3.]). It never touches math, comments, or verbatim.

This rule is enabled by default.

A hard-coded figure number instead of a cross-reference:

See Figure 3 for the results.
warning: hard-coded-reference
 --> example.tex:1:5
  |
1 | See Figure 3 for the results.
  |     ^^^^^^^^ hard-coded reference `Figure 3`; use `\ref`/`\cref` to a `\label` so the number stays in sync

Even tied with ~, the number is still hard-coded:

Table~1 lists the parameters.
warning: hard-coded-reference
 --> example.tex:1:1
  |
1 | Table~1 lists the parameters.
  | ^^^^^^^ hard-coded reference `Table~1`; use `\ref`/`\cref` to a `\label` so the number stays in sync

indented-docstrip-guard

Flag a syntactically complete %<…> marker in a .dtx file when it is preceded only by horizontal whitespace on its physical line. Docstrip recognizes guards only at column zero, so an indented near match is an ordinary comment and does not select or delimit generated code. No autofix is offered because activating a guard can change generated files.

This rule is enabled by default.

A docstrip guard indented by one space:

 %<*package>
\ProvidesPackage{example}
 %</package>
warning: indented-docstrip-guard
 --> example.dtx:1:2
  |
1 |  %<*package>
  |  ^^^^^^^^^^^ docstrip guards are recognized only at column zero
warning: indented-docstrip-guard
 --> example.dtx:3:2
  |
3 |  %</package>
  |  ^^^^^^^^^^^ docstrip guards are recognized only at column zero

inert-suppression

Flag a suppression directive that cannot take effect: skip with no following construct, on with no matching off, or a directive written on a .dtx documentation-margin line, where % is typeset prose rather than a comment. Also flag an off region left open at EOF; it currently suppresses through the end of the file, but the missing closer is usually accidental. Report-only: moving, deleting, or closing the directive requires knowing the boundary the author intended. Inline suppressions cannot hide this meta diagnostic; use [lint].ignore to disable the rule deliberately.

This rule is enabled by default.

An on directive with no matching open region does nothing:

% badness-lint on deprecated-command
{\bf text}
warning: inert-suppression
 --> example.tex:1:1
  |
1 | % badness-lint on deprecated-command
  | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ `on` has no matching `off`, so this directive closes no region

invalid-macrocode-frame

Flag a .dtx macrocode or macrocode* closing frame unless exactly four spaces separate its column-one % from \end{…}. The doc package scans for that literal physical delimiter, so a near match does not close the code chunk even though it looks like an ordinary environment to Badness. The safe autofix replaces only the malformed horizontal space with the required four spaces.

This rule is enabled by default.

A macrocode closer with only three spaces after %:

%    \begin{macrocode}
\def\example{value}
%   \end{macrocode}
error: invalid-macrocode-frame
 --> example.dtx:3:2
  |
3 | %   \end{macrocode}
  |  ^^^ `macrocode` closing frame requires exactly four spaces after `%`

After applying the fix:

%    \begin{macrocode}
\def\example{value}
%    \end{macrocode}

straight-quotes

Flag a literal ASCII double quote (") used for quotation. In LaTeX a straight " always sets a closing double quote, so an opening one comes out backwards; the correct forms are `` (two backticks) to open and '' (two apostrophes) to close. A quotation is reported once, spanning both quotes, and its fix rewrites the pair in one atomic edit – so a single editor code action repairs it from either end. A quote left unpaired (no closer before the paragraph ends) reports on its own. The fix is unsafe: it infers direction from context – a quote preceded by whitespace, a line break, an opening delimiter ((, [, {), a backtick, or the start of the document opens, anything else closes – and applies only under --unsafe-fixes or as an editor code action, since the guess can flip the typeset glyph. Single straight quotes (') are left alone (they are legitimately apostrophes), and comments, verbatim, math, TeX hex constants ("2D), and \pdfmapline font maps are never touched.

This rule is enabled by default.

Straight ASCII double quotes around a phrase:

He said "hello world" to me.
warning: straight-quotes
 --> example.tex:1:9
  |
1 | He said "hello world" to me.
  |         ^^^^^^^^^^^^^ straight double quotes; use `` `` `` (opening) and `''` (closing)

An opening quote after a parenthesis:

("quoted")
warning: straight-quotes
 --> example.tex:1:2
  |
1 | ("quoted")
  |  ^^^^^^^^ straight double quotes; use `` `` `` (opening) and `''` (closing)

swallowed-space

Flag a text-producing control word directly followed by a space that TeX eats, gluing the macro’s output to the next word (\LaTeX is renders “LaTeXis”) (ChkTeX 1). When TeX tokenizes a control word it discards following spaces, so the space never reaches the output. To stay conservative the rule fires only for a curated set of argument-less TeX-family logos (\LaTeX, \TeX, \BibTeX, …), only in text mode, and only when the next token is a word beginning with an alphanumeric character – a following period (\LaTeX . -> “LaTeX.”) is what the author wanted. The fix inserts {} after the control word (\LaTeX{} is), ending the macro name so the space survives; it is unsafe because it changes the typeset output, so --fix leaves it alone while --unsafe-fixes and the editor code action apply it.

This rule is enabled by default.

A logo swallows the following space, gluing it to the next word:

We used \LaTeX to typeset this.
warning: swallowed-space
 --> example.tex:1:15
  |
1 | We used \LaTeX to typeset this.
  |               ^ `\LaTeX` swallows the following space; add `{}` (`\LaTeX{}`) or `\ ` so it prints

space-before-command

Flag a plain space directly before a command that should hug the preceding word – \footnote, \footnotemark, \index, \label (ChkTeX 24/42). A space before \footnote sets a spurious space before the footnote mark (word \footnote{x} -> “word ¹”); a space before a zero-width \index/\label leaves a stray inter-word gap that can shift the recorded page. The fix deletes the space. It is unsafe – removing the space changes the typeset spacing – so --fix leaves it alone while --unsafe-fixes and the editor code action apply it. To stay conservative only the same-line WORD SPACE \cmd shape is flagged (a space at line start or after a brace is left alone), and math is skipped (an inter-token space is insignificant there), covering both $…$ and math environments like equation/align. For the zero-width \index/\label the fix is withheld unless the group is trailed by whitespace, a newline, or paragraph end, since otherwise the leading space is a real interword space to the following content.

This rule is enabled by default.

A space before a footnote sets a spurious space before the mark:

This is important \footnote{See the appendix.}
warning: space-before-command
 --> example.tex:1:18
  |
1 | This is important \footnote{See the appendix.}
  |                  ^ spurious space before `\footnote`; delete it so no stray space is typeset before the command

dash-length

Flag a dash of the wrong length for its context (ChkTeX 8). LaTeX sets a hyphen from -, an en dash from --, and an em dash from ---. Between two numbers a range takes an en dash, so 5-10 or 5---10 is flagged with an unsafe fix to -- (unsafe because it changes the typeset glyph and a hyphen between numbers is occasionally intentional). Between two words an en dash (--) is almost always a mistake, but whether a hyphen or an em dash was meant is ambiguous, so it is reported without a fix – except when it joins coordinate proper names (Barzilai--Borwein, Newton--Raphson), detected by an uppercase first letter on either flank, where the en dash is correct and the finding is suppressed. To stay conservative the rule only inspects a dash run that sits inside a single word with content on both sides and is the only dash run in that word, so dates (2020-01-15), ISBNs, spaced dashes, and option flags (--verbose) are left alone. Column spans in rule commands (\cline{1-3}, \cmidrule(lr){2-3}) and key arguments (\label{fig:1-3}, \cite{smith2020-1}) are specs and opaque identifiers rather than typeset ranges, so they are skipped too. The same applies to angle-delimited command and environment specifications such as Beamer’s \item<1-2> and \begin{onlyenv}<2-3>. Comments, verbatim, and math are never touched.

This rule is disabled by default; enable it with select.

A hyphen where a number range wants an en dash:

See pages 5-10 for the proof.
warning: dash-length
 --> example.tex:1:12
  |
1 | See pages 5-10 for the proof.
  |            ^ hyphen between numbers; use an en dash `--` for a number range

An en dash between words (ambiguous, so reported without a fix):

A well--known result.
warning: dash-length
 --> example.tex:1:7
  |
1 | A well--known result.
  |       ^^ en dash `--` between words; use a hyphen `-` for a compound or an em dash `---` for a break

times-variable

Flag a literal x used as a multiplication sign between two numbers, such as 640x200 or 3x3 (ChkTeX 29). TeX sets that x as an italic letter rather than the \times cross, so it reads wrong. The rule only fires when the whole word is digits x digits – one lowercase x with ASCII digits on both sides and nothing else – so ordinary words (matrix), spaced products (n x m), hex literals (0xFF, 0x12), and key arguments such as \label{fig:3x3} or \ref{fig:3x3} (where the x is part of an opaque identifier) are left alone. The fix is unsafe (a bare x between numbers is usually a cross but occasionally a real variable): inside math it rewrites the x to \times, and in text it wraps it as $\times$ so the result still compiles. So --fix leaves it alone; --unsafe-fixes and the editor code action apply it.

This rule is enabled by default.

A literal x as a multiplication sign in text (fixed to $\times$):

A 640x200 pixel image.
warning: times-variable
 --> example.tex:1:6
  |
1 | A 640x200 pixel image.
  |      ^ literal `x` as a multiplication sign between numbers; use `\times` for a cross

The same inside math mode (fixed to \times):

The grid is $640x200$ cells.
warning: times-variable
 --> example.tex:1:17
  |
1 | The grid is $640x200$ cells.
  |                 ^ literal `x` as a multiplication sign between numbers; use `\times` for a cross

math-operator-name

Flag a bare log-like function name (sin, cos, log, lim, and the rest of the LaTeX/amsmath set) written in math mode without its backslash, so TeX sets it as italic variables instead of the upright \sin operator with correct spacing (ChkTeX 35). It fires when the name starts a WORD and ends at a word boundary, catching both $sin x$ and the glued $sin(x)$, while leaving words that merely begin with one (since) alone and preferring the longest match (sinh over sin). To stay conservative it only fires inside math mode, never in a subscript or superscript, where max in x_{max} is almost always a label, and never inside a text-domain or unknown argument. The fix inserts the backslash (sin -> \sin); it is unsafe because it changes the typeset output (upright glyph and operator spacing) and a bare sin is occasionally a real product, so --fix leaves it alone while --unsafe-fixes and the editor code action apply it.

This rule is enabled by default.

A bare function name typesets as italic variables:

$sin x + cos x = 1$
warning: math-operator-name
 --> example.tex:1:2
  |
1 | $sin x + cos x = 1$
  |  ^^^ bare `sin` in math typesets as italic variables; use `\sin`
warning: math-operator-name
 --> example.tex:1:10
  |
1 | $sin x + cos x = 1$
  |          ^^^ bare `cos` in math typesets as italic variables; use `\cos`

It fires through the glued f(x) form too:

The limit $lim(x)$ diverges.
warning: math-operator-name
 --> example.tex:1:12
  |
1 | The limit $lim(x)$ diverges.
  |            ^^^ bare `lim` in math typesets as italic variables; use `\lim`

makeat-macro

Flag a macro whose name contains @ (\foo@bar, \p@, \@ifnextchar) used outside a \makeatletter/\makeatother region. There @ has its ordinary catcode, so it cannot be part of a control word: \foo@bar is read as \foo followed by the text @bar, not as a call to the internal macro \foo@bar. Usually the enclosing \makeatletter/\makeatother was forgotten. Because the formatter’s lexer already tracks \makeatletter state, this is decided exactly – an in-region name lexes as one token and is never flagged; only the split out-of-region form (control word abutting an @-word, or \@ abutting a letter-word) is. Report-only: a correct fix would mean wrapping the use in \makeatletter/\makeatother, not a tight local edit, so no autofix is offered. The end-of-sentence \@ (as in NASA\@.) is not flagged.

This rule is enabled by default.

An internal @ macro used without \makeatletter:

\my@command
warning: makeat-macro
 --> example.tex:1:1
  |
1 | \my@command
  | ^^^^^^^^^^^ `\my@command` uses `@` in a macro name outside a `\makeatletter` region; `@` is not a letter here, so this reads as `\my` followed by the text `@command`

A leading-@ macro (a \@-prefixed internal) outside a region:

\@ifstar{\StarredForm}{\PlainForm}
warning: makeat-macro
 --> example.tex:1:1
  |
1 | \@ifstar{\StarredForm}{\PlainForm}
  | ^^^^^^^^ `\@ifstar` uses `@` in a macro name outside a `\makeatletter` region; `@` is not a letter here, so this reads as `\@` followed by the text `ifstar`

sectioning-level-jump

Flag a structural heading that descends more than one level below the preceding structural heading – \section straight to \subsubsection, skipping \subsection (textidote’s sh:secskip). The active ladder follows the document class: \chapter is included only for classes known to provide it or when the source uses it, while unknown classes conservatively omit it. \paragraph and \subparagraph are transparent because technical papers commonly use them as run-in labels rather than outline subdivisions. Only downward jumps are flagged – climbing back up and repeated headings at one level are normal. The comparison is relative to the previous structural heading, never an absolute top level. Report-only: repairing a skip is a structural choice for the author, not a correct-by-construction edit.

This rule is enabled by default.

A heading that drops two levels at once (skipping \subsection):

\section{Introduction}
\subsubsection{Details}
warning: sectioning-level-jump
 --> example.tex:2:1
  |
2 | \subsubsection{Details}
  | ^^^^^^^^^^^^^^ `\subsubsection` skips a sectioning level after `\section` (expected `\subsection`)

missing-required-argument

Flag a command invoked with fewer {…} groups than the required arity in its curated built-in signature (ChkTeX warning 14, decided on the parse tree and signature database rather than line heuristics). TeX also accepts unbraced single-token arguments (\frac12), so the rule stays silent whenever a following token could still supply the missing argument and fires only at a hard boundary: the end of the enclosing group, math shell, or environment, an alignment &, a \\ line break, a blank line, or the end of the file. Contexts where a bare command is deliberate are skipped – macro-definition bodies (\newcommand{\bold}{\textbf}), arguments of unknown commands, standalone {…} scope groups, \let-style alias forms, and names the file itself redefines. Curated environment-local signatures take precedence over global signatures: inside parts, exam’s \part takes only optional points. Report-only: the missing argument’s content is the author’s to write, so no fix is correct by construction.

This rule is enabled by default.

A fraction missing its denominator:

$\frac{1}$
warning: missing-required-argument
 --> example.tex:1:2
  |
1 | $\frac{1}$
  |  ^^^^^ `\frac` is missing 1 of its 2 required arguments

A command left bare at the end of a group, with nothing to take:

\emph{see \textbf}
warning: missing-required-argument
 --> example.tex:1:11
  |
1 | \emph{see \textbf}
  |           ^^^^^^^ `\textbf` is missing its required argument

undefined-ref

Flag a \ref-family reference to a label defined nowhere in the document. Sound only when the label namespace is complete, so it stays silent unless the project view is closed (every include resolves to an analyzed file) and rooted. Inert on stdin or wherever no cross-file label resolution is available. No autofix.

This rule is enabled by default.

A reference to a label defined nowhere in the document:

\ref{sec:intro}
warning: undefined-ref
 --> example.tex:1:1
  |
1 | \ref{sec:intro}
  | ^^^^^^^^^^^^^^^ reference to undefined label `sec:intro`

undefined-citation

Flag a \cite-family key matching no entry in the document’s bibliography – the bibliographic analog of undefined-ref. Sound only over a closed, rooted namespace where every .bib resource resolves to an analyzed file; resource lookup honors BibTeX’s BIBINPUTS/TEXBIB search path. Suppressed entirely by a \nocite{*} wildcard (which marks every key as used). Inert without cross-file citation resolution. No autofix.

This rule is enabled by default.

A citation of a key that matches no bibliography entry:

\cite{knuth:1984}
warning: undefined-citation
 --> example.tex:1:1
  |
1 | \cite{knuth:1984}
  | ^^^^^^^^^^^^^^^^^ citation of undefined key `knuth:1984`

unreferenced-label

Flag a label definition unused by a \ref-family command anywhere in the document. A \eqref{A}--\eqref{D} range also uses labels between A and D when they occur in consecutive, singly labeled equation environments or numbered align and gather rows, including through literal included files with an unambiguous source order. Manual tags, suppressed numbers, and counter changes stop inference. Referencing a subequations group label also uses the labels in its enclosed math environments. The mirror of undefined-ref, and sound only when the label namespace is complete, so it stays silent unless the project view is closed (every include resolves to an analyzed file) and rooted. Inert on stdin or wherever no cross-file label resolution is available. Report-only: removing the dead label or adding a reference are both valid, so there is no autofix.

This rule is enabled by default.

A label that no \ref-family command ever targets:

\section{Intro}\label{sec:intro}
warning: unreferenced-label
 --> example.tex:1:16
  |
1 | \section{Intro}\label{sec:intro}
  |                ^^^^^^^^^^^^^^^^^ label `sec:intro` is never referenced

verbatim-trailing-text

Flag non-whitespace text after a verbatim-like environment’s \end{…} on the same line (ChkTeX warning 31). LaTeX closes a verbatim environment by scanning line by line to \end{verbatim} and then gobbling the rest of that line, so \end{verbatim} foo silently drops foo. Scoped to verbatim-like environments — read off the parse tree (an opaque VERBATIM_BODY, or a curated built-in verbatim name for the empty-body case) — because ordinary environments do not gobble their \end line. A trailing % comment is treated as trivia, not flagged. Report-only: whether to move or delete the swallowed text is the author’s call, so no fix is correct by construction.

This rule is enabled by default.

Text after \end{verbatim} is silently discarded by LaTeX:

\begin{verbatim}
sample
\end{verbatim} and more
warning: verbatim-trailing-text
 --> example.tex:3:16
  |
3 | \end{verbatim} and more
  |                ^^^^^^^^ text after `\end{verbatim}` on the same line is silently discarded

duplicate-package

Flag a package loaded more than once in the same file with \usepackage/\RequirePackage (which share one package namespace). LaTeX loads a given package only once; a second load is redundant and, when the options disagree, an option-clash error. A warning requires a prior load in the same conditional branch or an enclosing context (including an unconditional prior). Separate conditional tests are treated as uncertain and do not trigger a warning. Recognizes \if...\else...\fi and common macros with complete braced arguments, including \ifthenelse, \iftoggle, and \IfFileExists. Predicates are not evaluated, and coverage across branches is not combined. No autofix: removing a load can drop options the survivor lacks, and which load to keep is the author’s call. Class loads (\documentclass/\LoadClass) are a separate concern and are not flagged.

This rule is enabled by default.

The same package loaded twice:

\usepackage{amsmath}
\usepackage{amsmath}
warning: duplicate-package
 --> example.tex:2:1
  |
2 | \usepackage{amsmath}
  | ^^^^^^^^^^^^^^^^^^^^ package `amsmath` is loaded more than once

missing-provides

Flag a package or class source (.sty/.cls) that never identifies itself with the matching \ProvidesPackage/\ProvidesClass. Every well-formed package declares its identity so LaTeX can log it and honor date-based compatibility checks; a .sty carrying only \ProvidesClass (wrong kind) still counts as missing. The rule is inert for any other extension – a .tex has nothing to provide, and a .dtx hides its declaration inside guarded macrocode. No autofix: writing a correct \Provides… line (placement, date, version) is the author’s call.

This rule is enabled by default.

A package source with no self-identification (the docs are rendered against a .sty path):

\NeedsTeXFormat{LaTeX2e}
\RequirePackage{xcolor}
warning: missing-provides
 --> example.sty:1:1
  |
1 | \NeedsTeXFormat{LaTeX2e}
  | ^^^^^^^^^^^^^^^ package file lacks `\ProvidesPackage`

unknown-option

Flag a \usepackage/\RequirePackage option that the loaded package never declares with \DeclareOption, which LaTeX reports as an “Unknown option” error at compile time. Checked only against packages that are analyzed project files (a sibling .sty) — no option data ships for system packages — and only when the package’s declared set is trustworthy: a \DeclareOption* default handler, a key-value option processor (kvoptions, \ProcessKeyOptions, …), option forwarding, or an \input in the package silences the rule, as does a key=value option. Class loads (\documentclass) are not checked: an unknown class option is not an error, it becomes an unused global option. No autofix: dropping or renaming the option is the author’s call.

This rule is enabled by default.

With a sibling mypkg.sty:

\ProvidesPackage{mypkg}[2026/01/01 v1.0 Demo package]
\DeclareOption{draft}{}
\ProcessOptions\relax

Loading the sibling package with an option it never declares:

\usepackage[final]{mypkg}
warning: unknown-option
 --> example.tex:1:13
  |
1 | \usepackage[final]{mypkg}
  |             ^^^^^ unknown option `final` for package `mypkg`

redundant-script-braces

Flag braces around a single-token sub/superscript argument, which ^/_ bind without them (x^{2} is x^2). The autofix deletes the two braces and leaves the inner token untouched. It is withheld when dropping the braces would let the following character glue onto the argument and change meaning (x^{2}-3 stays braced — unspaced x^2-3 would re-lex 2-3 as one token; y_{\alpha}b stays braced — \alphab is one control word). It also leaves standard named math operators braced because commands such as \max are not valid unbraced script fields.

This rule is enabled by default.

Redundant braces around a single-token script argument:

$x^{2}$ and $y_{\alpha}$
help: redundant-script-braces
 --> example.tex:1:4
  |
1 | $x^{2}$ and $y_{\alpha}$
  |    ^^^ redundant braces around a single-token script argument
help: redundant-script-braces
 --> example.tex:1:16
  |
1 | $x^{2}$ and $y_{\alpha}$
  |                ^^^^^^^^ redundant braces around a single-token script argument

After applying the fix:

$x^2$ and $y_\alpha$

unclosed-math-delimiter

Flag a math opener the parser silently demoted to a plain token because no closer was reachable – a $ with no matching $, a \[/\( with no \]/\), or a \left with no \right. Such a shape is routine data in macro code (>{$} array columns, \expandafter\@tempa\[\@nil), so the parser tolerates it without a diagnostic; in prose it is almost always a dropped closer. To stay clear of the macro-code cases the rule is conservative: it reports only an opener in document prose, staying silent when it sits inside a brace group or optional argument (\newcommand{...}{$}, the >{$} column spec), an expl3 region, or a macrocode body. No autofix: the correction (insert a closer, or delete a stray opener) is ambiguous.

This rule is enabled by default.

An inline-math $ with no matching $:

Let $x = 1 be the base case.
warning: unclosed-math-delimiter
 --> example.tex:1:5
  |
1 | Let $x = 1 be the base case.
  |     ^ `$` has no matching `$` (unclosed inline math)

A display-math \[ with no matching \]:

The bound \[ x + y follows immediately.
warning: unclosed-math-delimiter
 --> example.tex:1:11
  |
1 | The bound \[ x + y follows immediately.
  |           ^^ `\[` has no matching `\]` (unclosed display math)

A \left with no matching \right:

$a + \left( b + c$
warning: unclosed-math-delimiter
 --> example.tex:1:6
  |
1 | $a + \left( b + c$
  |      ^^^^^ `\left` has no matching `\right`

label-before-caption

Flag a \label placed before the statement that establishes its intended counter: the outer \caption in a curated float (figure, table, and their starred forms), an explicit \captionof in a curated caption container (minipage), or the first \item in the standard numbered enumerate list. In either position, \label captures the previous \@currentlabel—usually an enclosing section number—so \ref silently prints an unrelated number. LaTeX gives no warning. The list case is limited to statement-level labels before the first item; labels after an item may belong to it, while itemize and description items do not step a reference counter. Attached custom item labels and complete Beamer overlay markers remain intact. The float case likewise skips labels nested in command arguments, and classifies nested counter steps conservatively. The fix moves the label just after the proven caption or item marker, and is Unsafe because it intentionally changes what \ref prints from an inferred intent.

This rule is enabled by default.

A \label above its \caption picks up the section counter, not the figure number:

\begin{figure}
  \includegraphics{plot}
  \label{fig:plot}
  \caption{A plot.}
\end{figure}
warning: label-before-caption
 --> example.tex:3:3
  |
3 |   \label{fig:plot}
  |   ^^^^^^^^^^^^^^^^ `\label` before the outer `\caption` in this `figure` does not capture the float number

A \label before the first \item has not seen the item counter step:

\begin{enumerate}
  \label{item:first}
  \item First
\end{enumerate}
warning: label-before-caption
 --> example.tex:2:3
  |
2 |   \label{item:first}
  |   ^^^^^^^^^^^^^^^^^^ `\label` before the first `\item` in this `enumerate` does not capture the item number

A \label above \captionof in a minipage likewise precedes the explicit counter step:

\begin{minipage}{\textwidth}
  \label{fig:plot}
  \captionof{figure}{A plot.}
\end{minipage}
warning: label-before-caption
 --> example.tex:2:3
  |
2 |   \label{fig:plot}
  |   ^^^^^^^^^^^^^^^^ `\label` before `\captionof` in this `minipage` does not capture the caption number

lonely-item

Flag \item written directly in the document environment, where LaTeX reports a lonely item because there is no list. The rule leaves items inside other environments, command arguments, low-level \list/\trivlist pairs, and standalone fragments alone because their list context may come from a custom definition or an including file. Report-only: the intended list type and boundaries cannot be inferred from the item.

This rule is enabled by default.

An item directly in the document body has no list:

\begin{document}
\item A
\end{document}
error: lonely-item
 --> example.tex:2:1
  |
2 | \item A
  | ^^^^^ `\item` has no enclosing list environment

Suppression

To suppress a rule at a single site, use a comment directive:

% badness-lint skip deprecated-command: legacy code, leave as-is
{\bf here}

The verb carries the scope. skip covers the next construct, off and on delimit a region, and skip-file covers the whole file wherever it sits:

% badness-lint off deprecated-command: legacy chapter
{\bf here}
{\it and here}
% badness-lint on deprecated-command

Naming the <id> is optional; leaving it out suppresses every rule over that same span. The : <reason> tail is optional everywhere.

% badness skip / off / on / skip-file do the same and turn off the formatter at the same time; see Formatting for the layout-only % badness-format spellings.

Parse diagnostics (rule id parse) are never suppressed by select/ignore.

BibTeX Linter Rules

badness lint runs a parallel set of built-in rules over each .bib file’s parse tree and reports a diagnostic for every finding. This page is the catalogue: one section per rule, keyed by its stable rule id. Bib rules share one id namespace with the LaTeX rules, so the same [lint] select/ignore (and --select/--ignore) target both.

Most rules are on by default. Each rule’s section states its default; enable an opt-in rule with select, or narrow the default set with select/ignore in the [lint] table (see the Configuration reference). Where a rewrite is unambiguous a rule carries an auto-fix: a safe fix (shown below as “After applying the fix”) is applied by badness lint --fix.

Each example below is linted live to produce its diagnostic and fixed output, so this page never drifts from the rules’ actual behavior.

duplicate-key

Flag a cite key defined by more than one entry in the same .bib file. Keys are compared case-insensitively, matching BibTeX, which silently keeps only one of the colliding entries; every definition after the first is flagged. No autofix: resolving the collision (rename vs delete) is the author’s call.

This rule is enabled by default.

The same cite key defined by two entries:

@misc{knuth84, title = {Draft}}
@book{knuth84, title = {Book}}
warning: duplicate-key
 --> references.bib:2:7
  |
2 | @book{knuth84, title = {Book}}
  |       ^^^^^^^ cite key `knuth84` is defined more than once

deprecated-suppression-syntax

Flag the retired @comment{badness-ignore <rule>} and @comment{badness-ignore-file [<rule>]} suppression spellings, which remain accepted for compatibility but are no longer documented. The Safe autofix rewrites only the family and verb to badness-lint skip or badness-lint skip-file; the selector, reason, delimiters, and remaining entry text stay byte-for-byte unchanged. This meta diagnostic is not silenced by the retired directive it reports; use [lint].ignore to disable the rule deliberately.

This rule is enabled by default.

A retired suppression directive in a structured comment:

@comment{badness-ignore unused-string: intentional}
@string{x = {X}}
warning: deprecated-suppression-syntax
 --> references.bib:1:10
  |
1 | @comment{badness-ignore unused-string: intentional}
  |          ^^^^^^^^^^^^^^ retired suppression syntax; use `@comment{badness-lint skip …}` instead

After applying the fix:

@comment{badness-lint skip unused-string: intentional}
@string{x = {X}}

inert-suppression

Flag a structured BibTeX suppression directive that cannot take effect: skip with no following entry, on with no matching off, or any badness-format directive, because the BibTeX formatter does not support format suppression. Also flag an off region left open at EOF; it currently suppresses through the end of the file, but the missing closer is usually accidental. Report-only: repairing the directive requires knowing the entry, boundary, or formatting policy the author intended. Inline suppressions cannot hide this meta diagnostic; use [lint].ignore to disable the rule deliberately.

This rule is enabled by default.

BibTeX recognizes the directive grammar, but its formatter has no suppression mechanism:

@comment{badness-format skip-file: preserve this file}
@book{key}
warning: inert-suppression
 --> references.bib:1:1
  |
1 | @comment{badness-format skip-file: preserve this file}
  | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ the BibTeX formatter does not support format suppression, so this directive does nothing

missing-required-field

Flag a regular entry lacking a field its type requires, per the biblatex data model. An alternation like date or year is satisfied by either, and classic-BibTeX aliases count (journal satisfies journaltitle). An entry type the built-in database does not know carries no signature and is never flagged. Report-only – field content cannot be invented.

This rule is enabled by default.

An @article without its required journaltitle:

@article{doe2020,
  author = {Doe, Jane},
  title  = {A study},
  year   = 2020
}
warning: missing-required-field
 --> references.bib:1:10
  |
1 | @article{doe2020,
  |          ^^^^^^^ entry `article` is missing required field `journaltitle`

unknown-field

Flag a field that is neither required nor optional for its entry type and carries no global field metadata – usually a typo, or data misplaced from another entry type. BibLaTeX silently ignores fields it does not know, so the mistake otherwise vanishes without a trace. Only entry types the built-in database knows are checked. Report-only – deleting the field would discard data.

This rule is enabled by default.

A typo’d field name (pubisher for publisher):

@book{turing50,
  author   = {Turing, Alan},
  title    = {A book},
  pubisher = {Elsevier},
  year     = 1950
}
warning: unknown-field
 --> references.bib:4:3
  |
4 |   pubisher = {Elsevier},
  |   ^^^^^^^^ unknown field `pubisher` on `book` entry

empty-field

Flag a field whose value is empty or whitespace-only (title = {}, note = ""). An empty field carries no data, and some styles still emit punctuation around it. The safe autofix deletes the field along with its separating comma.

This rule is enabled by default.

An empty note left behind by an edit:

@misc{knuth84,
  title = {Draft},
  note  = {}
}
warning: empty-field
 --> references.bib:3:3
  |
3 |   note  = {}
  |   ^^^^^^^^^^ field `note` is empty

After applying the fix:

@misc{knuth84,
  title = {Draft}
}

duplicate-field

Flag a field name appearing more than once on a single entry (names compared case-insensitively). BibTeX and Biber keep only one occurrence and silently discard the rest, so a duplicate is almost always a merge or copy-paste mistake; every occurrence after the first is flagged. When the repeated value is byte-identical to the kept one, a safe autofix deletes the redundant copy; when the values differ, which one wins is engine-dependent, so the finding is report-only.

This rule is enabled by default.

Two note fields with identical values – deleting the redundant copy is safe:

@misc{knuth84,
  note = {Draft},
  note = {Draft}
}
warning: duplicate-field
 --> references.bib:3:3
  |
3 |   note = {Draft}
  |   ^^^^ duplicate field `note` on `misc` entry

After applying the fix:

@misc{knuth84,
  note = {Draft}
}

Differing values are report-only (which copy the engine keeps is style-dependent, so dropping either would change meaning):

@misc{knuth84,
  note = {First draft},
  note = {Second draft}
}
warning: duplicate-field
 --> references.bib:3:3
  |
3 |   note = {Second draft}
  |   ^^^^ duplicate field `note` on `misc` entry

unused-string

Flag an @string macro defined in the file but never referenced by any field value. For the common self-contained .bib an unused macro is dead weight; in a multi-file bibliography it may be referenced from another .bib, so treat cross-file setups with care – cross-file @string resolution is not modeled yet. Report-only: deleting a definition is a meaning-level edit left to the author.

This rule is enabled by default.

A defined macro no field value references:

@string{cup = {Cambridge University Press}}
@book{turing50, title = {Draft}, publisher = {Springer}}
warning: unused-string
 --> references.bib:1:9
  |
1 | @string{cup = {Cambridge University Press}}
  |         ^^^ `@string` macro `cup` is defined but never used

undefined-string

Flag an @string macro used in a field value but defined nowhere in the file (the twelve month macros jan..dec are predefined). Usually a typo’d macro name or a missing @string definition; BibTeX errors on it at build time. In a multi-file bibliography the definition may live in another .bib, so a use resolved there is a false positive – cross-file @string resolution is not modeled yet. Report-only: the fix (define the macro or correct the name) is a meaning-level edit left to the author.

This rule is enabled by default.

A typo’d macro name (cpu for cup):

@string{cup = {Cambridge University Press}}
@book{turing50, title = {Draft}, publisher = cpu}
warning: undefined-string
 --> references.bib:2:46
  |
2 | @book{turing50, title = {Draft}, publisher = cpu}
  |                                              ^^^ `@string` macro `cpu` is used but never defined

title-capitalization

Flag an unprotected acronym or mid-word capital in a title-like field (title, booktitle, journaltitle, …). Many bibliography styles lowercase unprotected title text, so DNA renders as dna unless written {DNA}. Flagged are runs of two or more capitals and the camelCase brand pattern (a first capital mid-way through a lowercase-initial word, like iPhone); ordinary Title Case, name particles (McDonald), and mixed-case tokens (LaTeX) stay quiet, as does anything already inside a {...} group. Report-only – choosing what to protect is the author’s call.

This rule is enabled by default.

An unprotected acronym a title-lowercasing style would render as dna:

@article{watson53, title = {Molecular structure of DNA}}
warning: title-capitalization
 --> references.bib:1:52
  |
1 | @article{watson53, title = {Molecular structure of DNA}}
  |                                                    ^^^ unprotected capitals `DNA` in `title`; wrap in braces (`{DNA}`) to keep case under title-lowercasing styles

encoding-hints

Surface non-ASCII text in a field value as a hint (accented text is perfectly valid in a UTF-8 setup, hence not a warning). Raw non-ASCII renders correctly only when the file is UTF-8 and the document loads a matching input encoding (inputenc with pdfLaTeX, fontspec with Xe/LuaLaTeX); legacy toolchains may mangle it. Either confirm the encoding or use a LaTeX escape (\'e for é). Report-only – the right fix depends on the project’s toolchain.

This rule is enabled by default.

An accented name entered as raw UTF-8:

@article{erdos47, author = {Erdős, Paul}}
help: encoding-hints
 --> references.bib:1:32
  |
1 | @article{erdos47, author = {Erdős, Paul}}
  |                                ^ non-ASCII text `ő`; ensure the file is UTF-8 and the document loads an input encoding (inputenc/fontspec), or use a LaTeX escape

Suppression

BibTeX has no line-comment token, so per-site suppression rides a structured @comment entry instead of the LaTeX % directive. A plain directive suppresses one rule on the next entry:

@comment{badness-lint skip missing-required-field: publisher long gone}
@book{oldbook, title = {An Orphaned Book}}

The grammar is the LaTeX one, only the carrier differs. off and on delimit a region of entries, and skip-file covers the whole file wherever it sits:

@comment{badness-lint off missing-required-field: imported, incomplete by design}
@book{oldbook, title = {An Orphaned Book}}
@comment{badness-lint on missing-required-field}

Naming the <id> is optional; leaving it out suppresses every rule over that same span. Parse diagnostics (rule id parse) are never suppressed.

Benchmarks

These benchmarks compare the speed of Badness’s formatter, linter, and language server with other LaTeX tools, along with the language server’s memory use. The tools differ in formatting style, lint coverage, and editor features, so the timings alone cannot tell you which tool best suits your work.

Formatter

We compare badness with tex-fmt and latexindent on individual documents. Files that badness cannot format are excluded from every tool’s results.

  • badness: 0.23.0
  • tex-fmt: 0.5.7
  • latexindent: 3.24.7
  • backend: hyperfine (min runs: 3)
  • host: linux/x86_64, AMD Ryzen 9 7900 12-Core Processor
  • generated: 2026-09-21T14:09:42Z

Single-file results

Mean formatting time relative to badness on a logarithmic scale. The dashed line marks badness at 1; faster tools fall below it. Color distinguishes documents. Hover over a dot for times in milliseconds.
Data table

small.tex (baseline) (1233 bytes, 48 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness2.22801.55944.5042baseline
tex-fmt1.84471.40775.21161.2× faster
latexindent60.441158.346863.377327.1× slower

cv.tex (6273 bytes, 275 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness2.89682.19757.0406baseline
tex-fmt1.88361.41245.28851.5× faster
latexindent67.654864.680672.638923.4× slower

masters_dissertation.tex (95383 bytes, 2458 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness21.243419.953424.0947baseline
tex-fmt2.51752.11865.62268.4× faster
latexindent1619.68121616.03991624.468476.2× slower

phd_dissertation.tex (730369 bytes, 27482 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness206.0519198.8200214.6120baseline
tex-fmt9.34488.576312.387222.0× faster
latexindent25787.035625586.912225980.9774125.1× slower

Whole-project results

This comparison measures the time to check formatting across the .tex files of kks32/phd-thesis-template. Both tools compute the formatted output without writing changes to disk. latexindent is omitted because it has no recursive directory mode.

Mean time to check formatting across the thesis project, relative to badness on a logarithmic scale. The dashed line marks badness at 1; faster tools fall below it. Hover over a dot for times in milliseconds.
Data table

project (12 files) (47190 bytes, 1005 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness6.05454.89069.2864baseline
tex-fmt2.54432.14726.06912.4× faster

Linter

We compare badness lint with lacheck and chktex on the same individual documents. Each linter checks for a different set of problems. Neither comparison tool has a recursive directory mode, so this benchmark covers individual files only.

  • badness: 0.23.0
  • lacheck: 1.30
  • chktex: v1.7.9
  • backend: hyperfine (min runs: 3)
  • host: linux/x86_64, AMD Ryzen 9 7900 12-Core Processor
  • generated: 2026-09-21T14:09:42Z
Mean linting time relative to badness on a logarithmic scale. The dashed line marks badness at 1; faster tools fall below it. Color distinguishes documents. Hover over a dot for times in milliseconds.
Data table

small.tex (baseline) (1233 bytes, 48 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness2.52831.90144.2127baseline
lacheck6.28305.73617.50642.5× slower
chktex38.613833.097954.023015.3× slower

cv.tex (6273 bytes, 275 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness2.93122.35994.5520baseline
lacheck6.55335.92069.31712.2× slower
chktex37.978433.540346.252813.0× slower

masters_dissertation.tex (95383 bytes, 2458 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness58.965854.145770.2911baseline
lacheck8.36817.719510.95667.0× faster
chktex41.089736.149350.04361.4× faster

phd_dissertation.tex (730369 bytes, 27482 lines)

ToolMean (ms)Min (ms)Max (ms)Relative
badness201.2823191.1012207.7070baseline
lacheck21.000919.898323.03989.6× faster
chktex68.787462.448681.91852.9× faster

Language Server

We compare Badness with TexLab by opening the same five documents in the complete thesis project. The benchmark measures how long each server takes to start, respond to editor requests, and finish background work, as well as how much memory it uses.

  • Badness: 0.23.0
  • TexLab: 5.26.0
  • corpus: phd-thesis-template @ 3ce347686d75 (v2.4, 16 source files, 279004 bytes)
  • session: 3 fresh runs per server; 5 open files (37773 bytes)
  • sampling: every 0.15 s; quiet for 5.0 s; 60 s phase timeout
  • host: Linux x86_64, AMD Ryzen 9 7900 12-Core Processor (61.9 GiB RAM)
  • generated: 2026-09-21T14:17:30Z
  • navigation target: Aup91 in Chapter1/chapter1.tex at line 19

Speed

The startup measurements cover three waits:

  • Initialize measures the server’s response to the editor’s initialization request.
  • Workspace ready measures the time from process start until background indexing settles.
  • Open files ready measures the time from opening the documents until diagnostics arrive and background work settles.

After startup, we time requests for document symbols, hover information, definitions, references, and renaming. The chart shows the median response times. Tooltips and expandable tables include the 95th percentile (p95) and the number of results returned, which can differ between servers.

Readiness

Time until each server is ready, on a logarithmic scale. Dots show medians across fresh sessions, and color distinguishes the waits. Lower is faster. Hover over a dot for minimum and maximum times.
Data table
ServerWaitMedianMinMax
BadnessInitialize1.000 ms1.000 ms1.000 ms
BadnessWorkspace ready152.000 ms152.000 ms152.000 ms
BadnessOpen files ready33.000 ms33.000 ms35.000 ms
TexLabInitialize1.000 ms1.000 ms7.000 ms
TexLabWorkspace ready152.000 ms152.000 ms605.000 ms
TexLabOpen files ready40.000 ms38.000 ms42.000 ms

Warm requests

Warm request latency on a logarithmic scale. Dots show medians, and color distinguishes operations. Lower is faster. Hover over a dot for the 95th percentile, sample counts, and response details.
Data table
ServerRequestMedianp95Returned workSamples
BadnessDocument symbols0.224 ms0.315 ms5–16 symbols (median 14), 3 KiB180
BadnessHover0.115 ms0.154 ms1 result, 206 B180
BadnessGo to definition0.096 ms0.113 ms1 location in 1 file, 175 B60
BadnessFind references0.103 ms0.139 ms2 locations in 2 files, 346 B60
BadnessRename0.091 ms0.119 ms2 edits in 2 files, 418 B60
TexLabDocument symbols0.247 ms0.451 ms5–33 symbols (median 20), 5 KiB180
TexLabHover0.077 ms0.119 ms1 result, 207 B180
TexLabGo to definition0.068 ms0.080 ms1 location in 1 file, 371 B60
TexLabFind references0.065 ms0.077 ms2 locations in 2 files, 346 B60
TexLabRename0.066 ms0.080 ms2 edits in 2 files, 418 B60

Warm requests show medians across all samples, with p95 in tooltips and the table. Each target ran 20 measured rounds in each of 3 fresh sessions after 2 warmup rounds; symbols and hover span 3 files, while definition, references, and rename use Aup91 in Chapter1/chapter1.tex. Rename constructs the workspace edit but does not apply it.

Memory

The chart shows median memory use across three fresh sessions, including child processes. RSS counts resident memory, including shared pages in each process. The tooltips also show PSS, which divides shared pages among the processes using them to estimate their share of physical memory.

Median whole-process-tree RSS across fresh processes. Baseline follows initialization; settled follows the open-file workload. Peak is the largest sample through the timed requests. Tooltips also show PSS.
Data table
ServerMilestoneRSSPSSRelative settled RSS
BadnessBaseline9.1 MB7.2 MB-
BadnessSettled13.9 MB12.0 MBbaseline
BadnessPeak13.9 MB12.0 MB-
TexLabBaseline22.5 MB20.7 MB-
TexLabSettled29.6 MB27.7 MB2.13×
TexLabPeak29.6 MB27.7 MB-

Reproducibility

Run these commands from the repository root:

task bench:download  # Fetch the benchmark documents.
task bench          # Measure formatter and linter speed.
task bench:lsp      # Measure language-server speed and memory.

The scripts build badness in release mode. The formatter and linter comparison uses the tools available on PATH and skips any that are missing. Install hyperfine and jq for timing statistics; without them, the script uses a shell loop that reports only mean times. The language-server benchmark requires Linux, Python 3, and texlab. task bench:memory is an alias for task bench:lsp.

The commands write benches/benchmark_results.json and benches/memory_results.json. These committed files supply the charts, machine details, and tool versions shown above. Building the documentation reads these files without running the benchmarks. Neither benchmark runs in CI.

Documents

The individual documents are a committed small.tex baseline and three files from a pinned tex-fmt release: cv.tex, masters_dissertation.tex, and phd_dissertation.tex. The thesis project comes from a pinned revision of kks32/phd-thesis-template. benches/documents/download.sh records both pins.

The formatter and linter benchmarks skip any document that badness cannot format. For the project comparison, the script copies a fixed set of .tex files into a temporary directory, excluding unsupported files from both tools. This gives both formatters the same input files without interference from Git ignore rules. The language servers use the complete project, including its class, style, bibliography, and image files.

Formatter and linter commands

For individual documents, each formatter reads from standard input and writes to standard output:

ToolInvocation
badnessbadness format --no-config --stdin-filepath bench.tex
tex-fmttex-fmt --stdin
latexindentlatexindent -g /dev/null -

The project comparison includes directory traversal and uses check mode:

ToolInvocation
badnessbadness format --no-config --check <dir>
tex-fmttex-fmt --check --recursive <dir>

Each linter reads the document from its path:

ToolInvocation
badnessbadness lint --no-config <file>
chktexchktex -q <file>
lachecklacheck <file>

With hyperfine, each command gets one warmup and at least three measured runs. The script ignores exit codes because lint findings and formatting differences can produce nonzero exits. The commands and timing loop are defined in benches/compare_format.sh.

Language-server sessions

benches/compare_lsp_memory.sh starts three fresh sessions each of badness lsp and texlab run. In each session, the harness initializes the server, waits for background work to settle, opens five documents, and collects diagnostics using the server’s pull or push model. It then requests document symbols and citation or reference hovers and waits for background work to settle again.

The timed symbol and hover requests cover three chapter files. Definition, references, and rename use the Aup91 citation in Chapter1/chapter1.tex, whose entry is in References/references.bib. References include the declaration. Rename computes edits without applying them. Each request target gets two warmup rounds and 20 measured rounds per session. The chart aggregates these samples across all three sessions. The recorded results also include response sizes and counts of symbols, locations, edits, and affected files.

The harness samples the server and all descendant processes through Linux /proc every 150 ms. Background work has settled when CPU use stays below 5% of one core for five seconds. A phase fails if it does not settle within 60 seconds. Workspace and open-file readiness timings end at the start of their respective quiet periods.

Memory is recorded after initialization (Baseline) and after the open-file workload settles (Settled). Peak is the largest sample through the timed requests. The chart shows the median of each measurement across the three sessions, and the JSON file retains the measurements from each session.

Contributing to Badness

Thanks for your interest in Badness, a formatter, linter, and language server for LaTeX. This guide covers everything you need to build the project, run the tests, and get a change merged. Contributions of all sizes are welcome, from typo fixes to new lint rules and parser features.

Getting set up

Badness is a Rust workspace (edition 2024): the root package is the badness CLI/LSP/linter crate, and the publishable badness-parser and badness-formatter library crates live under crates/. The toolchain is pinned by rust-toolchain.toml, so a stable rustup install picks up the right version automatically. The published crates support Rust 1.89 and newer; CI checks that compatibility floor separately from the pinned development toolchain.

git clone https://github.com/jolars/badness
cd badness
cargo build

If you use Nix with devenv, the dev shell provides the full toolchain plus the profiling and benchmarking tools (perf, cargo-flamegraph, hyperfine, cargo-show-asm, cargo-llvm-cov) and the go-task runner. It loads automatically with direnv.

The task runner is go-task; task --list shows every available task. The most common ones are below, but every task maps to a plain cargo invocation if you’d rather not install it.

Building and testing

TaskEquivalentWhat it does
task buildcargo buildDev build.
task testcargo testRun the whole test suite.
task parser-propertiesRun parser losslessness properties at nightly depth.
task fmtcargo fmtFormat the code.
task lintcargo clippy --all-targets --all-features -- -D warningsClippy, warnings as errors.
task checkEverything CI runs: fmt-check, lint, test, wasm.

Run task check before opening a pull request; it mirrors CI exactly.

Badness uses insta for snapshot tests. When a change deliberately alters formatter or parser output, refresh snapshots with task snapshots and review the diff before committing.

The ordinary test suite runs 256 cases for each parser losslessness property. task parser-properties raises that to 4,096 cases, matching the scheduled nightly job. When proptest finds a counterexample, reduce it and preserve it as a readable regression test; the suite deliberately does not persist opaque seed files.

Performance is first-class. Benchmark before optimizing, and never regress losslessness for speed.

Checks that don’t run in CI

Two oracles need more than a Rust toolchain, so run them by hand when your change touches what they cover.

task typeset:check compiles tests/typeset/*.tex before and after formatting and diffs the typeset output. The CST oracles cannot see the one risk the key-value argument flag takes, where a space token is trivia to the CST and content to TeX, so run this when touching keyval signature data or the optional-argument lowering. It needs a TeX install.

task parse-compat runs texlab’s parser as a differential oracle over a corpus, skeletonizing both trees and comparing. It is a reference we measure against, not one we match, so a divergence is something to explain rather than automatically fix.

Project layout

Badness parses LaTeX into a lossless concrete syntax tree (CST) and builds three tools on top of it: a formatter (badness format), a linter (badness lint), and a language server (badness lsp). The architecture follows rust-analyzer: a hand-written, error-tolerant lexer and parser turn LaTeX into a flat token stream, then an event stream that a tree builder feeds into rowan; a semantic layer assigns meaning on top of the generic tree; and incremental recomputation is salsa-first.

The Architecture page in the book is the full tour, and it is worth reading before a non-trivial change.

Where things live:

  • crates/badness-parser — syntax layer, parser, semantic layer, the BibTeX pipeline, and the data/ signature artifacts.
  • crates/badness-formatter — the layout engine and the .bib formatter.
  • crates/badness-wasm — the wasm shim powering the docs playground.
  • src/ — the CLI, LSP, linter, and project layers, plus shim modules re-exporting the member crates at their old paths.

Both library crates must keep building for wasm32-unknown-unknown, so nothing in them may touch the filesystem, threads, or processes. A CI job guards this. Anything that needs the outside world belongs in the root crate.

Invariants

These properties are held by construction and enforced as test oracles. A change that breaks one is a bug, not a trade-off.

  • Losslessness: reconstruct(text) == text, byte for byte.
  • Idempotence: format(format(x)) == format(x).
  • The formatter is whitespace-only: it changes trivia (whitespace, newlines, comments, .dtx margins and guards) and nothing else. Content rewrites such as x^{2} → x^2 are linter autofixes, not layout.
  • Protected regions: verbatim-like content (verbatim, lstlisting, \verb, comments) is never altered by the formatter, apart from a document-wide line-terminator normalization.

A couple of ground rules keep the design coherent:

  • Semantic facts reach the parser only through a narrow, curated admission test; when in doubt, a fact belongs in the semantic layer. Parsing is the parser’s job; layout is the formatter’s job. Never paper over a parser mistake in the formatter.
  • New parser features need corpus and snapshot tests and a losslessness assertion.

Making a change

  • Prefer trunk-based development and atomic commits. Branch first for substantial changes; small fixes can go straight to main.
  • Follow Conventional Commits, for example feat(linter): add missing-required-argument rule or fix(parser): recover at unbalanced brace. The CHANGELOG.md is generated from the commit history by versionary, so a clear, well-scoped commit message is what shows up in the release notes. Don’t hand-edit CHANGELOG.md.
  • Keep commit subjects short (imperative mood, ideally under 60 characters) and use the body for rationale. Close issues with Fixes #123 in the body.
  • A rustfmt git hook rewrites unformatted files and aborts the commit, so run cargo fmt first. Clippy warnings are treated as errors.

Each workspace crate is its own versionary package with its own changelog and version. The root CLI tags bare v*; the members tag badness-parser-v* and badness-formatter-v*. Only the bare v* stream carries release assets.

Adding a lint rule

The add-lint-rule workflow automates this, but the shape is fixed:

  1. Implement Rule in a new src/linter/rules/<name>.rs, choosing node-shape, whole-file, or streaming dispatch, with an id, a default_severity, a description, and at least one triggering example. Emit a losslessness-safe fix where one is warranted, and set emits_fix accordingly. Rules are enabled by default; override default_enabled only for useful checks whose unavoidable false positives make them better suited to explicit selection.
  2. Register it in the three lockstep lists in src/linter/rules.rs: the module declaration, the re-export, and the entry in all_rules().
  3. Ship unit tests next to the rule and an integration test, plus a losslessness assertion on any fixture.
  4. Regenerate the rules reference with task docs:rules. Do not edit the rendered page by hand.

Generated data files

Several files in crates/badness-parser/data/ are generated from pinned upstream sources by scripts/gen_*.py and guarded by paired task …:check and :sync targets: cwl_signatures.json, the package and class name lists with package_metadata.json, and bib_fields.json. Re-sync them through their task rather than hand-editing the mechanical facts. signatures.json, colors.json, and tikz_libraries.json are curated by hand and may be edited directly.

Command entries in signatures.json may set completionKind to symbol for argument-free math and text symbols or logos, or keyword for argument-free control, spacing, and declaration commands. This metadata affects completion only. Omit it when the classification is unknown; an empty CWL signature does not establish zero arguments. Document definitions and project declarations override these classifications.

The published badness.toml schema is generated from the Rust configuration types. Regenerate badness.schema.json with UPDATE_EXPECTED=1 cargo test --test config_schema, and review the diff rather than editing the JSON by hand.

Windows CI bites twice

Line endings: the formatter emits LF and tests compare bytes against checked-in fixtures. When you add a fixture in a new extension under crates/*/tests/fixtures/** or crates/badness-parser/tests/corpus/**, add a matching … eol=lf line to .gitattributes. Never normalize line endings in code to pass a test; fix the attribute instead.

URIs: decode LSP URIs to filesystem paths only through uri_to_fs_path and path_to_uri in lsp.rs. Tests and snapshots must not assume / versus \.

Documentation

User-facing docs are an mdBook under docs/. Preview them locally with task docs:serve (live reload) or build them with task docs. The linter-rules reference and the benchmark page are generated; regenerate them with task docs:rules and task bench respectively rather than editing the rendered pages by hand.

Use task docs before auditing the production output. Its postbuild steps add canonical and social metadata, publish branding/og.png, adapt the navigation for keyboard access, and generate the sitemap. A plain mdbook build or the live preview does not run these steps. The standalone playground receives the same metadata as the book chapters. Redirect and helper pages stay out of the sitemap.

Test the documentation tools with cargo test --manifest-path docs/doc-utils/Cargo.toml. The benchmark renderer loads the vendored scripts in docs/src/vendor only when a page contains charts.

A note on AGENTS.md

The repo’s AGENTS.md is the operational contract for AI coding agents. It includes both repository-wide and subsystem-specific directives and is kept under 32 KiB so agents can load it as a single checklist of things not to break. For architectural rationale and tradeoffs, read the book’s Architecture page.

License

By contributing, you agree that your contributions are licensed under the project’s MIT License.

Architecture

Badness parses LaTeX into a lossless concrete syntax tree (CST). Its formatter, linter, and language server work from that tree, while a separate semantic layer interprets commands and environments. This separation lets the parser preserve source it cannot fully understand—a necessity for a language whose syntax can change as a document runs.

The design follows rust-analyzer, with a hand-written, error-tolerant parser, a rowan tree, and salsa queries for incremental recomputation. arity, a similar tool for R, was another influence. Build and test instructions live in Contributing.

From source to editor features

The parser reads a flat token stream and emits events for a separate tree builder:

text → lexer → token stream → parser → event stream → tree_builder → GreenNode

Start, Tok(idx), and Finish events describe the tree without constructing it during parsing. Diagnostics travel separately, keyed by byte range. The tree builder feeds the events into rowan and attaches trivia, retaining every byte of the source. A specialized SubTok event lets math parsing split a lexer token when TeX binds a script to a single character.

Each LaTeX node opening returns an event-layer Marker, which closing the node consumes. A DropBomb catches forgotten completions in debug builds without panicking again during unwinding. Marker::precede gives the same obligation to wrappers opened retroactively, such as paragraphs and scripted atoms. Comment binding instead moves a completed construct’s start over its documentation comment with extend_back, keeping its existing finish. The parser also checks that the final event stream balances before the tree builder receives it.

The LaTeX grammar keeps math bodies, script attachment, and paired \left/\right delimiters in grammar/math.rs. Named math environments use that same module’s body parser, so delimiter and script recovery stays shared. grammar/gates.rs owns the shape-gate policies, shared forward scan, and verdict caches. Its tests count token visits to check that repeated and nested queries stay linear; grammar callers only ask whether a delimiter can pair.

The formatter lowers the tree to a Doc intermediate representation and prints it. The linter collects diagnostics through a shared traversal. The language server uses salsa queries to combine syntax and semantic information across files.

A tree depends only on source text and explicit project declarations. Package scope, the runtime signature database, and the filesystem cannot change its shape. Parsing can therefore be repeated or cached without depending on the machine that happens to run it. Badness does not execute macros or typeset documents.

The crates

The Cargo workspace contains four crates. The root crate, badness, owns the CLI, linter, language server, configuration, and project discovery. It handles filesystem access and passes resolved inputs to the libraries.

badness-parser contains the LaTeX and BibTeX syntax trees, parsers, typed AST wrappers, and semantic models. Its data/ directory holds signature data, which build.rs turns into PHF tables. badness-formatter depends on that crate and provides the layout engine and both formatters. badness-wasm is an unpublished wasm-bindgen wrapper that powers the playground.

Both library crates build for wasm32-unknown-unknown, so their runtime code cannot use the filesystem, threads, or child processes. The formatter also runs in the dprint plugin. The plugin uses an empty runtime signature database, whereas the CLI can supply signatures scanned from neighboring .sty and .cls files. This accounts for an intentional difference between the two formatting entry points.

The root crate re-exports the libraries through modules such as src/parser.rs to preserve existing import paths. Its formatter and semantic modules also provide the disk-backed entry points that cannot live in the libraries.

The BibTeX side

BibTeX has a separate pipeline in bib/, with its own grammar, lexer, syntax kinds, typed AST, and semantic model. It follows the same lossless event-stream architecture as LaTeX and supports formatting, linting, completion, and document outlines.

% comments in .bib

Badness follows biber’s interpretation of % inside an entry: it begins a line comment between fields or value components, but remains ordinary text inside a braced or quoted value. Thus title = {50% off} retains its percent sign. Classic BibTeX rejects comments inside entries, and texlab does not model them; the latter difference is recorded in bib_parse_compat_allowlist.toml.

The lexer always emits a bare PERCENT token. The grammar knows whether it is reading a value or skipping trivia, so it decides whether to wrap the rest of the line in a COMMENT node. Braced groups, quoted strings, @comment bodies, and top-level junk retain % as ordinary text.

A percent sign inside a value also matters to formatting. BibTeX passes it to LaTeX, where it becomes a comment and makes the following line break significant. The formatter therefore preserves values containing an unescaped % byte for byte. Comparing syntax trees would not catch an unsafe line join here: the bibliography would still parse, but typeset differently.

Comments between fields stay attached to their surroundings when fields are sorted. A trailing comment remains on its field’s line, while an own-line comment moves with the following field. Comments after the final field appear above the closing delimiter. If an entry has nowhere suitable to place a comment, as can happen with @string, @preamble, or an entry without fields, the formatter preserves the entire block.

Inputs and configuration

The CLI discovers .tex, .sty, .cls, .dtx, .ins, and .bib files through ignore, respecting .gitignore and project excludes. It finds badness.toml by walking each input’s ancestors. Filesystem discovery and configuration resolution stay in the root crate; the libraries receive values such as FormatStyle and resolved declarations.

The file kind determines the lexer’s starting mode. Package sources (.sty, .cls, and .dtx) begin with @ treated as a letter, as though \makeatletter had already run. Document sources do not. Every file kind uses WrapMode::Reflow by default because reflow safety depends on the source’s structure, not its extension.

Project configuration controls formatting, lint selection, build-artifact locations, declarations, and excludes. Machine settings, such as the TeX installation and PDF viewer, belong in editor configuration. The language server merges editor settings with project settings and distinguishes an unset option from an explicit value.

The server caches resolved configuration by document directory. Before reusing an entry, it checks the existence, modification time, and length of the config files examined by the ancestor walk, including any fallback. This detects a changed or deleted config and a newly created nearer one even when the editor cannot register file watchers. Watcher notifications provide earlier invalidation when available.

The cache retains the candidate paths and Git boundary predicates, but validates them against the filesystem on each request. It checks a discovered source only once and checks an environment or global source separately.

Anchor validation must still detect symlinks that redirect discovery to another ancestor chain, even when both config files have identical timestamps and lengths. On Linux, an anchor whose spelling already matches its cached canonical path can use openat2 with O_PATH | O_DIRECTORY | O_CLOEXEC and RESOLVE_NO_SYMLINKS. Success proves that the path still resolves without following a symlink, replacing a readlink call per component with an open and close. The descriptor is closed immediately. Symlinks, noncanonical spellings, failed probes, and platforms without this operation use ordinary canonicalization. No client watcher capability or time-based validity window substitutes for these checks. See the Linux openat2 contract.

Declarations

Declarations let a project describe constructs whose meaning cannot be inferred reliably from source. Command declarations assign reference or citation semantics:

[commands.eqrefs]
like = "cref"

[commands.mycite]
like = "parencite"

Here, eqrefs takes cref’s comma-separated reference keys, and mycite takes parencite’s citation behavior. These declarations affect analysis and completion. They do not declare arity, change argument attachment, or select formatter layouts.

Environment declarations can assign built-in behavior or introduce delimiter spellings:

[environments.myenv]
like = "align"

[environments.eqnarray]
begin = ['\bea']
end = ['\eea']

[environments.mytheorem]
like = "theorem"
begin = ['\startmyenv']
end = ['\endmyenv']

An alias may supply just one delimiter because a literal \begin{X} or \end{X} can supply the other. TOML literal strings avoid escaping the backslash, which is optional in these spellings.

Environment declarations enter the parser through ParseCtx. The parser sees only the resolved declarations, never a full SignatureDb, and still checks whether delimiters can pair structurally. A declaration supplies a spelling; it cannot force an impossible pair into the tree.

The like targets come from closed, curated tables. Environment declarations copy a built-in environment signature; command declarations copy a reference or citation family. Neither resolves against CWL data or scanned definitions, and neither exposes an argument specification. Config loading rejects unknown targets, empty entries, invalid spellings, and collisions. Tables are keyed by category and name so layered configuration can merge individual declarations. Once validated, declared entries take precedence over scanned and built-in entries.

Both declaration categories share a high-durability salsa input, but reach their readers through separate queries: parse_declarations and semantic_declarations. Each query retains its previous result when its subset is unchanged. Renaming a citation alias therefore updates semantics without invalidating parses or their reparse caches. The LSP publishes declarations in the request dispatcher so switching between workspace roots cannot leave a handler using another project’s configuration.

Syntax and semantics

The syntax tree records what the source contains. The semantic layer interprets it using curated built-ins, CWL-derived signatures, and definitions scanned from source. Keeping those responsibilities separate is especially useful for generic LaTeX arguments: \foo{a}{b} might be a two-argument call or a command followed by two independent groups. The parser preserves the groups; semantics can interpret them when a signature is available.

Some curated or declared facts do influence parsing, but only when the source can disprove them. A delimiter alias, for example, must pass a shape gate before it can open an environment. If no valid closer is reachable, it remains an ordinary command. Generic arity cannot be checked this way: an incorrect arity can produce a lossless tree with the wrong attachment. It therefore stays out of the parser.

Semantic lookup can depend on the enclosing environment. SignatureDb::command_at searches the nearest environment with a local entry in environmentCommands, then falls back to the global signature. In exam’s parts environment, for example, \part accepts an optional points argument and has no sectioning role. The linter, outline, and label context share this interpretation, including in question files without a class declaration. Environment headers and closers retain their surrounding meaning. These local entries do not change parser grouping. The formatter resolves the same local signatures through Signatures::command_at, with scanned definitions taking precedence over the curated database.

The formatter also selects curated classCommands from one literal top-level \documentclass declaration. These signatures sit below loaded local package definitions and document definitions. For cas-sc and cas-dc, this identifies the key-value arguments of \affiliation without assigning that meaning to the same command in other classes. Selection reads only the CST; it neither changes parser grouping nor consults the TeX installation. Computed, nested, and multiple class declarations do not select a class signature.

Facts that authorize a rewrite need stronger evidence than facts used for completion. ContentKind::Keyval permits breaks after commas that had no following whitespace, so a wrong classification can change typeset output. Likewise, the curated labelKey flag says that an environment’s first optional argument can define a label. This cannot be inferred merely from key-value syntax: different processors can give a key named label different meanings. The semantic model accepts flat literal values, processes repeated entries in order, and treats a later dynamic value as unknown. A separate captionContainer flag identifies non-float environments in which \captionof can own a preceding label. Keeping these claims separate prevents one layout or lint classification from silently granting another.

The parser

The parser uses recursive descent over a flat token stream. It always produces a lossless tree, including for malformed source. When it cannot establish a construct’s structure statically, it leaves that construct generic rather than trying to execute TeX.

Sanctioned lexer modes

TeX lets source change how later characters are read. Badness recognizes a bounded set of common patterns where the source gives enough evidence to select a lexer mode. These modes cover package code, expl3, verbatim material, and literate .dtx sources without attempting general \catcode evaluation.

\makeatletter makes @ a letter. \ExplSyntaxOn and \ProvidesExpl* declarations enable expl3, where _ and : are letters. The flags are independent and can be active together. In .dtx files, a %<@@=…> guard or \ProvidesExpl* declaration anywhere in the file enables expl3 catcodes in every macrocode body.

Verbatim commands and environments capture their bodies as single tokens. Curated signatures describe built-ins, while a bounded two-pass definition scan recognizes user definitions from catcode-changing patterns and known definers such as \lstnewenvironment. The \NewDocumentEnvironment family also declares a verbatim body with a final c argument type. The scan recognizes that type only when every preceding argument has a brace or bracket shape the verbatim lexer can consume. Stars, token tests, and other unsupported header shapes withhold capture so they cannot shift its starting point. A c inside a default or delimiter is ordinary specification data. This keeps incomplete code examples opaque without interpreting the environment’s replacement text.

A command can also have one positional verbatim argument: \href captures its URL but leaves its visible text parsed. Such a capture requires the expected balanced group, and a local redefinition suppresses a colliding built-in mode. Short-verb declarations such as \MakeShortVerb{\|} allow |…| on one line to form an opaque token; .dtx mode and curated documentation classes enable | initially.

Definition bodies need different treatment from running document text. Inside curated command and environment definers, \begin and \end remain ordinary commands because a replacement body need not balance them. A control-symbol name in a definition, such as \DeclareRobustCommand\[, is definition data and cannot open display math. Expl3 regions likewise pass environment delimiters around as data, so the parser leaves them as commands and accepts an orphan \] without a diagnostic.

A .dtx macrocode body ends only at its literal frame line. Unmatched braces inside a chunk remain plain tokens because a definition can open a brace in one chunk and close it in another. The lexer also recognizes line-leading %<…> guards and ^^A comments on documentation-margin lines. These distinctions matter later when the formatter reconstructs the documentation’s % margins.

Other bounded rules prevent ordinary TeX data from becoming structure. The lexer isolates the delimiter after \left or \right. In numeric contexts, it recognizes backtick character constants, so \char`$ cannot open math. It also recognizes escaped character constants that occupy a whole alignment cell.

Shape gates

Before opening some constructs, the parser scans ahead to check that its normal walk can consume them. These checks are called shape gates. They keep delimiters used as macro data from swallowing unrelated source. A failed gate usually leaves ordinary syntax without a diagnostic: parser diagnostics can prevent formatting, so routine macro patterns should not trigger them.

For $, \[, and \(, a gate requires a reachable closer before an unbalanced brace, paragraph break, or end of file. Environment pairing asks a different question: would the environment escape the brace group in which it began? An escaping } demotes the opener, but reaching EOF does not. This preserves the useful diagnostic for a document with a missing \end.

Gates share Parser::gate_batch, which can settle several openers during one scan instead of repeatedly scanning nested source. Each policy must match the walk it guards, including recovery anchors, brace handling, and .dtx frames. State that can change an answer, such as the enclosing math flavor, belongs in the memoization key. DollarGate is not memoized because demoting $$ resumes parsing at its second dollar sign, where the same token position can pose a different question.

Most pairing gates count nested openers and environments independently. LeftRightGate instead uses one stack of brace, environment, and \left frames, because their order determines whether a \right can match. A frame mismatch invalidates outer pairs as well. Its anchors are the delimiters that end the surrounding math body; \left and \right themselves remain recognizable in definition and macrocode bodies, where package math commonly uses them.

Bracket gates reflect the fact that [ and ] are ordinary TeX characters. They attach an optional argument only when its closer is reachable along the optional-argument parse path. A command-adjacent nested [ can claim a closer, which then cannot close the outer argument. Environment delimiters and paragraph breaks normally stop the scan at any brace depth. The narrow long-text exception allows a paragraph break when both ends have the tight shape \cmd[…]{…}: the bracket directly follows the command, and its closer directly precedes a mandatory group. Signature arity does not supply this evidence.

In math, bracket gates also account for the enclosing delimiter. Within display math, a balanced $…$ can be nested inside an optional argument. Within inline $…$, a dollar sign at the bracket’s own level ends the surrounding math and prevents attachment. Macrocode brackets ignore braces already classified as unmatched chunk data, but stop at structural closing braces and respect the chunk’s frame.

All gates use the same interpretation of docstrip trivia. A guard-only line interrupts a run of newlines because docstrip removes that line entirely. A documentation margin, by contrast, leaves the line in the documentation stream, so a margin-only line can still form a paragraph break.

Environment aliases

A definition whose entire replacement body is \begin{X} or \end{X} can supply an environment alias. This lets \bea … \eea form the same environment shape as \begin{eqnarray} … \end{eqnarray}. An alias for one delimiter may pair with a literal spelling of the other. Projects can also provide aliases through declarations.

Inference is limited to the current file and curated target environments. Definitions themselves cannot open alias environments, and a positive gate must find a reachable closer before an alias can pair. More complicated definitions, such as argument-taking aliases or unresolved \let chains, remain generic.

Consumers resolve an alias from the parsed node. Looking only at raw spelling would confuse a command alias \bea with the unrelated environment name in \begin{bea}. Literal and alias delimiters share a target lookup, including in math parsing, while retaining separate closer indexes.

The conditional gate

A complete \if … \else/\or … \fi becomes a CONDITIONAL node with positional branches. Recognition uses a curated opener model shared with the linter, excluding if* macro families and definition operands that are not live control flow. The gate requires a reachable \fi at the appropriate brace, environment, and math nesting levels. It respects macrocode frames, stops at paragraph breaks at the conditional’s own level, and does not recognize conditionals in expl3 regions.

The node establishes the conditional’s extent, but does not claim to separate its test from its body. TeX’s scanners determine that boundary, and a static parser cannot identify it reliably. Even ast::Conditional::closer is fallible: a nested opener can be demoted when the walk applies its gate, allowing the walk to finish before the outer scan’s predicted closer.

Recursive descent, with Pratt local to math

Text parsing has no precedence rules. Math uses local precedence handling for script binding and \left…\right structure, but does not build an arithmetic expression tree. Ordinary characters, including arithmetic operators, stay in coalesced WORD tokens until a script requires a finer boundary.

For example, a,b^2 attaches the superscript to b. An unbraced script argument also consumes just one input character, so x^23_i becomes x^2 followed by 3_i. The parser emits byte-range sub-tokens for these boundaries. Elsewhere, a semantic view presents one virtual math atom per Unicode scalar without changing the CST.

The atom classifier combines generated unicode-math data with curated overrides. It supplies TeX spacing classes and a separate delimiter role. That separation matters because an atom’s spacing class does not prove it can pair: \sqrt’s Open class, for example, must not enter bracket accounting. Commands and characters share the lookup, and unknown commands default conservatively to Ord.

Curated environment signatures can select math parsing for an entire body. Math parsing and alignment layout remain separate claims: gathered is math, while aligned is both math and an alignment. A wrapper such as empheq can therefore parse as math while the formatter derives its grid layout from the body’s & and \\ structure.

Argument grouping and bracket policy

The parser greedily attaches trailing brace and bracket groups, subject to the bracket gates. The semantic layer later interprets that attachment using the available arity. A lone * directly after a command and before an argument can join the command as a starred-variant marker.

Curated built-in slots may refine how a group’s contents are parsed through ArgumentDomain::Math or Text. A positional matcher skips omitted optionals and aligns the attached groups with those slots. A matched math group uses the math parser; unknown, unmatched, and excess groups use generic parsing. Attachment remains greedy, and CWL data, scanned definitions, and project declarations cannot supply these domains.

Verbatim slots require a related distinction. A verbatim argument is captured by the lexer, preventing characters such as % from becoming syntax. ContentKind::Opaque, by contrast, tells the formatter to preserve whitespace in an already parsed group. A companion slot matcher accounts for captured arguments so later groups retain their proper positions.

Expl3 provides an argument specification in the command spelling itself. The parser can therefore attach \tl_set:Nn \l_a {x} according to Nn, keeping \l_a and {x} as arguments of the same call. Greedy attachment would instead attach {x} to \l_a. A token scan in grammar/expl3.rs plans the attachment, and the walk replays that plan. Control-sequence arguments keep their own bare COMMAND nodes; groups remain ordinary GROUP nodes.

Underivable heads, including w and D forms, colonless names, and \::n expansion drivers, retain greedy grouping. The scan also declines when math, lexer-mode changes, docstrip boundaries, or unreachable closers prevent it from matching the walk. A blank line inside a brace group allows the consumed prefix to be committed. Matching braces are indexed once per frame so nested calls do not repeatedly scan the same groups.

Attachment mistakes can survive both losslessness and formatter idempotence checks. An independent oracle therefore compares grammar attachment with semantic::expl3 argument consumption. The semantic model also handles calls that fall back to greedy parsing. Texlab has no argspec model, so expl3 regions are allowlisted in the differential gauge.

Trivia attachment

Trivia normally belongs to the nearest enclosing node. A contiguous run of own-line % comments immediately before a command or environment instead binds to it as a DOC_COMMENT. A trailing comment stays where it is, and a blank line breaks the forward attachment.

LaTeX has no separate documentation-comment marker like Rust’s ///. Stopping at blank lines keeps a license header from becoming documentation for the first command. Node kind alone determines which constructs can receive comments; attachment does not consult signatures.

Error recovery

An error produces a diagnostic alongside the tree. The parser recovers at structural anchors such as \begin, \end{…}, a blank line, }, $, &, and \\, always consuming input rather than looping on an unexpected token.

Losslessness tests reconstruct arbitrary UTF-8 and generated malformed syntax through both parsers. LaTeX cases cover documents, packages, .dtx sources, and explicit declarations. These tests require byte-for-byte reconstruction, not an absence of diagnostics. Curated snapshots and differential comparisons with texlab check the structure that reconstruction alone cannot establish.

Incrementality

Salsa tracks dependencies across files and queries. Its database stores rowan GreenNodes and creates red cursors on demand, because red trees cannot be shared across threads. parsed_document stores green trees under no_eq, unsafe(non_salsa_values); parsing remains a pure function of source and declarations.

Source text is an Arc<str>, shared by the live buffer, worker jobs, salsa, and read snapshots. Passing a document between them increments a reference count. upsert_file avoids a salsa write when text is unchanged, and text_is_current checks whether a read job still describes the current buffer. Both compare pointers first, then contents: a disk reread may have equal text in a new allocation.

Inputs have durability levels that reflect how often they change. Text has LOW durability, paths and declarations have HIGH, and the ProjectFiles membership input has MEDIUM. Typing can then invalidate text-dependent queries without making stable project data appear changed.

workspace_project derives an equality-comparable project value from ProjectFiles. Keyless queries use it for include and package graphs, label and citation resolution, package options, and signature scopes. Membership belongs in an input so each database snapshot contains the right file set. Keeping one current project value also avoids retaining interned copies of every historical membership and their dependent query results.

Intra-file reparse

Within a changed file, parser::reparse can reuse most of the previous green tree. It tries progressively broader edits: replacing one ordinary token, replacing a protected body token, reparsing a math node, or reparsing a limited prose region. Each tier uses the ordinary lexer or parser and must establish that the edit cannot change structure outside its replacement. If none can do so, the caller parses the whole file.

A successful reparse must produce exactly the same green tree and SyntaxError vector as a full parse of the edited text. Every failed proof returns None. This lets guards remain conservative: an unfamiliar construct costs a full parse without weakening correctness. The implementation does not resume from lexer checkpoints or preserve grammar restart state.

Token and protected-body splices

The token tier relexes a leaf and checks that it remains one token of the same kind. Probes on either side check that it still separates from its neighbors. In \foo1ab, changing 1ab to aab would merge the two tokens into \fooaab, so that edit cannot use a leaf splice. The tier also rules out changes to the definition scan and to any grammar decision that reads the token’s text.

A source-scanning test tracks those text reads in the grammar and lexer predicates. Each read must be unreachable from a spliced leaf or protected by a guard. This covers facts such as a statement’s terminating semicolon, a starred command marker, and an environment name. Newline edits and definition-sensitive positions decline, as do oversized boundary probes.

Math leaves need an additional check because script parsing can split one lexer WORD into several CST leaves. The tier reconstructs the coalesced word and checks its outer boundaries. A leaf that supplies a script or its base must remain one Unicode scalar; an unscripted prefix or remainder may change length. An edit that moves this partition goes to the math tier.

Protected bodies cannot be relexed alone: the lexer needs the opener to enter verbatim mode. That tier therefore relexes the whole enclosing node, including its delimiters. The unedited fragment must reproduce its original tokens, and the edited fragment must reproduce the same sequence with only the body token changed. The closer must be present inside the fragment, ruling out an unterminated capture. This catches an inserted \end{verbatim} or an unbalanced brace in a captured URL without duplicating the lexer’s capture rules.

The locality argument also depends on raw body bytes bypassing lexer state updates. Lexer tests check that property and the counterexample where breaking a capture changes later lexing. Newlines are safe within an intact capture, so pressing Enter in a listing can still use a leaf splice. The replacement shares every green node outside the path from the leaf to the root.

Fragment lexing uses the base parse’s ParseCtx and full-file .dtx facts, including implicit expl3 mode. An edit that changes a full-file signal declines. Reproducing the old fragment does not authorize inventing missing entry state.

Math and prose regions

A math edit that changes shape reparses the outermost enclosing INLINE_MATH, DISPLAY_MATH, or math environment. Its delimiters establish math mode and include the enclosing gates that the edit could invalidate. An isolated parse must first reproduce the old node. The edited parse must then yield one node of the same kind spanning the whole fragment, and a right-boundary probe checks that recovery cannot consume the unchanged suffix.

This tier admits state-neutral math syntax. Control sequences, comments, environment names, definition-sensitive positions, lexer-mode ambiguity, and .dtx source decline. Fragment diagnostics are replaced, prefix diagnostics are retained, and suffix diagnostics shift with the edit. A base whose diagnostics are out of source order declines because a local splice cannot reproduce a global recovery-stack reorder.

The region tier covers two prose shapes: an edit across several direct leaves of one top-level paragraph, and an edit to the blank-line seam between two paragraphs. It first reparses the old fragment under the base’s exact context. Edits may touch only direct prose or trivia and insert no structural or catcode-sensitive spelling. Unchanged commands can remain within the fragment because their state transitions are unchanged.

A single paragraph uses a node splice. Removing a seam may merge two paragraphs, so that case rebuilds ROOT from shared green children. Both paths replace fragment diagnostics and shift those after it. Seam splicing declines .dtx until its column-sensitive documentation layer can be accounted for.

Blank lines alone do not prove that arbitrary regions are independent. A wider region tier would also need to account for forward-looking gates outside the fragment, verify boundaries against unchanged neighbors, and show that the replacement’s tokens and diagnostics compose with untouched siblings. The current restrictions provide that locality without additional parser state.

The reparse cache and edit chain

The previous text, tree, diagnostics, and pending edits live beside salsa, accessed through the IncrementalDb reparse methods. They cannot be ordinary inputs: invalidating the previous parse on every write would discard the value needed for reuse. Reading this cache inside parsed_document is sound because it cannot change the query’s result. A missing, stale, or evicted entry merely causes a full parse.

The cache stores a new base only after every fallible step has completed, so a panic or cancellation cannot leave mismatched text and tree. It drains the consumed prefix of the edit chain rather than clearing the whole chain, because another edit may arrive during parsing. That prefix is drained even if reuse fails: once the base changes, those edits no longer describe a transformation from it. Eviction favors entries that have already benefited from reparsing, preventing a project-wide query from displacing every open buffer with files parsed only once.

The language server supplies edits directly. apply_content_changes returns the clamped byte offsets it used, with each edit relative to the text produced by its predecessors. A whole-buffer replacement returns None. Disk reloads and other writes without an edit chain also force a full parse; the query does not spend time rediscovering edits by diffing whole texts.

Every upsert_file is followed by reparse_stage_edits, passing the available chain or None. Staging must follow the write so an in-flight query cannot see edits for text that has not arrived. It must also happen when the write is skipped: a buffer can return to the text salsa already holds while still having an edit chain relative to the older reparse base.

Reparse validation

Every tier returns through a shared finish function. In debug builds, it compares the result with a full parse. In every build, it checks that the tree spans exactly the edited text and falls back if the length differs. Seeded hazard cases and corpus edits check both the tree and diagnostics; deliberate failures test that the oracles detect divergence.

task reparse-corpora:check runs the shared harness in tests/support/reparse_harness.rs across the pinned corpora, using each file’s normal lexer configuration. It checks a minimum splice rate and compares exact per-tier tallies against tests/reparse_baselines/. Both are needed: a tier that refuses every edit can pass every equivalence assertion, and a move between tiers can leave the overall splice rate unchanged.

The direct benchmark in benches/reparse.rs declares the required tier and a speedup floor for each case, comparing it with a full parse in a release build. task bench:gate runs these checks alongside the end-to-end keystroke benchmark, which also measures buffer updates. Together they test correctness, useful coverage, and whether reuse actually saves work.

Typed AST wrappers

Thin AstNode and AstToken wrappers provide typed, read-only access to the rowan tree. Their accessors describe positions and tolerate greedy attachment. They do not consult signatures: a Command::title() accessor would be misleading because \section and \newcommand share the same syntax kind.

The formatter uses wrappers for field access and raw nodes for structural dispatch. Meaning remains in the semantic layer, where callers have the context to interpret an attached group.

The formatter

The formatter lowers the CST into a Wadler/Prettier-style Doc intermediate representation. A separate printer chooses flat or broken forms according to the available width. Lowering can thus describe the possible layouts without committing to a line break before the printer knows the current column.

Whitespace and content

The formatter changes trivia while preserving non-trivia content. It normally replaces whitespace runs with break primitives, then lets the printer choose spaces, line breaks, and indentation. Comments and protected regions retain their contents, apart from configured line-ending normalization.

Content rewrites belong to linter fixes. Removing script braces from x^{2} or replacing $$…$$ with \[…\] requires a separate meaning-preservation argument. Keeping such edits out of formatting makes its content guarantee checkable independently of individual layout rules. Fix application likewise does not invoke the formatter.

Formatting can change token boundaries and CST shape. Inserting insignificant math whitespace, for example, splits a coalesced WORD into several tokens. The content oracle therefore compares concatenated non-trivia text rather than token boundaries. Parse stability is not a formatter invariant.

Trivia-invariant layout

Idempotence requires the second formatting pass to make the same layout choices as the first. That becomes difficult if a rule reads a distinction the formatter can erase or create. A lone newline is the usual example: reflow can turn alpha\nbeta into alpha beta, or insert a newline when a line exceeds the width. Treating that newline as evidence of a structural boundary lets the first pass change the second pass’s decision.

Ordinary layout rules therefore read only trivia properties their output preserves: blank-line presence, comment presence and own-line status, and .dtx margins or guards at column zero. If a predicate satisfies P(fmt(x)) == P(x), using it cannot by itself make the next pass choose a different layout.

The Gap type enforces this distinction at the lowering boundary. Its variants are Glued, Space { flat }, Blank, and Comment; there is no Newline variant. A lone newline and a single space have the same flat rendering. Wider authored whitespace can remain in flat only because its readers reproduce it unchanged. Width calculations and break selection use this normalized view.

Some policies intentionally preserve authored lines. They use WideGap, which also exposes newline counts, and need a fixed-point argument for every layout they can emit. These Tier 2 cases include preservation modes, unresolved statement layouts, and a few narrow rules within reflow. A preservation rule has a straightforward argument when it re-emits a newline in the same place: the next pass sees the same break and preserves it again.

The command-only-line rule is one such case. Curated block commands have a positive signature property, but an unknown \mymacro may still occupy its own authored line. The rule preserves breaks around that line without moving them. If width wrapping strands a command on its own line, preserving that break on the next pass agrees with the original fill. The rule is excluded from signature-proven prose arguments, where a newly forced child break could instead change the enclosing group’s layout.

Opaque groups under Reflow are otherwise width-driven: they stay flat when they fit and wrap at existing gaps when they do not. A glued junction cannot acquire a break. Delimiter padding can become a newline only when its flat spelling is the single space that newline reproduces. Interior blank lines, comments, embedded newlines, and forced child breaks select a block form; edge blank lines cannot select it because that form trims them away.

Recognized Beamer overlay bodies are a Tier 2 exception. The formatter matches the complete body slots of \only, \uncover, \visible, \invisible, \onslide, \action, \alt, and \temporal, including bounded angle specifications left generic by the parser. These bodies keep authored line breaks and inline gaps while their indentation normalizes. Width does not reflow them. Each emitted gap retains its line structure on the next pass, including at the braces, so both inline and multiline bodies are fixed points. Unmatched syntax and groups beyond the wrapper’s body slots keep ordinary group layout.

Signature-matched opaque braced environment arguments keep their top-level words together. A value such as {section in head/foot} should not split at spaces to keep an earlier option list flat. Lone source newlines normalize to spaces; comments, paragraph breaks, and nested blocks retain their structural layout.

A narrow exception preserves a newline after \\ in a structurally plain, command-only text group. Its block framing and row breaks recur on the next pass, giving the rule a fixed point. Macro-like groups and virtual .dtx documentation streams do not use it. Outside ordinary reflow, delimited-group rules may also retain the single-line versus multiline distinction when each emitted form preserves that distinction.

formatter::perturb tests these properties by varying trivia without changing TeX content. check_trivia_convergence requires every variant to reach a fixed point, parse cleanly, reconstruct losslessly, and retain its non-trivia content. check_trivia_invariance asks the stronger question of whether every variant formats identically. The latter also reports intentional line-preserving cases, so badness debug format --checks trivia-strict serves as a survey rather than a universal pass condition.

Paragraph line breaks

WrapMode determines how the printer handles paragraph breaks. The default, Reflow, fills lines to the configured width. Stable keeps acceptable authored breaks while balancing overflow, changes, displacement, and raggedness. Preserve keeps authored breaks. Sentence places one sentence on each line regardless of width. Semantic keeps authored breaks and sentence boundaries, then fills each remaining run to the configured width. The two modes share sentence detection and citation handling, using language-specific abbreviation profiles resolved from the format configuration. Width decisions stay in the printer so indentation and documentation margins count toward the limit.

Semantic wrapping preserves the breaks it inserts on subsequent passes. Each width-broken run already fits at the same indentation, and sentence breaks and authored breaks remain hard boundaries. These boundaries use Ir::PreservedLine: the printer rejects a flat layout containing one, but lowering does not treat it as a structural block break. Otherwise, reparsing a width break inside a nested prose argument could select a different enclosing optional or group layout and change its indentation. Sentence boundaries use the same representation because they become authored breaks on the next pass. Increasing the width therefore does not join previously formatted lines. The fixture invariants exercise this fixed point, including nested arguments and documentation margins.

A zero line_width disables width-driven breaking throughout the formatter, including BibTeX and display math. The shared width helpers resolve zero to the printer’s effectively unbounded width; structural breaks remain active. Stable uses a zero target to lower authored breaks to hard breaks instead of optimizing toward an artificial target. Making them explicit in the IR also prevents enclosing groups from flattening them away.

Display math has a separate MathWrap setting for single-formula bodies, whose default follows the effective paragraph mode. Its breaking policy keeps multiplicative terms together and allows continuations at additive operators. A top-level \mid separates the following condition from equation-chain alignment. These choices can leave a cohesive term slightly over width instead of separating a short operator fragment from it.

Statement bodies

A TikZ picture contains statements rather than prose. The curated statementBody flag identifies this behavior in the TikZ and pgfplots families. The parser wraps runs ending at a top-level semicolon in STATEMENT nodes, supplying the extent the formatter needs without parsing the full path language.

Under Reflow, each statement starts a line, and its continuation lines hang one indentation step beneath the head. The semicolon determines the statement on every pass, so width wrapping cannot change that boundary. Leading comments and comment-terminated command-only prefixes stay at body indentation. Content without a terminating semicolon retains its authored-line policy. Other wrap modes flatten the statement wrappers before laying out the body.

Breaks within a statement use semantic::tikz::statement_glue. It marks gaps that belong inside a unit, keeping a path operator with its argument, at with its surrounding operands, and a coordinate with its operation. Option lists can break after commas but retain phrases such as loop above. Comments stop these rules, and unrecognized syntax uses ordinary layout. This model belongs in semantics because a coordinate-like spelling can also be a node name or prose; interpreting it need not change the syntax tree.

The flag also asserts that whitespace between statements is insignificant. The formatter can therefore split even a glued …;\draw boundary. This is a specific permission to introduce whitespace, checked by tests/typeset/statement_seams.tex. Only curated signatures and declarations that inherit them can supply the claim. The nearest environment determines the body policy, so an itemize nested inside a node label retains its own layout. statementBody remains separate from code, which identifies .dtx macrocode and its lexer regime.

Algorithm2e uses a separate algorithm2e environment flag and argument content kind. Its formatter reads direct \; control-symbol tokens as statement terminators, keeping nested arguments and math intact. Curated control-flow arguments form indented blocks, and ordinary text spacing normalizes within each statement. This does not add parser nodes or change argument attachment. Environment-local signatures keep shared names such as \For out of unrelated prose; their complete braced argument shapes must match before the algorithm2e meaning applies. The typeset check in tests/typeset/algorithm2e.tex covers glued statement boundaries and nested control flow. A candidate environment also needs a complete control-flow call in its direct body before top-level \; tokens become terminators. Without that proof, an algorithm float can belong to the separate algorithm package, where \; retains its spacing meaning.

Reflow is safe by construction

A file extension cannot establish whether reflow is safe. Package files can contain ordinary prose, and document files can contain macro code whose spaces matter. Each formatting path instead checks the structure it is about to lay out, independently of the selected wrap mode.

Fully margined documentation environments in .dtx can be formatted as virtual LaTeX. Their DOC_MARGIN tokens remain in the CST, but lowering omits them while laying out content and restores % on generated content lines and % on empty lines. Such a block owns its margins, so surrounding documentation prose does not add a second prefix. Alignment grids measure the virtual content before the printer accounts for the margin.

This path requires the environment to own its closing line. Guards, macrocode, protected bodies, and mixed margins prevent entry. Other layout paths also refuse subtrees whose margins or guards they cannot preserve. Dropping a margin could turn a ^^A documentation comment into source content. A final detector checks whether reflow has allowed content to escape its margin and, if so, re-lowers the paragraph with byte-faithful preservation. The literal framing lines around macrocode remain intact.

Optional arguments, tables, and math spacing

Textual optional arguments fill each line at existing top-level comma-space boundaries, with indented continuations. Delimiters stay attached wherever the source has no intervening space; an existing single-space edge may become a newline. This preserves the argument’s leading and trailing space tokens. Comments retain their binding, including authored % markers beside delimiters.

A proven key-value argument instead uses a grouped layout: flat when it fits, one entry per line otherwise. Width selects the form; the number of keys or a trailing comma does not force expansion.

A comma followed by authored whitespace supplies a break opportunity. Breaking after a glued comma introduces a TeX space token, so it requires a signature that identifies the argument as a key-value list. The same permission lets mandatory groups in commands such as \pgfkeys, \tikzset, and \lstset use the segmented layout. Nested groups protect their internal commas. Because mandatory groups often hold typeset text, this classification must come from curated signatures; mechanical CWL %keyvals marks cannot establish it for a brace group.

Table layout is also formatter-owned. The renderer reads static column specifications such as {lcr} and aligns cells accordingly, falling back to left alignment for specifications it cannot model. The curated align flag normally selects grid layout, but a top-level & can establish an alignment in an otherwise unknown environment.

Math whitespace has a different safety argument: TeX discards ordinary catcode-10 whitespace delivered directly to a math list. That does not make all whitespace beneath a math node insignificant. A macro can inspect spaces in its arguments or replay them as text. The formatter therefore enters a command’s arguments only at signature-proven Math slots, leaving text, unknown, unmatched, and excess arguments unchanged. A scanned redefinition shadows a built-in with unknown domains and restores preservation. Typeset fixtures test macros that preserve argument spaces or branch on them.

Within direct math content and proven math slots, lowering uses the shared virtual-atom view. It spaces binary and relation operators, retains compound relations such as :=, and treats a binary atom without a left operand as unary. Scripts keep punctuation operators compact, as in i=1, while retaining spaces around control-word operators. Delimiter-edge gaps disappear, yielding \Gamma(x). A fully glued slash stays glued; a gap on either side becomes symmetric, as in a / b. The atom classifier’s delimiter role supplies nesting information shared with the linter.

Conditionals

A paired conditional stays flat if the whole construct fits and breaks at all dividers otherwise. If any divider is glued to adjacent content, as in \ifmmode y\else z\fi, the formatter preserves the authored bytes because adding a space can change TeX’s output. WrapMode::Preserve also retains the original line breaks.

Branch contents use the nearest non-conditional ancestor’s policy: prose contexts reflow, while group-like contexts preserve. Documentation comments remain part of the lowered content even when parsing attaches them inside the conditional. There is no separate body indentation because the CST does not establish where the conditional’s test ends and its body begins.

expl3 code formatting

Within expl3, spaces and tabs are ignored characters and ~ supplies a space token. The formatter can therefore control inter-token whitespace independently of WrapMode. Its layout follows the LaTeX Project’s style guide: separate steps on separate lines, consistent spacing and brace placement, and indentation that reveals the call structure. Naming and expandability rules belong to the linter.

A derivable argspec identifies a call and the elements its slots consume. That unit supplies a structural statement boundary that survives width wrapping. When the argument scan cannot resolve a call, the formatter falls back to its authored physical line and preserves that boundary on subsequent passes.

The lexer and formatter share toggle-name recognition, but the formatter also requires a toggle to be a top-level statement before taking control of layout. A toggle mentioned as data may never execute. The lexer’s broader recognition can still produce a lossless tree, whereas formatting that region as expl3 could remove meaningful spaces.

Expl3 conditionals expose their T and F branches as attached groups. A conditional at the start of a statement places each branch on its own line, one indentation step beneath the call. A trailing conditional expands only when width requires it. Both decisions read the same attached structure, including calls with intervening single-token arguments such as \int_compare:nNnTF {a} = {1} {T} {F}.

Line endings

The printer builds output with \n. A final pass applies FormatStyle::line_ending: auto follows the source, lf and crlf select a fixed spelling, and native follows the platform. Keeping this separate from layout prevents line-ending preferences from influencing break placement.

The lexer treats CRLF as one physical line ending, including when a backslash captures it in a CONTROL_SYMBOL. LF and CRLF thus produce the same token-kind structure while their trees retain the original bytes.

Line-ending normalization also applies to protected regions. Otherwise a CRLF source could retain CRLF inside verbatim bodies while the printer emitted LF everywhere else. Only line terminators change; the rest of each protected region is preserved.

Comment directives

badness_parser::directives resolves suppression comments into sorted, non-overlapping byte ranges for formatting and linting. The shared resolver lives in the parser crate because it is pure tree analysis needed by both consumers.

% badness-format, % badness-lint, and % badness select formatting, linting, or both. They share the verbs skip, off, on, and skip-file. The retired % badness-ignore spelling remains supported. BibTeX carries lint directives in @comment{...} entries because % does not form a line comment between entries.

The resolver records each directive’s range and outcome, including dangling skip, unmatched on, unclosed off, and unsupported forms. The inert-suppression rule reads these results instead of parsing comments again. Directive-like text on .dtx documentation margins, and format directives in BibTeX carriers, are recorded as unsupported without suppressing anything.

Suppression requires containment in the resolved range. Mere overlap could otherwise suppress an ancestor containing most of the document. Anchors follow skip_target and are clamped at the previous directive boundary, keeping adjacent regions distinct. The formatter emits suppressed nodes from source; placement may adjust first-line indentation, but their interior bytes remain intact.

The linter

The linter reads the shared CST and semantic model without consulting ambient machine state. Each rule supplies the description and examples used to generate the LaTeX and BibTeX rule references.

Rules and dispatch

A Rule declares a stable kebab-case identifier, severity, whether it is enabled by default, and whether it can emit a fix. Rules are Send + Sync, so the same registry can serve parallel CLI work and the language server’s read pool.

Rules share one traversal. A node rule subscribes to syntax kinds and runs when the driver encounters matching elements. A whole-file rule runs after the walk, using semantic or project information. A streaming rule receives elements in document order, which suits checks that track a toggle or the preceding heading. The registry compiles node subscriptions into a dispatch table indexed by SyntaxKind and reuses it across files. Rule selection and ignores are applied as a post-filter, keeping configuration out of the driver.

RuleContext assembles the file’s tree, semantic model, and available project resolution. Missing cross-file information is represented by None, leaving the corresponding rules inactive. It also shares indexes for conditional branch paths and effective text or math mode, so rules do not derive them repeatedly.

The mode index partitions token ranges into Math, Text, and Unknown. Explicit math establishes math mode, and curated positional argument domains can override the surrounding mode. Unknown commands, unmatched groups, and uncurated slots remain unknown. A rule that needs math requires Math; one that needs text requires Text. A fix whose meaning changes by mode must skip unknown regions.

Signature-dependent rules are similarly conservative. For example, missing-required-argument reads curated signatures, including environment-local meanings, and skips names redefined in the file. Bulk CWL arities and ambient package discovery cannot establish that a required argument is missing.

expl3 semantic checks

Expl3 rules share a lazy Expl3Index in RuleContext. It reads attached arguments through typed accessors and semantic slot shapes, retaining the raw argspec letters needed to check variant compatibility. A command-shaped node alone does not establish that the command executes.

The index starts in top-level code and enters recognized unexpanded definition bodies and trailing T/F branches. Other arguments stay opaque. Incomplete calls and expansion wrappers leave subsequent sibling consumption unknown. In .dtx, only macrocode bodies count as code, including implicit expl3 regions identified by colon-bearing control words.

Within known definitions, a source-mapped token view tracks parameter counts and removes one level of doubled hashes per enclosing definition. Outer parameter substitutions remain unknown. This distinguishes ##5 in a nested message from the surrounding function’s #5 without expanding macros. Character-level spans also distinguish parameter five followed by 1 in #51.

The shared facts support checks for incompatible variants, protected predicate definitions, and invalid message parameters. These rules report warnings without fixes because the intended signature or message is unknown.

Autofixes

A Fix contains one or more edits applied atomically, including across files. The apply engine is a pure function of source, fixes, and applicability flags, shared by the CLI and editor code actions. It rejects malformed or overlapping fixes. lint --fix repeats linting and application until it reaches a fixed point.

A fix must stand on its own as a raw edit: formatting will not run inside its application to repair spacing or attachment. A rule can report a problem while withholding a fix when the source does not justify a safe rewrite. For example, redundant-script-braces retains braces where removal could change binding, and around operators such as \max, whose \mathop expansion cannot serve as an unbraced script field.

Safe fixes preserve meaning and are eligible for lint --fix. Fixes that may change typeset output are Unsafe and require --unsafe-fixes or an explicit editor code action. Neither kind is responsible for satisfying line width. Inline suppression uses the shared comment directives, with an optional rule name to narrow the scope.

The language server

Editor navigation depends on the local project and TeX installation, so the language server can read metadata that the parser and formatter cannot. It keeps that information separate from parser shape and formatter signature resolution. The server uses lsp-server and lsp-types, with a synchronous main loop and thread pool that accommodate salsa’s unwind-based cancellation.

Read jobs send ordinary responses directly to the transport writer through WorkerSender. Diagnostics still return to the main loop for version checks, and client edit requests return there for request-ID allocation. Incoming jobs remain ordered through the single writer, including declaration publication.

Diagnostics supersession cancels only the old analysis snapshot through its Salsa cancellation token. The worker may already have queued a rename or another request against the updated text, so global cancellation at dispatch would discard a current response without another edit. Database writes still cancel all outstanding readers. A superseded diagnostics job may finish unwinding after its replacement starts; its completion cannot release the replacement’s slot.

The live buffer

An open document is an immutable TextBuffer containing an Arc<str>, the negotiated position encoding, and a lazily initialized LineTable. The main loop and worker jobs share it through Arc<TextBuffer>. Each job therefore keeps a consistent text and index even if a later edit has already produced a new buffer. Handlers use line_index() to share the table for that document version.

Cross-file citation renames obtain a TextBuffer from the salsa file_buffer query, keyed by source file and negotiated encoding. The query shares the snapshot’s text allocation and invalidates when that text changes. It keeps the table and text together even for a bibliography with no open editor buffer. Rename requests create URI and index data only for files with matching keys; repeated requests reuse those indexes without rescanning the source.

LineTable stores line-start offsets and a flag identifying lines with non-ASCII bytes. LineIndex pairs that table with its source text and answers position queries. ASCII columns are byte distances; UTF-16 queries scan the relevant line when necessary. The table must always describe the paired text, since a mismatched pair can return incorrect positions without failing.

When an edit produces a new buffer, LineTable::patch updates an initialized table. Starts before the edit remain unchanged, starts after it shift by the byte delta, and the boundaries are rescanned. Badness treats bare \r as a line ending, so boundary checks must account for a CRLF pair being split or joined. Inserting x into a\r\nb between \r and \n, for example, creates an additional line. A table that has not yet been requested stays lazy. Debug builds compare every patched table with a fresh scan.

Read jobs check the captured text against the database through text_is_current. Worker writes are processed in order; only text-free analysis requests can coalesce. The edit chain used for incremental parsing travels with the buffer update, as described under Intra-file reparse.

Expl3 completion

The pinned TeXstudio expl3-commands.cwl contributes a sorted expl3Names list through task cwl:sync. The build script embeds this list as static data. It supplies names only: expl3’s single-token and parameter-text arguments do not fit the brace-based signature database, so this catalog never assigns formatter or linter signatures.

The shared semantic::expl3::calls reader identifies complete calls, literal operands, and recognized executable bodies. Both semantic lint rules and the completion symbol collector use it. The collector records functions, variables, constants, conditional forms, and generated variants without interpreting macros or asserting execution order. Separate Salsa queries cache the position-free symbol sets and merge the current file with transitively loaded local packages and classes. Name-preserving edits backdate the merged symbol query.

A separate completion mode index follows the lexer’s toggle names and .dtx macrocode mode transitions; it does not use the formatter’s layout ownership gates. Built-in expl3 suggestions and local names requiring _ or : appear only at expl3 command positions. Command completions carry explicit replacement edits for the full name, excluding the backslash, using the live buffer’s negotiated position encoding. Fresh-parse fallbacks use the same collector and mode classifier for the captured document.

Project and installation data

Shipped CTAN metadata maps package names to descriptions and catalog identifiers for hover and completion. The optional TEXMF index adds installed .sty, .cls, and .dtx files for links, definitions, and completion. It discovers roots with kpsewhich -var-value and caches the index under a distribution fingerprint. The index is controlled by editor settings and never supplies formatter signatures.

Existing .aux files provide label numbers and table-of-contents entries. A dedicated line scanner reads them and follows \@input chains; the LaTeX parser is unsuitable because these files are written under \makeatletter. Modification time and length detect fresh compile results without requiring a watcher. This data enriches label hover and document symbols, while the formatter remains independent of it.

Bibliography resolution separates pure extraction of resource names from path lookup. A local file wins. Plain BIBINPUTS and TEXBIB entries provide a fallback, and kpsewhich --progname=bibtex --format=bib can resolve the full Kpathsea grammar. The CLI loads the result as a citation dependency. The LSP publishes the mapping from the written path to the actual path as an explicit salsa input, so queries depend on resolved data rather than reading the environment themselves.

Citation completion returns the bibliography namespace with a filterText containing each key, title, and author list. The client can then match any of those fields using standard LSP filtering.

Badness does not run TeX engines or parse .synctex.gz. Forward search launches a configured PDF viewer in response to a user action, without blocking a read worker while the viewer runs. Filesystem paths and document URIs pass through uri_to_fs_path and path_to_uri to retain Windows drive handling.

Validation

The parser and formatter have separate correctness obligations. Parser tests require byte-for-byte reconstruction even for malformed input. Formatter tests require unchanged non-trivia content, preserved protected regions, and idempotence. Incremental reparsing adds exact agreement with a full parse, including diagnostics. These properties are checked together where the subsystems meet.

Two broader checks address what those invariants cannot prove. The texlab differential oracle compares simplified tree structures over real source corpora. Differences need an explanation, but texlab is a reference rather than a required byte target. Typeset fixtures test whether whitespace changes alter TeX’s output, something CST comparisons cannot determine. task typeset:check compiles fixtures before and after formatting and compares the results. It is required when changing key-value signature behavior or optional-argument lowering and runs separately from default CI.

Changelog

0.27.0 (2026-10-08)

Features

  • linter: flag extra math linebreaks (ddd4895), closes #203

Bug Fixes

  • formatter: separate multiline math siblings (d8d4531), fixes #204
  • formatter: put intertext on separate lines (7da63df), fixes #202
  • formatter: split leading labels in math grids (561dff9), fixes #201
  • formatter: format CAS affiliations as keyvals (9b550a9)

Dependencies

  • updated crates/badness-formatter to v0.10.2
  • updated crates/badness-parser to v0.11.1

0.26.0 (2026-10-06)

Features

  • docs: add homepage link headers (a5b01eb)
  • docs: serve markdown to agents (9b97446)
  • lint lonely items in document body (7b73d15)
  • config: support extending config files (a376f19)
  • lsp: distinguish symbol completions (92ed0e4), closes #190
  • lsp: show command signatures in completion (9d5d913), fixes #190

Bug Fixes

  • tests: avoid Nix build race and panic wording (a0eefb2)
  • linter: infer equation ranges across files (7fb6654), closes #195
  • lint: count referenced subequations labels (516a421), refs #195
  • formatter: preserve Beamer overlay body lines (1576e17)
  • docs: omit decorative logos from markdown (7e82888)
  • docs: scope worker to content pages (f25e029)
  • docs: keep markdown chapter links usable (246d691)
  • constrain label completion to key range (19230cc)
  • replace full reference completion keys (cc2ab0e)
  • lint: honor simple equation reference ranges (1b43349), refs #195
  • format: normalize algorithm2e statement spacing (aaff9b9), fixes #193
  • lsp: broaden command completion kinds (5474ca2), refs #190

Dependencies

  • updated crates/badness-formatter to v0.10.1
  • updated crates/badness-parser to v0.11.0

0.25.0 (2026-09-27)

Features

  • lsp: support source file and folder renames (8640b6d)

Bug Fixes

  • formatter: keep environment argument values intact (8b50184)
  • formatter: fill textual optional arguments (34b86fe)
  • docs: repair social previews and accessibility (ced6e51)
  • lsp: publish IPC advertisements atomically (b124210)
  • update salsa (c65ca08)
  • linter: allow implicit braces (21389da), fixes #186
  • match exclusions across symlinked project roots (8dc3b45)

Dependencies

  • updated crates/badness-formatter to v0.10.0

0.24.0 (2026-09-21)

Breaking Changes

The semantic wrap mode now considers line-width and reflows if the line width is exceeded. To get the old behavior, set

[format]
wrap = "semantic"
line-width = 0 # disable line width limit

Features

  • lsp: add expl3 completion (17ab28b)
  • format: honor width in semantic wrapping (fc6feee), closes #183
  • lint: add expl3 semantic checks (7e1ad2c)

Bug Fixes

  • linter: avoid uncertain duplicate warnings (3dfe033), fixes #185
  • fix rename race (b6d7b7f)
  • parser: capture xparse c environment bodies (a3558a5), fixes #184
  • lint: accept exam question parts without arguments (3e4ff7c), fixes #181
  • formatter: preserve glued opaque environments (cab55f6)

Performance Improvements

  • lsp: reduce repeated hover lookup work (6ef4bf8)
  • lsp: reduce configuration validation overhead (ae34296)
  • lsp: reduce warm rename overhead (132a112)
  • build: enable ThinLTO with one codegen unit (2e9414c)

Dependencies

  • updated crates/badness-formatter to v0.9.0
  • updated crates/badness-parser to v0.10.0

0.23.0 (2026-09-03)

Features

  • config: publish JSON schema (f83cacd)
  • lint: support opt-in rules (8cc742a)

Bug Fixes

  • linter: skip positional verbatim arguments (08b5243), fixes #175
  • linter: ignore Beamer overlay ranges (d57390f), fixes #174
  • formatter: stabilize long optionals (04d8ecf)
  • parser: parse long mixed optionals (38adc16)
  • formatter: format empheq as math (e7a2140), fixes #172
  • linter: retain braces around math operators (946bf2b)

Dependencies

  • updated crates/badness-formatter to v0.8.3
  • updated crates/badness-parser to v0.9.0

0.22.1 (2026-08-31)

Bug Fixes

  • formatter: attach citations to sentences (57e9bb8), fixes #163
  • formatter: normalize commented env args (6154eb9)

Dependencies

  • updated crates/badness-formatter to v0.8.2
  • updated crates/badness-parser to v0.8.1

0.22.0 (2026-08-28)

Features

  • lsp: add Beamer frame symbols (11f2a76)

Bug Fixes

  • accept Windows short paths in BIBINPUTS (3a243cd)
  • resolve bibliography search paths (c867a8e), fixes #161
  • formatter: align macro statement prefixes (4e7efa5), fixes #158
  • formatter: prioritize math breakpoints (9f85294)
  • formatter: preserve postfix limit signs (658ef51)
  • loosen test (5a6eeb8)

Dependencies

  • updated crates/badness-formatter to v0.8.1
  • updated crates/badness-parser to v0.8.0

0.21.0 (2026-08-27)

Features

  • formatter: close display math lines (7394f25)
  • formatter: glue inline command arguments (f1093ff)
  • formatter: split environment keyvals (3ea1412)
  • editors: add Zed extension (c898e44)

Bug Fixes

  • formatter: preserve BibTeX lint directives (e0567a9), fixes #159
  • formatter: split commented begin tails (5d30318)
  • formatter: preserve authored line-break rows (4886414)
  • lint: remove \paragraph rule from headings lint (0343e90)
  • lint: retire mismatched delimiter rule (a034f0b)
  • formatter: keep labels with headings (7ba5fa5)
  • formatter: match omitted environment slots (c00340f)
  • formatter: honor control-word boundaries (4ee6fc6)
  • formatter: recognize unary math signs (6afa14f)

Performance Improvements

  • formatter: cache group break state (4a37963)

Dependencies

  • updated crates/badness-formatter to v0.8.0
  • updated crates/badness-parser to v0.7.0

0.20.0 (2026-08-26)

Features

  • formatter: configure item indentation (b869f10), closes #150
  • formatter: indent commented begin arguments (1da8350)
  • report inert suppression directives (70e57f7)
  • lint: check minipage caption labels (6ea84bd)
  • lint: flag indented docstrip guards (9aabc58)
  • drop MSRV to 1.89 (b5c6250)
  • linter: warn on retired suppressions (7c30f9d)

Bug Fixes

  • formatter: close environment lines (c72d71d)
  • formatter: treat gathered as math (b31668c)
  • formatter: align nested math environments (eeebbbf)
  • support Rust 1.94 (a7afd67)

Dependencies

  • updated crates/badness-formatter to v0.7.0
  • updated crates/badness-parser to v0.6.0

0.19.0 (2026-08-25)

Features

  • linter: check labels before list items (779e9cc)
  • lint: detect invalid macrocode frames (9be2612)
  • ci: scan arXiv source projects (c5e256b), ref #152
  • formatter: separate section headings (a75a46e)
  • collect labels from environment options (df6f1f3)
  • add command declarations (c9effb8)
  • lsp: add table column refactor (def2998)
  • formatter: align row terminators (7277610)
  • lint: detect extra alignment tabs (1129686)

Bug Fixes

  • formatter: keep scripted colon relations tight (77833dc)
  • formatter: preserve colon relations (dbb64a2)
  • lsp: refresh cached config (27f22d0)
  • parser: close citation command set (e10aa54)
  • linter: handle nested subfigure labels (bb0fed5)
  • lsp: scope range formatting to selection (4439eeb), fixes #149
  • formatter: wrap long citation lists (8dbca9e)
  • honor curated ref/cite families in declarations (a3aa4ba)

Performance Improvements

  • firewall declarations by parse and semantic tier (ef61239)

Dependencies

  • updated crates/badness-formatter to v0.6.0
  • updated crates/badness-parser to v0.5.0

0.18.0 (2026-08-24)

Features

  • add LSP memory benchmark (9a975d7)
  • formatter: glue item overlays (e7b3efe)
  • parser: reparse math fragments (4638eb5)
  • formatter: unify math spacing (170da11)
  • add semantic math atom classification (3889311)
  • parser: add argument domains (f7f4b01)

Bug Fixes

  • parser: consume CRLF control symbols atomically (20dd182)
  • parser: parse alignment char constants (78f3cd9)
  • parser: guard expl3 mode boundaries (08c0387)
  • formatter: preserve trailing control newline (2a7977b), fixes #141
  • lsp: replace project snapshot generations (a1194cb)
  • lsp: gate query logging (f212680)
  • parser: tighten catcode signal (834c7b6)
  • parser: pass plain braces through optionals (82b30f1)
  • parser: parse href URLs verbatim (03ad84f)
  • formatter: preserve glued forced blocks (4938f6c)
  • formatter: align virtual dtx tables (e0a2cde)
  • parser: preserve argument mode semantics (63aefa8)
  • formatter: bound relation alignment (a560698)
  • formatter: preserve math line breaks (a1ad967)
  • parser: honor TeX script atom boundaries (e6131b6)

Dependencies

  • updated crates/badness-formatter to v0.5.0
  • updated crates/badness-parser to v0.4.0

0.17.0 (2026-08-20)

Breaking changes

  • parser: pair one-sided environment aliases (d757cdc), closes #117
  • parser: arity-directed expl3 attachment (#119) (5f2f9d8)

Features

  • formatter: format DTX doc environments (ac98c95), fixes #127
  • parser: intra-file incremental reparse (#130) (393e0c3)
  • bench: time the keystroke pipeline (b1b2b0e)
  • parser: pair one-sided environment aliases (d757cdc), closes #117
  • parser: arity-directed expl3 attachment (#119) (5f2f9d8)
  • formatter: wrap picture statements at TikZ unit boundaries (5079f96)
  • formatter: hang statement continuations in picture bodies (5266aba)

Bug Fixes

  • parser: isolate command definition names (8d6e274), fixes #133
  • formatter: preserve TeX line semantics (f16188b), fixes #132
  • formatter: handle mixed dtx doc regions (6edfd81), fixes #126
  • preserve dtx documentation math (de0c54c), fixes #138
  • formatter: preserve guarded dtx paragraphs (0373420), fixes #123

Performance Improvements

  • lsp: share document text and line index (3d4a5f8)

Dependencies

  • updated crates/badness-formatter to v0.4.0
  • updated crates/badness-parser to v0.3.0

0.16.0 (2026-08-14)

Features

  • config: declare environments in badness.toml (#115) (a80b5af)
  • formatter: lay out picture bodies as statements (b437091), closes #114
  • linter: add % badness-lint suppression directives (c03114d), refs #114
  • formatter: add suppression comment directives (1810cde), refs #114
  • linter: add blank-line-in-keyval (0758ea4)
  • formatter: segment a mandatory keyval group (d79ec73)
  • formatter: width-driven layout for opaque brace groups (03022a8)
  • cli: add the strict trivia-invariance check (3796d7a)
  • linter: add label-before-caption rule (e8c5b7f)
  • linter: colorize the pretty lint report (e4faf16)
  • parser: pair user-defined environment delimiters (2bbff60), closes #109
  • cli: read stdin from -, not a bare terminal (5bb4788), closes #111
  • formatter: lay conditionals out all-or-nothing (ed84bfe)
  • parser: gated CONDITIONAL node for \if…\else…\or…\fi (e0ca4ef)
  • bib: parse and preserve % comments (e005cc9)
  • skill: add formatter-fixture for construct coverage (a52095a)
  • lsp: add inverse search over IPC (3b829eb)
  • lsp: add textDocument/forwardSearch (8028a40)
  • formatter: explode sibling-attached expl3 branches (d3fc51a)
  • formatter: expand optional arguments to the width (4c28ba4)
  • formatter: add optional serde and schema features (80726c7)
  • formatter: reflow doc-margined out-of-region expl3 runs (aa9445a)
  • formatter: reflow dtx prose around margined blocks (4f118a2)
  • semantic: add curated block-level command property (c54a5ff)

Bug Fixes

  • formatter: break a keyval group’s glued opener (507a982)
  • scripts: keep pdflatex stdout off hyperref’s .out (1e7c4c1)
  • formatter: body a \begin tail past the declared arity (bd7028e)
  • parser: break every gate’s run at a docstrip guard (d682e8f)
  • formatter: guard a prose argument’s edge comments (8976815)
  • formatter: break around curated block-level commands (09b8d4f)
  • linter: pair straight quotes into one finding (e1ff0d1)
  • parser: harden environment-alias pairing (84f11a2)
  • project: link a subfile to its parent document (df3e66c), closes #112
  • lsp: spell decoded URI paths with native separators (f199bd1)
  • formatter: break around sectioning commands (f4be809)
  • formatter: stop deleting ] inside prose arguments (7d2799f)
  • formatter: make optional fallbacks deterministic (9d2095c)
  • formatter: correct what lower_conditional assumes (902dbd9)

Performance Improvements

  • cli: use Histogram for the --check diff, measured on the corpora (29a678a)
  • cli: pick Patience over Histogram for the --check diff (595abb3)
  • cli: diff --check with Histogram, write it buffered (f34abb8)
  • lint: render pretty snippets from a line window (9dffc3d)
  • parser: one batch driver for all nine shape gates (#113) (9e01ee5)
  • parser: bound the environment-alias closer scan (ae83909)
  • formatter: gate the doc-margin scans on cx.is_dtx (4e7babf)
  • parser: answer on_doc_margin_line from a pre-scan (930380b)

Dependencies

  • updated crates/badness-formatter to v0.3.0
  • updated crates/badness-parser to v0.2.0

0.15.0 (2026-08-07)

Features

  • formatter: default every file kind to reflow (ba9f2f9)
  • formatter: add a line-ending style (373a16c)
  • wasm: add badness-wasm playground shim crate (6486890)

Bug Fixes

  • formatter: hug detonating atoms in fallback fills (db63ddb)
  • formatter: accept a relation as an expl3 N slot (4a3d92b), closes #106
  • formatter: gate the expl3 forced-break dispatch in fallback lines (7437f69)
  • linter: skip parameter-template keys in key scans (928aa4a), closes #104
  • formatter: keep fitting math segments flat (903ec3e)

Dependencies

  • updated crates/badness-formatter to v0.2.0
  • updated crates/badness-parser to v0.1.1

0.14.0 (2026-08-06)

Features

  • cli: diff changed files in format --check (eed1537)
  • semantic: curate filecontents and ltxdockit verbatim envs (d805548), closes #98
  • formatter: respace flush expl3 argument braces (c44125c)
  • formatter: add trivia-perturbation invariance oracle (#103) (f22d668)
  • packaging: publish badness-bin to the AUR on release (82a64c2)
  • lint: add --output json machine-readable findings (92b04bd)
  • formatter: explode expl3 conditional branches (R4) (bebdbde)
  • parser: recognize package-defined verbatim envs (696f109)
  • build: bundle man pages and completion in tarballs (8268c88)

Bug Fixes

  • formatter: render expl3 conditionals all-or-nothing (c8b3aef)
  • formatter: keep annotated expl3 branches on the exploded path (ac07506), refs #101
  • formatter: drop expl3 sibling break coupling (836ed83), closes #101
  • packaging: tolerate pre-0.14 release tarballs in PKGBUILD (750493b)
  • installer: detect musl/libc in installation script (eb195a1)
  • formatter: pin forced expl3 block body to break mode (349ccdd)
  • formatter: stabilize trailing expl3 hang group (72a6d35), closes #96
  • npm: fall back to musl when glibc build fails (bd1891e)
  • formatter: stabilize trailing expl3 conditional (b5ed902), refs #96
  • parser: pair \left/\right inside macro code (20a59ef), closes #95
  • parser: bound math bracket gate at dollar closer (703f5f5), closes #99
  • formatter: keep expl3 parameter runs tight (47d6277), exception #1 and #2
  • formatter: sticky-break fill for expl3 statements (107ecb0), closes #94
  • linter: stop TikZ/pgf false positives (aca172f)
  • linter: drop “Part”, gate hard-coded-reference item labels (274c124)
  • linter: gate space-before-command on trailing break (2f66cae)
  • linter: skip hex constants, font maps in straight-quotes (2fb9ee3)
  • linter: skip redefined commands in deprecated/primitive (0490f16)
  • linter: skip starred headings in sectioning-level-jump (befe30e)
  • parser: parse array and tikzcd bodies as math (9db4e8f)

0.13.0 (2026-08-01)

Features

  • semantic: curate codeexample as verbatim env (4fb98f7)
  • lexer: infer expl3 in toggle-less .dtx (caba767)
  • lsp: link package docs via texdoc in hover (40b2665)
  • lsp: filter cite completion by title and author (509cd4a)
  • linter: express and apply cross-file fixes (e9071f8)
  • vscode: add feature toggles for the LSP (bb3f13c), closes #86
  • formatter: align user environments on & (ee88953), closes #84
  • formatter: hang expl3 attached brace arguments (c25c91b)
  • formatter: hang \item continuations under preserve (5edeea7), closes #82
  • incremental: mark SourceFile.path HIGH durability (bca096a)
  • linter: flag unclosed math delimiters as likely typos (4731c18)

Bug Fixes

  • parser: parse .code.tex under package flavor (2694a86)
  • linter: withhold deprecated-command fix in reference position (668647a)
  • linter: skip compound-logo swallowed space (0a546b7)
  • linter: ignore \string-prefixed package loads (e20504e)
  • linter: skip \texttt dashes and \foreach ranges (f605ee5)
  • linter: skip citation locators and env titles in hard-coded-reference (620bde5)
  • linter: skip xypic @ DSL in makeat-macro (81e4c81)
  • parser: attach math optional across balanced $...$ (b5c6c50)
  • formatter: scope preserve spacing collapse to prose (0999610)
  • formatter: normalize inner spacing under preserve (be8d5ba)
  • linter: skip prose rules in Lua code and doc placeholders (bcd4f44)

Performance Improvements

  • line-index: precompute wide-char table, own no text (d9c526f)

0.12.0 (2026-07-30)

Features

  • formatter: break leading \label onto its own line (18b17c5)
  • add stable-diff paragraph wrapping (#41) (5734632)
  • format: indent expl3 continuation groups one step (d53f064)
  • config: add BADNESS_CONFIG env var for config path (b8756e8)
  • cli: add hidden debug format check command (da5959f)
  • config: add global user config fallback (41c4570), closes #40
  • formatter: add math-wrap display-math break policy (dbba5eb), closes #42

Bug Fixes

  • formatter: pin inter-argument docstrip guards to column 0 (3556c79), closes #78
  • formatter: keep stable-wrap break mask aligned on Nil atoms (d4211b1)
  • parser: shape-gate unclosed \left instead of erroring (29aa319), closes #77
  • parser: close the latex2e format-error buckets from the smoke test (#80) (bb9484e)
  • parser: stop the lexer hiding braces, guards, and short-verb bars (#79) (a6e2119), refs #71
  • parser: stop environments escaping their brace group (#75) (7d82f35), refs #71
  • parser: scope the math gates’ paragraph-break anchor to the body’s own level (#74) (ef89b76), closes #70
  • gate expl3 relayout to top-level toggles (81a1a92), closes #69
  • parser: suppress braced verbatim on redefinition (513e963)
  • formatter: preserve fully-guarded expl3 chunks (5d2e46b), closes #72
  • parser: treat escaped backtick char constant as data (d7edf4b), refs #71
  • format: keep glued brace opener on its line (97e7abb)
  • format: treat sign after opener as unary in math (94af6dd)
  • format: keep guarded expl3 code groups broken (417c480), closes #61
  • format: keep doc-margin math environments verbatim (d8e3864), refs #61
  • format: make trailing comments zero-width in expl3 code (db53fc8)
  • format: keep doc-commented expl3 statements whole (fb61e15)
  • parser: treat expl3 regions and v-arg names as macro data (e9deecf), issue #60
  • parser: add ^^A, v-arg, and char-constant lexing (ff2b516)
  • parser: gate \[/\( and isolate \def-family names (bb55149), closes #65
  • lexer: accept comment tail on macrocode end frame (fa01c29), closes #62
  • parser: gate \begin/\end on a name-shaped group (0dfeb0b)
  • parser: gate text-mode brackets on a reachable closer (47f92f9), refs #60
  • lexer: add l3doc to curated doc classes (e34cd7f)
  • parser: gate dollar math on a reachable closer (4a01a6b)
  • formatter: feed run-final trivia to expl3 run separator (ad2ce81), closes #58
  • formatter: own only macrocode bodies in dtx expl3 regions (d0abf21)
  • formatter: refine expl3 group style in regions (1cda543), refs #57
  • parse doc short verbs and chunked dtx macro code (37fbf9c), refs #57
  • formatter: strip script braces before operator atoms (d627794), closes #56
  • parser: count nested optionals in math bracket gate (7b98255), closes #55
  • parser: widen definition bodies to hooks and \newcommand (ab7d809)
  • formatter: keep trailing comments riding their line (71baa30), closes #54
  • parser: make verb delimiter capture opt-in (4788af9), closes #53
  • formatter: keep own-line comments in list bodies (f468619), closes #48
  • formatter: keep grid indent on doc-commented rules (e0ccee0), closes #49
  • formatter: collapse fitting multi-line optional args (cf15d18), closes #47
  • formatter: lift \begin-line % in every env layout (9c7fe0e), closes #38
  • parser: accept split \begin/\end in env definitions (e4f711d), closes #45
  • parser: stop math brackets parsing as optional args (0eb49d5), closes #43
  • formatter: improve display-math break-point detection (9e97826), refs #42
  • formatter: keep multi-line math LHS off the relation column (b7fd62b), closes #39
  • formatter: keep trailing % glued to a block segment (5369c22), closes #38
  • vscode: swap npm-run-all for npm-run-all2 (b631714)

0.11.0 (2026-07-20)

Breaking changes

  • lsp: move texmf config to editor settings (2f83a84)

Features

  • lsp: move texmf config to editor settings (2f83a84)
  • formatter: hang nested blocks in align grids (5103bab)
  • linter: document bib rules in –explain and docs (a2a742c), closes #24

Bug Fixes

  • tests: adapt to lsp-server 0.10 Response API (5f1e889)
  • linter: skip script labels inside argument groups (1f8c0eb), closes #37
  • parser: name blank line as math terminator (2751787), ref #35
  • linter: ignore key arguments in dash-length (506e5f0)
  • linter: ignore rule-command spans in dash-length (adeecf6), closes #34
  • linter: ignore key arguments in math-shape rules (387810a), closes #25
  • parser: keep unmatched [ a plain atom in math (c185c13), closes #23
  • linter: ignore labels in exclusive conditional branches (a6e0c22)
  • linter: ignore package loads in exclusive branches (d16d3d1), closes #27
  • linter: target the whole construct for DOC_COMMENT-bound suppressions (cd647fa), fixes #26

0.10.0 (2026-07-15)

Features

  • add --force-exclude to format and lint (05c8e32)

0.9.0 (2026-07-14)

Features

  • linter: make fixes carry multiple atomic edits (7f52ef5)
  • linter: add diagnostic related information (ada447b)
  • parser: add a release-mode stuck-loop step limiter (f93b1e5)
  • lsp: add diagnostic tags and rule doc links (906b0df)

Bug Fixes

  • lsp: recover poisoned db mutexes instead of panicking (5844ff2)
  • bib: resolve field aliases in missing-required check (f20283b)

0.8.0 (2026-07-11)

Features

  • lsp: show source package in macro hover (eaf030e)
  • lint: add unknown-option rule for local packages (f040033)
  • bib: document links for doi and url fields (e1f15f7)
  • completion: argument-value enum completion (f363499)
  • bench: add linter speed benchmark vs lacheck and chktex (e6a821e)
  • bench: add whole-project folder benchmark (620c7cc)
  • bib: typed AST wrapper layer for BibTeX CST (0abe33a)
  • ast: typed AstNode/AstToken wrapper layer (35eae44)

Performance Improvements

  • cli: parallelize lint –fix across files (133a4c3)

0.7.0 (2026-07-08)

Features

  • lsp: selection ranges from CST hierarchy (2aff55f)
  • formatter: column-spec-aware table alignment (9ab94ba)
  • linter: package-aware duplicate and provides lints (758fac3)
  • semantic: recognize package metadata and options (0cd95ce)
  • lsp: color and TikZ/PGF library completion (1a881f3)
  • lint: add unreferenced-label rule (4d975a1)
  • lint: add verbatim-trailing-text rule (a11358d)
  • lint: flag line-break tie in missing-nonbreaking-space (de2d51f)
  • lint: autofix obsolete-environment eqnarray to align (aa26b13)
  • lint: add missing-required-argument rule (5206ee6)
  • lsp: references, rename, goto-def for user macros (fdeb0e9)
  • lsp: negotiate client capabilities at initialize (36b6ed2)
  • lsp: change-environment refactor command (1f27fab)
  • lsp: glossary/acronym key completion (f73f138)
  • lsp: signature help for command arguments (0c5f649)
  • lsp: label hover and symbol numbers from .aux (3efb7aa)
  • semantic: classify what a \label labels (01a8b0b)
  • project: scan .aux for label numbers and toc (ed48898)
  • config: add [build] section with aux-dir (ad8cc8a)
  • lsp: go-to-definition for include/package file arguments (99927ea)
  • lsp: resolve packages via TEXMF index and CTAN metadata (24ba5c7)

Bug Fixes

  • bib: tighten title-capitalization camelCase heuristic (91de065)
  • ci: rename aux.rs, allow option-ext MPL-2.0 (cc1c834)

Performance Improvements

  • formatter: parallelize the CLI format paths (d38b4d6)
  • linter: cache registry, stream rewalkers, parallelize CLI (5c3813a)
  • signature: bake CTAN metadata via phf, not runtime parse (f635d23)

0.6.0 (2026-07-06)

Features

  • completion: complete \usepackage/\documentclass names (2457147)
  • completion: add baked package/class name lists (ff4906d)
  • lsp: add document links (915aea6)
  • lsp: highlight matching \begin/\end pair (d643518)
  • lsp: re-indent on close via onTypeFormatting (5972340)
  • parser: parse math environments in math mode (9097be3)
  • formatter: implement sentence and semantic wrap modes (17003ba)
  • linter: add hard-coded-reference rule (da66c29)
  • linter: add sectioning-level-jump rule (6ac6def)
  • linter: add makeat-macro rule (2ae6d07)
  • linter: add space-before-command rule (36d5fa3)
  • linter: add abbreviation-spacing rule (2fea8db)
  • linter: add swallowed-space rule (c48aa20)
  • linter: add primitive-command rule (94da7ca)
  • linter: add math-operator-name rule (17cc5f2)
  • linter: add times-variable rule (52de07a)
  • linter: add dash-length rule (a6218e0)
  • linter: add straight-quotes rule for ASCII quotes (adff4ba)
  • linter: add ellipsis rule for literal … (488ebdd)
  • linter: generate rules reference from metadata (74e2234)
  • math: normalize operator spacing (36c9314)
  • semantic: keep built-in over delegating arity-0 redef (9fd50d8)
  • add title, author, date, thanks to signatures db (3c537d1)
  • formatter: stack binary chains under the relation too (0777920)
  • formatter: align relation chains in display math (e69a72e)
  • formatter: join alignment-cell continuation lines (cd3e590)
  • semantic: resolve packages to .dtx sources (249e68e)

Bug Fixes

  • parser: point unclosed-delimiter errors at the opener (1029351)
  • formatter: tight spacing and no paren breaks in display math (7112b8c)
  • linter: allow en dash between proper names in dash-length (2ab4342)
  • formatter: peel over-attached cell off table rules (7c91ac9)

Reverts

  • “feat(formatter): stack binary chains under the relation too” (4a6988b)

0.5.0 (2026-07-01)

Features

  • lsp: add range formatting support (5ad2827)
  • lsp: add workspace symbols support (eb8a111)
  • formatter: format expl3 code (catcode 9/10 model) (ac4ff31)
  • lsp: watch on-disk tex/bib/config and reanalyze (b551c01)
  • dtx: reflow documentation prose under reflow (be57646)
  • lsp: outline entries for dtx documented macros (cba0b01)
  • lsp: add textDocument/documentHighlight (404069b)
  • bench: add formatter speed bench vs tex-fmt & latexindent (82ddeb5)
  • format: reflow brace-group bodies as statements (bb976e0)
  • lsp: discover and apply badness.toml per document (e56a8af)
  • lint: add missing-nonbreaking-space (tie before cite/ref) (4d75da4)
  • lsp: surface linter autofixes as code actions (13c727e)
  • lsp: resolve completion items with signature and citation detail (f9892e6)
  • lsp: add hover for commands, environments and citations (3c6047c)
  • lsp: add pull diagnostics (a73fd7b)

Bug Fixes

  • lsp: honor excludes for siblings (7a50529)

Performance Improvements

  • signature: bake CWL tier into a build-time phf map (a920d4a)

0.4.0 (2026-06-23)

Features

  • semantic: mark the cross-reference family inline (c7c77a7)
  • semantic: ingest CWL corpus as a bulk signature tier (4740bf5)
  • lint: don’t withold lints that disturbs alignment (8ea1efc)
  • bib: diagnose missing field separator; fix value trivia attachment (e14751c)
  • bib: autofix duplicate-field when values are identical (c34bd78)
  • bib: duplicate-field lint rule (f2f6d60)
  • lsp: rename labels and citation keys (textDocument/rename + prepareRename) (7b1d01b)
  • config: badness.toml configuration (CLI) (8c68ca2)
  • project: package load graph + package signatures into scope (f8e6bc7)
  • semantic: doc/ltxdoc prose↔code association query (a52f17c)
  • file-kind: .ins installation-script support (plain code, Preserve) (85c9c7a)
  • formatter: .dtx two-layer formatting (foundation, Preserve) (6c7861f)
  • semantic: doc/ltxdoc signatures + DOC_COMMENT node (M3) (95ec2a2)
  • parser: lex expl3 syntax mode (_/: as letters) (c98e2e8)
  • parser: lex .dtx docstrip guards as GUARD tokens (M2) (b09c507)
  • parser: parse .dtx docstrip surface syntax (M0+M1) (8e54604)
  • lsp: add textDocument/foldingRange (f0ea513)

Bug Fixes

  • cli: fix file-detection in cli linter (7821b6a)

0.3.0 (2026-06-21)

Features

  • lsp: add textDocument/references (find references) (2ef3606)
  • sty/cls: format and lint LaTeX package/class sources (54692cf)
  • lsp: bib-aware completion and \cite key completion (493ad41)
  • bib: add generator to sync bib_fields.json with biblatex data model (189de08)
  • bib: align entry-type required fields to the data model (35b81d9)
  • bib: derive field/entry DB from biblatex’s canonical data model (55a6883)
  • bib: recognize the full standard biblatex field set (e2c2639)
  • semantic: flag user verbatim environments via begin-code catcode scanning (eefc1a1)
  • semantic: scan \def-defined verbatim commands and helper chains (6cad9c1)
  • semantic: flag user verbatim-argument commands via definition scanning (19ef5f1)
  • lsp: go-to-definition for refs and citations (2535199)
  • cli: –stdin-filepath routes lint stdin to the bib pipeline (f8a4831)
  • cli: –stdin-filepath routes format stdin to the bib pipeline (96f1b80)
  • lsp: cross-file project assembly — undefined-ref/citation fire live (38b7f2c)
  • bib: Phase 4 — incremental, LSP, and project-graph integration (b593bdc)
  • bib: linter rules + CLI wiring (Phase 3) (571c2d3)
  • bib: field & entry sorting (Phase 2c) (438a61d)
  • bib: value reflow (Phase 2b) — wrap long field values by category (3cfed27)
  • bib: formatter (Phase 2) — lower bib CST to shared Wadler IR (de48afd)
  • bib: semantic model + field/entry signature DB (b59befc)
  • bib: differential parse oracle vs texlab + phased roadmap (d7360b6)
  • bib: first-stab BibTeX/BibLaTeX parser (6f38675)
  • lsp: add basic completion (20903b7)
  • linter: autofix infra + dollar-display-math $$→[ fix (216f590)
  • linter: obsolete-environment, dollar-display-math, mismatched-delimiter lints (8f89b51)
  • formatter: break wide display math at top-level operators (716612f)
  • linter: cross-file label resolution + undefined-ref / duplicate-label (270a035)
  • formatter: keep appendix environment body flush like document (b1a55f7)
  • formatter: collapse cite-family key lists deterministically (d88e7e3)
  • semantic: extract unbraced \newcommand\foo definition form (f2472d5)
  • parser: bind leading comments into the following construct (0afabeb)
  • lsp: add document symbols (5547650)
  • parser: don’t wrap a lone block environment in a PARAGRAPH (b4a46fe)
  • formatter: use latexindent-style desc hang (46ab231)
  • formatter: reflow inline prose commands inline, not as blocks (5d706b2)
  • collapse blanklines into 1 (b19d8da)
  • formatter: grid-align comments and rule lines; enable tables (4cbb183)
  • formatter: lower display math as an indented block (5e2cefc)
  • parser: lex verbatim-argument commands; fix multi-line VERB formatting (73cf04c)
  • cli: add badness parse command (7735a75)
  • formatter: align itemize blocks (47a2b19)
  • formatter: don’t indent document environment (3cd0d04)
  • linter: add rule layer with duplicate-label and deprecated-command (4aaee37)
  • align & columns in align/matrix environments (d5abdca)
  • match \left … \right delimiter pairs in math (3079875)
  • add structured math model and math formatting (02802f6)
  • support argument-taking verbatim environments (ab8eb74)
  • add file-walk for formatter (1603230)

Bug Fixes

  • lsp: handle Windows file URIs in path completion (5b38f45)
  • formatter: keep a trailing % on the \begin header line (e02413f)
  • formatter: ass JSS/Sweave verbatim environments to signatures (21b5e61)
  • don’t reflow single % (be49170)
  • formatter: don’t push % to next line (de271ae)
  • formatter: fall back when an alignment cell contains a comment (918c592)
  • parser: don’t treat comment-only lines as paragraph breaks (3c83c01)
  • formatter: keep command-only lines on their own line under reflow (739a32f)
  • linter: migrate render.rs to annotate-snippets 0.12 API (602d835)

0.2.0 (2026-06-12)

Features

  • add vscode and open vsx extensions (975f1e4)
  • npm: package for npm (b3a576f)

0.1.0 (2026-06-12)

Breaking changes

Features

  • formatter: reflow signature-marked prose arguments (18c99ee)
  • lsp: ra-style writer/threadpool, cancellation, incremental sync (8628f92)
  • lsp: reuse cached salsa tree for formatting (30cd2d5)
  • implement semantic group scanning (4f5e9ca)
  • parser: model \ line break as a LINE_BREAK node (651e1c5)
  • formatter: paragraph reflow via a Wadler Fill node (0cbe264)
  • semantic: add built-in signature database (e9bf2de)
  • rename fmt to format (1fedc1b)
  • linter: add minimal badness lint command (443fa6a)
  • lsp: add minimal lsp server (7e6f4fe)
  • formatter: indent multi-line group/argument bodies (5e66038)
  • parser: differential parse oracle vs texlab (25e065c)
  • lsp: add semantic model and reference support (61707c1)
  • build project graph (cc81a29)
  • incremental: salsa harness for cached parsing (67a1948)
  • formatter: environment-body indentation (5b3d1b5)
  • formatter: whitespace normalization (first real rule) (00385eb)
  • formatter: Phase 2 formatter MVP — identity round-trip (ab2ef57)
  • parser: Phase 1 recursive-descent grammar with error recovery (511352c)

Bug Fixes

  • attach arguments to environment (a6772d2)
  • parser: stop $-math at group and \end anchors (1319fd8)