Engineering plans / Proposed · October 4, 2026

Rust hot paths: native kernels for ingest

Garage’s ingest is Python end to end: walk, extract, gate, chunk, attribute, store. Profiling it shows that three small, self-contained pieces of that path take most of the CPU the pipeline spends per document, and that each of them is a pure function of bytes with no I/O, no database and no network. They are the shape of code that Rust does well and Python does badly: character-by-character scans and regex passes over text, and glob matching over every filename a walk sees.

This plan moves those three pieces into one Rust library, libgarage_native, called from Python through ctypes the way libtesseract and libpq already are. The Python implementations stay as the reference and as the fallback: a venv without the library, a PyPI install, or GARAGE_NATIVE=0 runs the pure-Python path, and a conformance test holds the two to identical output. The chunk goldens that already pin the splitters’ output byte for byte are the acceptance test, so the stored chunker ids, and with them every vector in the corpus, survive the port.

The plan deliberately stops where the measurements stop. After these ports the ingest is bound by the gRPC round trip and the Postgres transaction each document already pays, and by the PDF extractors, none of which a Rust kernel changes. Section 2 lists what is not in scope and why.

1. Where the time goes

1.1 Method

Two measurements, both on 2026-10-04 against main at db351e8, CPython 3.14.7, a Linux x86-64 container. Absolute numbers will differ on an Apple silicon Mac; the shares and ratios are what this plan rests on, and milestone 0 re-takes them on a Mac before any Rust is written.

  1. A whole ingest under cProfile. ingest_source() over tests/corpora/second-brain-os (250 Markdown files, 1.3 MB) with the RecordingGateway from tests/test_second_brain_corpus.py, so storage costs nothing and what remains is the pipeline’s own CPU.
  2. Each candidate in isolation. time.perf_counter, best of three, on synthetic inputs sized like real files: generated prose, Markdown with headings and fenced code, Python classes, a single-line minified JavaScript bundle, timestamped log lines, and a synthetic tree of 30k files.

Milestone 0 commits the harness as tools/bench/ so these tables can be regenerated with one command.

1.2 Share of ingest CPU on the corpus

Stage Share of the run Notes
extract/quality.py::assess 25.6% _alpha_ratio alone is 19.5%: a Python generator over every character of a 64 KB sample
ingest/chunking.py::chunk_text 17.4% MarkdownHeaderSplitter.split_text is 12.7%: "".join(filter(str.isprintable, ...)) per line
attribute/resolver.py::resolve 16.9% pathrules.classify_path 8.7%, git.find_repo_root 7.8% (a stat per ancestor, per file)
extract/dispatch.py::extract 16.8% reading, YAML front matter (8%), normalize_text
scan plus walk 14.4% the tree is walked twice, once to count and once to ingest

Per document that is about 2.6 ms of Python. For comparison, a real ingest also pays one gRPC CheckDocumentStat round trip and one Postgres transaction per document, so the Python share is roughly half of the per-document cost today and the I/O half is untouched by this plan.

1.3 The candidates in isolation

Function Input Time Throughput
quality.assess 100 KB prose (64 KB sample) 6.0 ms 11 MB/s
quality.assess 64 KB of log lines 9.1 ms 7 MB/s
chunk_text MARKDOWN 100 KB 7.7 ms 13 MB/s
MarkdownHeaderSplitter.split_text 100 KB 4.4 ms 23 MB/s
chunk_text PROSE 100 KB 1.0 ms 97 MB/s
chunk_text CODE (.py) 100 KB 0.9 ms 113 MB/s
chunk_text CODE (.js, one line) 50 KB 60 ms 0.8 MB/s
RecursiveSplitter on one unbroken token 160 KB 194 ms 0.8 MB/s
normalize_text (for scale) 100 KB 0.27 ms 367 MB/s
sha256 (C, for scale) 100 KB 0.24 ms 424 MB/s

The last code row is the splitter’s fallback to the empty-string separator (splitters.py, _split_keeping_separator): a line with no space becomes one piece per character, and _merge then pops them one at a time with current = current[1:]. It is linear, but a hundred times slower than the normal path, and code trees are full of such files: bundled JavaScript, SVG path data, lock files, base64 blobs in source.

1.4 The walk

A walk over a synthetic tree of 30k files (24k of them indexable), under cProfile:

  Time Share
walker.walk in total 3.28 s 100%
_matches_any (the diagnostic patterns) 2.13 s 65%
of which fnmatch.fnmatch 1.51 s 46%
posix.stat 0.21 s 6%
os.walk itself 0.07 s 2%
os.walk plus a stat per file, no predicates 0.19 s the syscall floor

Each filename is matched against the twenty DIAGNOSTIC_FILE_PATTERNS with fnmatch, which lower-cases and compiles on every call; each directory name against DIAGNOSTIC_DIR_PATTERNS and DEPENDENCY_PATH_FRAGMENTS through pathlib. Per file that is 13 µs for is_diagnostic_file and 16 µs for is_dependency_dir, against 7 µs for the stat. The scanner (ingest/scanner.py) runs the same predicates over the same tree before the ingest does, so a 192k-file source pays this twice.

1.5 What is not hot

Measured and found fine, so left alone: normalize_text (367 MB/s), the SHA-256 hashing (C), git log parsing (430k lines/s, once per repository), Messages attributedBody decoding (about 1 µs per message), Reciprocal Rank Fusion (it runs in SQL), protobuf (the upb backend).

Two things that are hot but belong to other plans: PDF extraction runs at 175 pages/s under pypdf and 5 pages/s where it escalates to pdfplumber, and the backfill renders each vector to text with str(float) per dimension (0.5 to 1.7 ms per chunk). The first is a library choice, the second a binary adaptation in psycopg; neither is a kernel to port.

2. Scope

In scope, in the order the measurements rank them:

  1. The quality gate, extract/quality.py::assess. A pure function of a string.
  2. The splitters and chunking, ingest/splitters.py (RecursiveSplitter, MarkdownHeaderSplitter, the per-language separator tables) and ingest/chunking.py (_locate, the char offsets). Pure functions of a string and a few integers.
  3. The walker’s per-file pass, ingest/walker.py: directory pruning, the filename predicates, the stat, the cloud-placeholder check (extract/placeholder.py), and the WalkStats counters, as one native pass that returns candidate records. Shared with scanner.scan_filesystem.
  4. Optional, last: the LangExtract alignment in enrich/langextract/resolver.py (_best_lcs_spans, a pure-Python triple loop costing 3 to 40 ms per fuzzily aligned fact) and tokenizer.py::RegexTokenizer.tokenize (0.9 MB/s). Textbook kernels, but the facts pass is bound by local LLM inference at seconds per chunk, so they are a few percent of its wall time. Measured first, ported only if a real facts run shows otherwise.

Out of scope, with the reason:

3. Design

3.1 One library with a C ABI, called through ctypes

The alternative is a PyO3 extension module. It loses on four counts here:

Crate: native/garage_native/ (Cargo.toml, Cargo.lock, src/), built by //native:garage_native as rust_shared_library, producing libgarage_native.dylib / .so / .dll. Edition 2024, as the toolchain declares.

3.2 Loading and the fallback

garage_rag/native/__init__.py grows accelerator() -> ctypes.CDLL | None, cached:

  1. GARAGE_NATIVE=0 in the environment: None, always. Tests and bug reports use it to force Python.
  2. loaded_library("garage_native"): the copy dyld already mapped (the app, its XPC services, the launchers).
  3. GARAGE_NATIVE_LIBRARY=<path>: an explicit path, for CI and development.
  4. Otherwise None.

Before it is used, the library’s garage_native_abi_version() must equal the constant in the Python module; a mismatch logs once and returns None. A stale dylib must never produce different chunks than the Python it ships with.

Each ported function keeps its Python signature and body, and gains a first line:

if (lib := accelerator()) is not None:
    return _native_assess(lib, text, sample_bytes=sample_bytes)

The Python body below it is the specification. A behavior change lands there first, with its golden update, and the Rust side follows. The library can be removed from a build and nothing breaks but the speed.

3.3 Data across the boundary

Text goes in as UTF-8 bytes with a length. A Python str holding lone surrogates cannot be encoded; that document takes the Python path (the extractors never produce one, but the fallback costs nothing).

3.4 Semantics that must survive exactly

This is where a port goes wrong, so each item gets a test in milestone 0 before any Rust exists.

3.5 Build, packaging and CI

4. Milestones

M0: baseline and conformance harness, no Rust

M1: the crate, loading and packaging

M2: the quality gate

M3: the splitters and chunking

M4: the walk

M5, optional: LangExtract alignment

5. Risks

6. Decisions to confirm

  1. C ABI through ctypes rather than PyO3 (3.1). Recommended: C ABI.
  2. Crate name and place: native/garage_native, libgarage_native.
  3. Whether M0’s Python fixes land first. Recommended: yes; they are cheap, they stand on their own, and they decide how much of M4 is worth building.
  4. Whether PyPI ever ships the library, or stays pure Python with the fallback.
  5. Whether M5 happens at all, or the facts pass is left to the LLM’s clock.