Simplification and robustness plan
Findings from reading main at db351e8 on 2026-10-04, with line numbers from that commit. Three
themes, each one a chain of PR-sized steps that leaves the app working after every step:
| Theme | One line |
|---|---|
| One facade | Every reader and every writer of the corpus goes through GarageService; the app stops parsing psql output and the ingest worker stops silently falling back to direct Postgres. |
| Extraction contract | Extractors are registered by media type (image/* → Tesseract), not by a hand-kept list of suffixes, with one result shape, one error hierarchy and one name per extractor. |
| Schema | Drop what nothing writes, remember every ingest outcome in one place, bound the bookkeeping tables, apply migrations once, and tune the local-first Postgres for what it actually does. |
Two defects found on the way are being fixed separately on main (see §5): extractor_revision()
raises KeyError for .eml/.emlx, and a PlaceholderFile raised inside extract() is counted as
an ingest error instead of a placeholder.
0. Where we start
What exists and shapes the design. Each item is one thing the plan changes.
The facade has leaks.
- The Swift app reads models, sources and corpus stats by running
psql -tAF\tand splitting the output (macapp/Sources/GarageApp/Services/PostgresService.swift:213-241, 744, 778, 1014), beside theListModels,ListSourcesandGetStatsRPCs that return the same data. It also appliesdata/sqlitself (:897-938), beside theInitDbRPC. - The ingest worker’s storage gateway is chosen by environment sniffing
(
ingest/gateway.py:1008-1050): gRPC whenGARAGE_GRPC_SOCKETorGARAGE_GRPC_PORTis inos.environ, else direct Postgres. The app never passes a socket inIngestOptions(AppState.swift:1130-1136), so the choice rests onconfigureHelpershaving run, andmergeConfigurationmirrors options intoos.environonly when Python is already ready (GarageXPCServiceBase.swift:286), with no replay. AconfigureHelpersthat arrives early leaves ingest writing to Postgres directly, with nothing logged. - The database URL reaches the ingest helper by three routes:
updateConfiguration,setDatabaseURL, andIngestOptions.databaseUrlinside the options JSON (IngestService.swift:412-420). - Nine handlers in
service/server.pyhold their own SQLAlchemy queries (Search,ListDocuments,GetDocument,ListSources,ListModels,GetStats,PersistScan,GetEmbeddingBatches,UpdateEmbeddings); the rest are thin overops/.cli.py:541andmcp_server/server.py:390carry their own copies of some of those reads. PersistDocumentmultiplexes seven gateway methods behind a stringaction(server.py:1285-1368), and progress streams are told apart by string phases that Swift compares in a dozen places (GarageGRPCService+Operations.swift:70, 142, 171;AppState.swift:993, 1172, 1200-1203;IngestModels.swift:88-96).- Connection state is a cached enum the app re-syncs only from the Status page’s refresh button
(
GarageGRPCService.swift:358-366). If the helper hosting the server dies,statusstays.running,call()never restarts it, and a ping failure closes the shared channel under other in-flight RPCs (:352). The embed worker has two code paths, in-processbackfill_modelandGetEmbeddingBatches/UpdateEmbeddings, which disagree on batch size, the width check, andON CONFLICT(embed/ollama.py:175vsserver.py:1462); Swift calls onlyembedTexts, so the second path is dead from the app. - The gRPC server runs ten worker threads over a pool of five connections plus five overflow
(
server.py:1516,db/engine.py:41-43).Backfill,EnrichFactsandScaneach hold a worker for their whole run, andGetStatusreportsneeds_migrationon any connection error, sincepending_migrationsreturns every file on an exception (db/migrate.py:217).
Extraction is keyed on suffix tables that drift.
extract/dispatch.pyholds eleven suffix sets and a parallel_EXTRACTOR_MODULEStable that must list every extractor by hand; it missed_email(theKeyError).ingest/chunking.py:37keeps its own suffix map that includes.soland.cob, whichCODE_EXTENSIONSlacks, so those branches are unreachable.ingest/classify.py:32,ingest/scanner.py:375, 448, 533,extract/messages.py:67,extract/nicknames.py:150andextract/image.py:41each hold more.- Extractors name themselves one way and the revision table another:
documents.extractorholdstesseract,pypdf,pypdf+pdfplumber,python-docx,openpyxl, whileextractor_revisionreportsimage:2,pdf:1,docx:1. Nothing joins the two. documents.mimeexists in the schema and the model and is written asNoneon every row (ingest/gateway.py:564).- The error contract is four exceptions in two hierarchies:
NoTextFound,UnsupportedFileandExtractionErrorunder one root,PlaceholderFileunderException; the pipeline’s extract step has fourexceptclauses (ingest/pipeline.py:256-302) and still misses one. - A scanned PDF is flagged
meta.likely_scanned(extract/pdf.py:175-183) and nothing acts on it; there is no way for the PDF extractor to hand a page to the image extractor. Office extractors skip embedded images for the same reason. - The Messages path (
ingest/conversations.py) bypasses classify, attribute, the quality gate and the chunk cap: a second pipeline for one source kind.
The schema carries dead weight and forgets things.
conversationsandmessages(005_conversations.sql) are never read or written; the only reference is a row count inschema_summary. Their design (one other participant, a synthesized document) does not match how threads are stored. The need behind them is real, though: pulling every message one person sent, to read their tone, is a per-message question, and today the per-message facts the readers have in hand (ChatMessage.sent_at,guid, the sender handle,extract/messages.py:79-96; a mail’sDateandFrom) survive only as rendered text inside the chunk.chunks.direction/sender(014) carry the handle but no time and noauthorslink.ingest_statehas four values; onlyokandextract_failedare written.author_identities.kindallowsimessage_handleandhandle, which nothing writes.- A document whose re-extraction fails flips to
extract_failedand keeps its chunks (gateway.py:382), but search filtersd.state = 'ok'(search/hybrid.py:120), so a transient parser error makes a previously good document vanish from results until the next successful run. - A file the quality gate rejects is deleted and its outcome forgotten (
gateway.py:400-421), so it is re-read, re-extracted and re-rejected on every run.ingest_outcomesremembers onlyno_textandfailed. ingest_seengets one row per (run, uri) and is never pruned; neither isingest_runs.apply_migrationsruns every SQL file on every call and consultsschema_migrationsonly to report what is pending; the Swift copy skips applied versions. A migration that is edited after it was applied is silently applied twice or never, depending on who runs it.- Index overlap:
chunks_doc (document_id)is covered bychunks_ord_unique (document_id, ord);documents_class (corpus_class)bydocuments_class_trust (corpus_class, trust_tier);chunks_sha (chunk_sha256)serves no query (chunk reuse compares hashes in Python,gateway.py:619-633). chunks.tsvisto_tsvector('english', text)for every chunk, code included, whiledocuments.langis stored and unused.- Every document commit runs with default
synchronous_commit = onagainst a cluster that only this app uses, and no session setsstatement_timeoutorlock_timeout.
1. One facade: GarageService
Goal: the only process that opens a Postgres connection is the one hosting GarageService (plus the
CLI and tests, which are the same code in-process). Everything else, the app, the ingest helper, the
embed helper, the MCP helper, speaks protobuf to it. Then the server can own the connection pool,
the session settings, the migrations and the invariants, and nothing can bypass them.
1.1 The app reads only through the RPCs
- Replace
PostgresService.listRegisteredModels,listRegisteredSourcesandfetchCorpusStatswithListModels,ListSourcesandGetStats(the Swift client already has them inGarageGRPCService+Operations.swift). Keeppsqlfor what only it can do:pg_isready-style reachability, backup and restore, and the launcher’s “is the cluster up” check. - Replace
PostgresService.applyMigrationswithInitDb. Postgres must be up before the server starts, and the server needs the schema before it answers; the order is already cluster → helper → server →InitDbon the Python side, so the Swift copy is a second implementation with a different skip rule, not a bootstrap necessity.GetStatus.db_status = needs_migrationis the signal the app acts on. - Delete the
psqloutput parsing andRegisteredModel/RegisteredSource/CorpusStatsdecoding once nothing uses them. Test: a unit test overPostgresServiceasserting it no longer has anySELECTtext.
1.2 The ingest worker persists only through the facade
- Make the gateway choice explicit.
ingest_xpctakes astorageargument that is one ofgrpc(socket | host:port, token)ordirect(database_url); the helper passes the gRPC one and the CLI passesdirect. Remove theos.environsniffing fromget_storage_gateway(gateway.py:1008-1050); a worker with no explicit storage fails atBeginIngestSession, loudly, instead of opening Postgres. - One route for configuration.
IngestOptionscarries the gRPC address and token (it already carries the token); dropsetDatabaseURLandIngestOptions.databaseUrl, and stop putting the database URL in the ingest and embed helpers’updateConfigurationat all, since they no longer connect. The “Database Connection” self test in those two helpers becomes a “gRPC Connection” test that makes a realPingwith the token. - Replay configuration.
mergeConfigurationstores the options andpythonDidBecomeReadyapplies whatever arrived before Python was up (GarageXPCServiceBase.swift:280-292). This is a bug fix on its own and can land first. - Fold the per-worker embed path into the one that exists. Swift calls only
embedTexts; backfill runs inside the server. Either deleteembed_via_grpc,GetEmbeddingBatchesandUpdateEmbeddings(reserve the field numbers), or makebackfill_modelthe one implementation and have the RPC path call it with a gateway, as ingest does. Deleting is the simpler choice and the plan assumes it;test_embed_xpc.pygoes with it.
1.3 Server internals: handlers translate, modules decide
- Move the nine inline queries into
ops/(ops/corpus.py:search,list_documents,get_document,list_sources,list_models,stats;ops/sources.py:persist_scan). Each returns a dataclass; the handler, the CLI and the MCP tools all call the same function.cli.py:541andmcp_server/server.py:390stop carrying their own copies. This is the rule the module docstring already states (server.py:1-9); it is just not true yet. - One translation layer,
service/convert.py, withto_proto(dataclass)/from_proto(message)pairs and a round-trip test per message (test_grpc_serialization.pyalready does this for some). Empty-string-means-None lives there, once, instead of in each handler (server.py:1322-1368does it forreplaceand not for the other actions). - Replace
PersistDocumentRequest.actionwith aoneof outcome { Placeholder; ExtractFailed; NoText; Rejected; Seen; RefreshMetadata; Replace }. Each arm carries only its fields, so the server cannot receive atitleon aseenor discard amtimeon aplaceholder(which it does today,server.py:1286). Proto3oneofis wire-compatible to add beside the string field; the string field is reserved after the app ships with the new one. - Typed progress. Give
ScanStatus,BackfillStatus,EnrichFactsStatusand the ingest progress JSON aPhaseenum (STARTED,PROGRESS,DOCUMENT,FINISHED,SKIPPED,CANCELLED,FAILED) and move the ingest progress message intogarage.protoso Swift decodes it with the generated code rather than a hand-mirrored struct (IngestModels.swift:52-85). Swift then switches on the enum in one place per stream instead of comparing strings in six. - Drop
formatted_outputfrom responses. It is presentation, every client formats differently, and it is the only reason some handlers importsnippetand f-string summaries. The CLI builds its text from the dataclass. - Shared enums. Declare
CorpusClass,TrustTier,SourceKindandAuthorRoleas proto enums and use them in every message that carries one.db/models.pykeeps the Python enums but a test asserts they match the proto; Swift dropsCorpusTaxonomy.swift:7-8,SourcesView.swift:38-40andFactPrompts.swift:109.
1.4 Connection lifecycle the app can trust
- One
call()path.search,listDocumentsandgetDocumentget their own copies of the auto-start and a different error mapping (GarageGRPCService.swift:379-381, 432-434, 462-464). Route them throughcall()so cancellation and errors mean the same thing everywhere. - Status from facts, not memory. Replace the cached
statusenum with a watchdog: aPingevery few seconds while any view that needs the backend is open, plus the helper’s XPCisServerRunning. A failed ping marks the server down and triggers one restart attempt with backoff;call()waits on readiness rather than on the enum.start()returning at once while.starting(:186) becomes “await the in-flight start”. - Never close the shared channel under in-flight calls. A failed ping marks the channel suspect and
the next
call()rebuilds it after the current calls finish;cleanupChannelfromdeiniton the main actor goes. - Size the server for its streams.
max_workersgrows to the number of long streams the app can run at once plus a reserve for unary reads (the app runs backfill and enrich-facts on dedicated runners, so at least two long streams), and the engine pool matches.GetStatusdistinguishes “cannot connect” from “needs migration” (pending_migrationsraises, the handler reportsdb_status = unreachable), so the app shows the right remedy. - Session settings in one place.
session_scopetakes arole(read,ingest,maintenance) that setsstatement_timeoutfor reads,synchronous_commit = offfor ingest commits (one document per transaction on a local cluster; durability across a crash is a re-ingest of the last document, which the stat skip already handles), andlock_timeoutfor DDL. See §3.6.
1.5 Helper configuration is one struct
- One
HelperConfigurationvalue (gRPC address, token, models directory, manifest, MCP executable) built in one place (GarageGRPCService.helperConfiguration()already exists,:134-147), sent to every helper byconfigureHelpers, replayed on helper restart and on Python readiness. The database URL is in it only for the helper that hosts the server and the MCP helper. - The in-process (sandboxed) host shares one Python with the ingest delegate; the server’s
os.chdirand ingest’sreset_settings/reset_engineaffect each other (GarageXPCServiceDelegate.swift:56-58,GarageIngestXPCServiceDelegate.swift:274-283). With §1.2 ingest no longer touches the engine, and the server stops callingchdir(it takes the working directory as a setting instead), which removes the shared state.
Order: replay fix (1.2 bullet 3) → explicit gateway (1.2) → app reads over RPC (1.1) → ops extraction and convert layer (1.3) → oneof and enums (1.3) → lifecycle (1.4) → helper config (1.5). Each is one PR; 1.1 and 1.2 are independent of each other.
2. Extraction contract: media types in, one result shape out
Goal: adding a format means registering one extractor under the media types it handles; the
walker, the revision table, documents.mime, the chunker’s language choice and the scanner all read
the same registry. image/* goes to Tesseract because the registry says so, not because eleven
suffixes are listed in the right set.
2.1 Identify the file once
extract/media.py:
@dataclass(frozen=True)
class MediaType:
type: str # "image/png", "application/pdf", "text/x-python", "message/rfc822"
source: str # "suffix" | "name" | "magic" | "uttype"
def identify(path: Path, *, size: int) -> MediaType | None
- Two phases, because the walker must not read.
identify_by_name(path)is suffix and known-name lookup only (free, no I/O), from one table that maps suffix → media type; it is what the walker and the scanner call, and an ambiguous suffix (a.msgis OLE, a.docmay be RTF, an.xmlmay be an Atom feed) returns a provisional type that counts as indexable.identify(path)runs in the extract step, after pruning, the size check and the budgetedensure_local, and may read the first 4 KiB once to settle an ambiguous case; a cloud stub is therefore never downloaded by identification, and a dataless read is never attempted on the walker’s thread (walker.py:246-268,materialize.py). On macOS,UTType(filenameExtension:)and, in the second phase,UTType(contentsOf:)are reachable from Python through the same ctypes pattern asextract/imageio.py; off macOS the table andmimetypesstand in. is_indexable(path)becomesidentify_by_name(path) is not None and registry.handles(it). The walker keeps its “no I/O before pruning” property by construction.documents.mimegets the identified type. That column finally means something: it drives the Documents page’s icon,rag_get_document’s hint to the client, and per-type stats.
2.2 One registry, one protocol
class Extractor(Protocol):
name: str # the one name: "image", "pdf", "docx", "email", "code", ...
version: str # bumps retry remembered outcomes
media_types: tuple[str, ...] # "image/*", "application/pdf", "text/markdown", ...
def extract(self, path: Path, media: MediaType, ctx: ExtractContext) -> ExtractResult: ...
REGISTRY: list[Extractor] # first match on media type wins; "text/*" is the fallback
def extractor_for(media: MediaType) -> Extractor
def extractor_revision(media: MediaType) -> str # f"{e.name}:{e.version}", derived, cannot be missing
ExtractResult.extractorisExtractor.name; the engine that did the work (pypdf,pypdf+pdfplumber,tesseract) moves tometa.engine.documents.extractorthen joinsextractor_revision, and_indexed_by_a_retired_extractor(pipeline.py:154) becomes a comparison ofdocuments.extractor || ':' || extractor_versionagainst the registry instead of a per-suffix special case.ExtractResultgains structured side outputs besidetext:comments(author, time, body, anchor, parent) andattachments(name, media type, bytes hash), which §3.9 persists asmessagesrows anddocument_shares. An extractor that has none leaves them empty; the pipeline never parses them out of the text.ExtractContextcarries settings (max_file_bytes,ocr_min_chars) and adelegate(media_type, path_or_bytes)callback, so an extractor can hand a sub-document to another: the PDF extractor rasterizes a page it found empty and delegates it asimage/png; the DOCX extractor delegates embedded images. Thelikely_scannedflag becomes a behavior. Delegation is bounded by the same budget settings (pages per PDF, images per document) so a photo album in a.pptxdoes not become an OCR run.- One error hierarchy under
ExtractionOutcome:NoText,Unsupported,Placeholder,Failed. The pipeline’s extract step becomes oneexcept ExtractionOutcome as outcomethat maps to the gateway’s oneof arm (§1.3).PlaceholderFilemoves under it (the “deliberately distinct” note inplaceholder.py:66-70is about the remedy, which the subclass still expresses). - Lazy imports stay: each registry entry names a module and the attribute to load, as
_EXTRACTOR_MODULESdoes today, so a walk over a code tree still never importspdfplumber.
2.3 Every other suffix table reads the registry
ingest/chunking.py:37(_CODE_LANGUAGES) becomes a map from media type (text/x-python) to splitter language, and its two unreachable entries either get media types or go.ingest/classify.py:32(_DOC_EXTENSIONS_IN_CODE_TREES) asks “is this media type prose” of the registry entry (ContentKindlives on the extractor, so the question is answerable without extracting).ingest/scanner.pycounts by the sameidentify, so a scan’sexpected_elementsand an ingest’sseenagree by construction (today maildir counts.msg/.mboxthe walker never indexes,:448).extract/messages.py:67,nicknames.py:150andscanner.py:375share oneSQLITE_SUFFIXES.- Swift has no suffix tables to delete; it never decides what is indexable.
2.4 Sources produce documents through one pipeline
The Messages path is a second pipeline because its input is not a file. Make the unit of work a
DocumentCandidate (uri, media type, a way to get bytes or text, stat) produced by a Producer
per source kind: the walker yields file candidates; the sqlite producer yields one candidate per
thread with media_type = "application/x-garage-thread" and the rendered text already in hand.
ingest_one then runs the same steps for both: stat skip, extract (a no-op extractor for
pre-rendered text), quality gate, chunk, classify, attribute, persist. The thread-specific parts
(one chunk per message, direction/sender) become the CONVERSATION chunker, which is where they
belong. Then communications get the chunk cap and attribution evidence like everything else, and
_ingest_source’s if kind == "sqlite" branch (pipeline.py:546-566) goes.
Order: identify + documents.mime (2.1) → registry and derived revision (2.2, replaces the
hand-kept table and lands the one-name rule; a migration rewrites documents.extractor from the
old engine names) → error hierarchy (2.2) → the other tables (2.3) → delegation for scanned PDFs
(2.2, with a budget setting and the schema regen) → producers (2.4). The two bug fixes in §5 land
first and the registry step deletes the table they patch.
3. Schema: resilient, bounded, tuned
Goal: the schema says what the code does, every ingest outcome has one home, bookkeeping cannot grow
without bound, migrations apply once, and the server is configured for a single-user local corpus.
Migrations keep the repo’s rules: idempotent, re-applicable, data/sql is truth and
db/models.py mirrors it, and docs/schema.md is updated with each.
3.1 Drop what nothing writes (015_drop_unused.sql)
DROP TABLE messages, conversations(in that order:messages.conversation_idreferencesconversations,005_conversations.sql:48, so the parent cannot go first) and both models. A thread is the document that holds it, and nothing reads or writes either table, so both are empty on every install. §3.8 creates a newmessageswith a different shape; dropping the old one here keeps that migration a plainCREATE TABLE.ingest_state: removeembed_partialandplaceholder. Postgres cannot drop enum values, so recreate the type: addingest_state_v2 ('ok', 'extract_failed'), alter the column with aUSING, drop the old type, rename. Or, with §3.3, dropdocuments.stateentirely.author_identities_kind_check: keep only the kinds code writes, plus whateverself_identitiesconfig may list (a test enumerates both).- Drop
chunks_doc,documents_classandchunks_sha(covered or unused, §0). documents.mimestays and is populated (§2.1).
3.2 Migrations apply once (db/migrate.py)
apply_migrationsskips files whose stem is inschema_migrations, as the Swift copy already does; idempotency remains the safety net, not the mechanism.- Add
checksumtoschema_migrations. An applied file whose checksum changed is an error that names the file: edit a migration, add a new one. A test hashesdata/sqland fails when a committed migration’s checksum moves after it is tagged (keep adata/sql/CHECKSUMSfile the test regenerates on demand). - Each file runs in its own transaction (
BEGIN … COMMITaround the file, withlock_timeoutset so a migration that needs an exclusive lock fails fast against a running backfill instead of queuing behind it and blocking every reader). The extension file stays outside SQLAlchemy as today. - Replace the f-string
INSERT(migrate.py:160-162) with a parameter.pending_migrationsraises on a connection error instead of returning every file (§1.4 uses the distinction). - With §1.1 the Python applier is the only one; the Swift copy is deleted.
3.3 One home for ingest outcomes (016_ingest_outcomes.sql)
Today a file’s last outcome lives in two places with different rules: documents.state/error
for a document that exists, ingest_outcomes for one that does not, and nowhere for rejected.
ingest_outcomesbecomes the record of the last attempt for every (source, uri), whatever the result:outcome IN ('indexed', 'unchanged', 'no_text', 'rejected', 'failed', 'placeholder', 'unsupported'), withbyte_size,mtime,source_sha256,extractor_revision,error,run_id,recorded_at.rejectedis remembered, so the quality gate runs once per file version instead of once per run.documents.stateanddocuments.errorgo. A document row exists only for indexed content, and search’sd.state = 'ok'filter goes with it. A failed re-extraction recordsfailediningest_outcomesand leaves the document and its chunks searchable; the Documents page shows the outcome beside the document by joining on (source_id, uri). A document whose file is gone is still reconcile’s job, not an outcome.check_statreads one row from each table and_is_settledbecomes “the outcome row’s stat matches and its revision is current”, for every outcome butplaceholder. A placeholder is a file whose bytes were not fetched, often because that run’s materialization budget was spent; its stat does not change when the budget is renewed, so it is never settled and is offered toensure_localon every run (cheap: no read happens unless the budget allows one). The pipeline’s step 1 and step 3 skip rules collapse to one. Test: a run with an exhausted budget followed by a run with a renewed one indexes the file.
3.4 Bound the bookkeeping (017_last_seen.sql)
- Replace
ingest_seenwithingest_outcomes.last_seen_run_id: “seen in run R” is “the row’s last-seen run is R or later”, which §3.3 writes anyway. Reconcile deletes documents whose row’s last-seen run is older than the source’s last completed run, under the same completed-run guard as today (ingest/reconcile.py). One fewer table, no per-run row fan-out, and the walk’srecord_seenbecomes an update on a row that exists. - Runs can overlap:
begin_sessiononly inserts aningest_runsrow (gateway.py:220-240), so a scheduled maintenance ingest and a manual one can walk the same source at once. Every outcome write therefore setslast_seen_run_id = GREATEST(last_seen_run_id, :run_id), never a plain assignment, so an older run finishing late cannot move a marker backwards and make reconcile delete a document both runs saw. Test: two interleaved runs over the fixture corpus, the older one recording last, leave every document with the newer run’s id and reconcile deletes nothing. This depends on §3.3 having a row for every indexed URI (today success deletes the outcome row), so it lands after it. - Keep the last N
ingest_runsper source (N = 20, a setting) and delete older ones infinalize_session.fact_runsis per (document, prompt) and already bounded. scan_detailsstays onsources; it is one row per source.
3.5 Search-side columns
chunks.tsvuses thesimpleconfiguration forcodechunks andenglishfor the rest:to_tsvector(CASE WHEN chunker LIKE 'code%' THEN 'simple' ELSE 'english' END::regconfig, text)is immutable enough for a generated column when written with the literal regconfig casts. Stemming identifiers (parser→pars) hurts code recall and helps nothing. The query side changes in the same step, sincetsquery_exprbuilds every query withenglish(search/hybrid.py:115-117) and anenglishquery forrunwould no longer match asimplevector holdingrunning. The FTS CTE binds both queries and matches each chunk with its own configuration:(c.chunker LIKE 'code%' AND c.tsv @@ :q_simple) OR (c.chunker NOT LIKE 'code%' AND c.tsv @@ :q_english), ranked byts_rank_cdagainst the matching query. Test intest_postgres.py: a prose chunk withrunningand a code chunk withrunningare both found byrunin prose andrunningin code, andrundoes not match the code chunk.documents.langis either used to pick the configuration or dropped; the plan drops it until an extractor sets it.facts.tsvserves onlyListFacts; keep it, it is small.- The
hnsw_bqfirst stage (hybrid.py:229-259) andindex_kind = 'none'are reachable only for models wider than 4000 dims without MRL;models.jsonhas none. Leave the code, add atest_postgres.pycase so it stays working, since it is the only path that is never exercised by the app.
3.6 Postgres session and cluster settings
Set by the server (§1.4) per session role, and by the app in postgresql.conf for the bundled
cluster:
| Setting | Where | Why |
|---|---|---|
synchronous_commit = off |
ingest and backfill sessions | one commit per document or batch; losing the last one on a crash is a re-ingest the stat skip handles |
statement_timeout = 30s |
read sessions (search, lists) | a runaway query cannot wedge the UI; the app’s 30 s client deadline becomes a server-side fact |
lock_timeout = 5s |
migrations, DROP TABLE emb_* |
fail fast behind a backfill instead of queuing and blocking every reader |
hnsw.ef_search |
search sessions (exists) | unchanged |
maintenance_work_mem = 512MB |
backfill session before CREATE INDEX |
HNSW builds are memory-bound; the index on a new model table is built after the first backfill, not before (create_embedding_table builds it empty today, so every insert is an index insert) |
jit = off |
cluster | short OLTP queries lose to JIT warm-up |
shared_buffers, effective_cache_size |
cluster, from physical RAM | the bundled cluster ships Postgres defaults sized for a shared host |
The index-after-backfill change is the one with a visible payoff: build the per-model table without
its HNSW index, backfill, then CREATE INDEX once (emb_tables.create_embedding_table,
registry.index_ddl). embedding_models gets index_built boolean so search falls back to exact
KNN (SET LOCAL enable_indexscan, or no ORDER BY operator hint needed; pgvector does exact scans
without an index) until it is built, and GetStats reports it.
3.7 Constraints that catch bugs
chunks:CHECK (char_end IS NULL OR char_end >= char_start),CHECK ((direction IS NULL) = (sender IS NULL)).facts: same offsets check;CHECK (char_start IS NOT NULL)once the extractor drops ungrounded facts, as the comment in006_facts.sqlsays it does.documents:CHECK (content IS NOT NULL)once §3.3 makes a document row mean indexed content;CHECK (extractor <> '').ingest_outcomes:CHECK (outcome <> 'failed' OR error IS NOT NULL).embedding_models:CHECK (index_kind <> 'hnsw_bq' OR storage_kind = 'vector').
3.8 Messages as rows, keyed to their chunks (018_messages.sql)
The question “what did this person write, and how” needs each message as a row with its author,
its time and its chunks. A message is the unit of authorship; a chunk is the unit of embedding;
one message has one or many chunks (a text has one, a long mail has several), so the link runs
from chunk to message, as chunks.fact_id (007) runs from chunk to fact. Threads are a relation
between messages, since mail threads are trees and span files, while a Messages thread is one file.
CREATE TABLE messages (
id bigserial PRIMARY KEY,
document_id bigint NOT NULL REFERENCES documents(id) ON DELETE CASCADE, -- the thread file, or the .eml
author_id bigint REFERENCES authors(id) ON DELETE SET NULL, -- resolved sender
direction text NOT NULL CHECK (direction IN ('sent', 'received')),
sender text, -- the raw handle or address
sent_at timestamptz, -- NULL when the source has no usable time (a mail with no Date)
external_id text, -- chat.db guid, mail Message-ID
in_reply_to text, -- mail: the parent's Message-ID, as written
thread_key text NOT NULL, -- chat: the chat guid; mail: the root Message-ID, else 'doc:<document_id>'
subject text,
meta jsonb NOT NULL DEFAULT '{}'::jsonb,
CONSTRAINT messages_external_unique UNIQUE (document_id, external_id)
);
CREATE INDEX messages_author_sent ON messages (author_id, sent_at DESC NULLS LAST);
CREATE INDEX messages_thread ON messages (thread_key, sent_at);
CREATE INDEX messages_external ON messages (external_id);
ALTER TABLE chunks ADD COLUMN message_id bigint REFERENCES messages(id) ON DELETE CASCADE;
CREATE INDEX chunks_message ON chunks (message_id) WHERE message_id IS NOT NULL;
- Messages threads. One
documentsrow per thread as today; onemessagesrow per message withthread_key = chat guid; each message chunk carriesmessage_id. A message long enough to split still has one row and several chunks. - Mail. One
documentsrow per.eml(as today) and onemessagesrow for it, withexternal_id = Message-ID,in_reply_tofrom the header, andthread_keyfrom the firstReferencesentry, else its own Message-ID, elsedoc:<document_id>so a message with no identifiers is its own thread rather than grouped with every other such message. A mail with no parseableDate(which the extractor already accepts,test_mail_extract.py:122-133) getssent_at = NULL, never an invented time; per-author queries order withNULLS LAST. The body’s chunks all carry itsmessage_id. The thread is thenWHERE thread_key = ?ordered bysent_at, and the tree is a self-join ofin_reply_toonexternal_id, resolved at query time so ingest order does not matter and a parent that was never ingested is just a missing row. A subject-based fallback for mailers that dropReferencesis ametanote, not a column, until it is needed. - Quoted text. A reply carries the earlier messages quoted below it; those words are not the
sender’s and must not count toward their tone. The mail extractor (§2.2) splits the new text from
the quoted tail (
>prefixes,On … wrote:lines, Outlook’sFrom:block). The new text’s chunks getmessage_id; the quoted chunks stay searchable but carry nomessage_id,directionorsender, so per-author queries never see them andrag_search’s direction filter keeps to words people actually wrote.documents.contentkeeps the whole mail forrag_get_document. - No duplicate uniqueness on
Message-ID. The same mail is often in two mailboxes (Sent and a folder, or two accounts), which is two documents;(document_id, external_id)is unique,external_idalone is only indexed. author_idis the resolved sender, through the sameauthor_identitieslookup attribution uses fordocument_authors(phone/email→ author, create on miss,is_selffor the owner). That is what makes “every message from X” one index scan whatever handle X used in which app, and what lets tone be studied per person rather than per address. Recipients stay ondocument_authors(sender/recipient/ccroles) as today; a per-message recipient table is not needed while a mail is a document.chunks.directionandchunks.senderstay for the search filter (both engines filter on them inWHERE, which is why014made them columns). The migration back-fills what is stored: for chats, one row per message chunk fromdirection/senderand the rendered timestamp line; for mail, one row per document frommeta.message_idandmeta.email_date(extract/mail.py:159-168keeps those two and nothing else:ReferencesandIn-Reply-Toare not inmetaor in the rendered header, soin_reply_toandthread_keycannot be recovered from the database). The back-filled mail rows getthread_key = Message-IDandin_reply_to = NULL; the mail extractor’sVERSIONis bumped in the same step so every.emlstill on disk re-extracts on the next ingest with the headers, the quoted-text split and the real thread key, and a mail whose file is gone keeps its back-filled row. Thenchunks.message_idis set.- The queries this enables, each a join with no text parsing: messages by author in a date range,
with text (
chunks.text) and vectors (emb_* ON chunk_id) for clustering by tone; a mail thread in order, across files; a per-person summary inrag_list_authors; the v1.5 memory and triples work gets a per-message anchor instead of a document-level one. - Order within §3:
messageslands after the outcomes change (§3.3) so it is written by the one pipeline (§2.4) and not by thesqliteside branch; the quoted-text split in the mail extractor lands with it, since the rows are wrong without it.
3.9 Versions, comments and shares (019_versions.sql, 020_shares.sql)
Three more things a document is besides its current text. Each is a side table keyed to
documents, messages or authors, with a chunk link only where its text should be searchable,
so they add to §3.3 and §3.8 without reshaping them. Each is built when a source produces it; the
tables are designed now so the ones that land first do not have to move.
Versions. Today replace_document overwrites: the previous text is gone, ingested_at is the
only history, and chunk reuse by hash (gateway.py:619-633) is the one version-aware thing, since
an unchanged chunk keeps its id and vectors.
CREATE TABLE document_versions (
id bigserial PRIMARY KEY,
document_id bigint NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
version_no int NOT NULL, -- 1.. per document, head is max
content_sha256 bytea NOT NULL,
source_sha256 bytea,
byte_size bigint,
mtime timestamptz,
content text NOT NULL, -- TOAST-compressed by Postgres
extractor text NOT NULL,
extractor_version text NOT NULL,
author_id bigint REFERENCES authors(id) ON DELETE SET NULL, -- who made this version, when known
external_id text, -- provider revision id, git commit, ...
observed_at timestamptz NOT NULL DEFAULT now(),
meta jsonb NOT NULL DEFAULT '{}'::jsonb,
CONSTRAINT document_versions_no_unique UNIQUE (document_id, version_no)
);
CREATE INDEX document_versions_sha ON document_versions (document_id, content_sha256);
-
Uniqueness is on
(document_id, version_no)only. A revert (A → B → A) is a real third version whose hash repeats the first; a unique on the hash would reject the insert and, insidereplace_document’s transaction, roll back the head update with it. The hash index serves “has this document ever held this text”. Test: the A → B → A sequence yields three versions. documentsstays the head: its columns are the current version, andchunks,factsandmessageshang off the head only. Search and backfill never see old versions, so embedding cost is unchanged.rag_get_document(version=)and a Documents-page history readdocument_versions; a diff between two versions is computed, not stored.- A version is written when
replace_documentsees a newcontent_sha256for an existing document (version_no + 1, the old head copied first on the migration that adds the table). That is the only version signal today: ingest observed a change. Others plug in throughexternal_idandauthor_idwhen a source has them: a cloud provider’s revision list, git for a tracked file (the commit, with the committer asauthor_id, which §0’s attribution already resolves), Office tracked changes (a.docxwith revision marks yields the accepted text as head and the authors of the marks inmeta), an edited iMessage (chat.dbkeeps edit history; it goes inmessages.meta.edits, not here, since a message is not a document). - Retention is a setting (
ingest.keep_versions, default 10 per document, 0 keeps none), applied inreplace_document; adocument_versionsrow is never written forcodedocuments by default, since git is the better history and the corpus-wide churn would dwarf everything else.
Comments. A comment is a message about a document, anchored to part of it: it has an author,
a time, a body, a parent and a resolution. Rather than a third table of authored text, it is a
messages row with kind = 'comment' and an anchor, so “everything X wrote” stays one query
across texts, mail and comments, and the chunk link gives it search and vectors for free.
ALTER TABLE messages ADD COLUMN kind text NOT NULL DEFAULT 'message'
CHECK (kind IN ('message', 'comment', 'reaction'));
ALTER TABLE messages ADD COLUMN anchor jsonb; -- {"page": 3}, {"paragraph": 12, "quote": "..."}, {"cell": "B7"}, {"char_start":..,"char_end":..}
ALTER TABLE messages ADD COLUMN resolved boolean; -- comments only; NULL for messages
-- in_reply_to / thread_key already give replies and the comment thread; document_id is the commented document.
- Local sources that carry them, and which extractor reports them (an
ExtractResult.commentslist, §2.2, persisted byreplace_documentbeside the chunks):.docxcomments and tracked changes (word/comments.xml,w:ins/w:delwithw:author/w:date),.pptxcomments,.xlsxcell notes (author in the note), PDF annotations (/Annotswith/T,/Contents,/M), Messages tapbacks (associated_message_type,kind = 'reaction', one row, no chunk), and mail, where a reply already is the comment (in_reply_to). - Anchoring is
jsonbbecause every format anchors differently and nothing filters on it; when the anchor can be mapped to the head’s text,char_start/char_endare added so a comment can be shown beside the chunk it is about, as facts are. - A comment on a version that is no longer the head keeps
anchor.version_noso the Documents page can show it against the right text. - A comment is a communication whatever its document is. Every off-box check keys on the
parent document’s
corpus_classtoday: backfill’s pending query (embed/ollama.py:103-111),rag_ask’s result check (mcp_server/agent.py:456-466), fact extraction. A comment on adocumentorcodedocument is someone’s message to the owner, and under those rules its chunk would be eligible for an off-box provider. The content rule therefore becomes “a chunk is a communication when its document is, or when it carries amessage_id”, applied in one place (egress.allows_communicationsgains the chunk predicate, andpending_chunks_sql, the search filter the agent uses, andenrich/facts.pytake it from there rather than each writingd.corpus_class = 'communication').test_embed_egress.pyandtest_egress_block.pyget a mixed case: adocument-class fixture with two comments, an off-box model, and the assertion that the body chunks are posted and the comment chunks withheld, for backfill,rag_askand facts. This step does not land without those tests.
Shares. Who a document went to, or came from, and by what channel. This is the relation that
turns a received trust tier from a path rule into evidence, and answers “what have I sent X” and
“what did X send me”.
CREATE TABLE document_shares (
id bigserial PRIMARY KEY,
document_id bigint NOT NULL REFERENCES documents(id) ON DELETE CASCADE,
direction text NOT NULL CHECK (direction IN ('sent', 'received', 'shared')), -- shared: a live shared folder or link
author_id bigint REFERENCES authors(id) ON DELETE SET NULL, -- the counterparty
channel text NOT NULL CHECK (channel IN ('mail', 'messages', 'airdrop', 'download', 'dropbox', 'icloud', 'drive', 'link')),
message_id bigint REFERENCES messages(id) ON DELETE SET NULL, -- the mail or text that carried it
external_id text, -- share link id, provider share id
shared_at timestamptz,
meta jsonb NOT NULL DEFAULT '{}'::jsonb,
-- One row per share event. NULLS NOT DISTINCT (Postgres 15+; the bundled server is 18, CI runs
-- pg18) so a download or AirDrop with no provider id and no resolved person still collapses to
-- one row on every rescan, while two mails carrying the same file stay two rows (message_id).
CONSTRAINT document_shares_unique
UNIQUE NULLS NOT DISTINCT (document_id, channel, message_id, external_id, author_id)
);
CREATE INDEX document_shares_author ON document_shares (author_id, shared_at DESC);
CREATE INDEX document_shares_message ON document_shares (message_id);
- Local signals, in the order they are cheap to get: a mail attachment (the carrying message is
the
messagesrow; the attachment’s bytes are hashed and matched to an existing document bysource_sha256, or ingested asmail:<Message-ID>/<part>when nothing else holds them); a Messages attachment (chat.db’sattachmenttable and~/Library/Messages/Attachments, same matching); thecom.apple.metadata:kMDItemWhereFromsxattr, which records the URL a download came from and, for AirDrop and Mail saves, the sender (channel = 'download' | 'airdrop', the counterparty resolved throughauthor_identitieswhen it is an address); a Dropbox or iCloud shared folder or a Drive file with collaborators, where the client exposes it locally (direction = 'shared', re-observed on each scan so a revoked share goes). - Attribution reads it: a document whose only share row is
receivedfrom X with no other evidence getstrust_tier = receivedand X as asenderondocument_authors, withevidence = 'share:mail:<Message-ID>', which is the diagnosable formdocs/attribution.mdwants. - Shares are relations between people and documents, and threads are relations between messages; both are what Apache AGE in the bundled Postgres is for. These tables stay the relational truth, and a graph view is projected from them when a query needs a traversal, never the other way round.
Order within §3: versions can land any time after §3.3 (it only touches replace_document);
comments need §3.8 and the extractor side-outputs of §2.2; shares need §3.8 for message_id and
the attachment matching in the mail and Messages readers.
Order: migrations-apply-once (3.2, no schema change, unblocks safe edits) → drop unused (3.1)
→ outcomes (3.3, with the gateway oneof from §1.3 so the wire change and the table change are one
PR) → last-seen (3.4) → tsv config and lang (3.5) → session roles and index-after-backfill (3.6)
→ messages (3.8) → versions, comments, shares (3.9) → constraints (3.7, last, once the code
guarantees them).
4. Cross-cutting: one source for each truth
| Truth | Today | After |
|---|---|---|
| Corpus class, trust tier, source kind, author role | SQL enums, Python enums, three Swift string lists | proto enums (§1.3); SQL and Python checked by test; Swift uses the generated enums |
| Progress phases | strings compared in six Swift sites and three Python ops | proto Phase enum (§1.3) |
| What is indexable and how | eleven suffix sets plus six satellites | extract/media.py + the registry (§2) |
| Extractor name | documents.extractor vs _EXTRACTOR_MODULES |
Extractor.name (§2.2) |
| Model catalog | models.json, GarageConfigLoader with three decode shapes and name heuristics |
models.json read by Python; Swift gets ListModels plus a ListCatalog RPC that returns the catalog with is_embedding, dims, provider resolved (closes the catalog item of #27) |
| Config defaults | config/__init__.py and GarageConfigLoader.swift:439-487 |
Swift reads settings through GetSetting; the defaults exist once, in Python |
| Schema version | Python applies everything each time; Swift skips applied | Python applies once with checksums (§3.2); Swift calls InitDb |
| Database URL for helpers | three routes to ingest | one HelperConfiguration (§1.5), and only two helpers get the URL at all |
5. Already underway
Two defects confirmed while reading, fixed on main in their own session and PR, independent of
this plan:
extract/dispatch.py:267_EXTRACTOR_MODULEShas no_emailentry;extractor_revision()raisesKeyErrorfor.eml/.emlx, so a mail file that fails extraction or has no text crashes the per-file handler instead of being remembered. Fix: add the entry and a test that every extractorextractor_forcan return is in the table. §2.2 later derives the revision from the registry, which makes the table and the test unnecessary.PlaceholderFileis a plainExceptionandextract()raises it fromcheck_materialized;ingest_one’s extract step does not catch it, so the file is counted as an error and recorded asseeninstead of as a placeholder. Fix: catch it in step 3 as step 2 does, with a pipeline test. §2.2’s error hierarchy makes it oneexcept.
6. Sequencing
Each row is one PR, with the tests that guard it. Rows in the same group are independent of each other.
| # | Change | Guard |
|---|---|---|
| 1 | §5 defects | new unit tests |
| 2 | Replay helper configuration when Python becomes ready (§1.2) | GarageXPCServiceBase unit test with a stubbed runtime |
| 3 | Migrations apply once, with checksums and per-file transactions (§3.2) | test_migrate.py, test_postgres.py |
| 4 | Explicit storage choice for the ingest worker, delete the env sniffing (§1.2) | test_ingest_gateway.py, test_ingest_xpc.py |
| 5 | App reads models, sources, stats and runs InitDb over gRPC; delete the psql reads and the Swift migrator (§1.1) |
Swift unit tests; test_grpc_operations.py |
| 6 | Move inline handler queries into ops/; service/convert.py with round-trip tests (§1.3) |
test_grpc_documents.py, test_grpc_serialization.py, test_cli_commands.py |
| 7 | identify() + documents.mime populated (§2.1) |
test_fixture_corpus.py asserts a mime per document |
| 8 | Extractor registry, derived revision, one name, migration rewriting old names (§2.2) | test_image_extract.py, test_mail_extract.py, test_postgres.py for the rename |
| 9 | One error hierarchy; one except in the pipeline (§2.2) |
pipeline tests over a mocked gateway |
| 10 | Drop the old messages and conversations, unused enum values and indexes (§3.1) |
test_postgres.py migration re-apply |
| 11 | PersistDocument oneof + ingest_outcomes as the one outcome record, documents.state dropped (§1.3, §3.3) |
test_ingest_gateway.py, test_postgres.py, test_grpc_serialization.py |
| 12 | ingest_seen → run_id on outcomes; prune ingest_runs (§3.4) |
reconcile tests, test_postgres.py |
| 13 | Proto enums and Phase; ingest progress in proto; Swift switches on them (§1.3) |
enum-parity test; Swift presentation tests |
| 14 | Connection watchdog, one call() path, server sized for streams, unreachable vs needs_migration (§1.4) |
test_grpc_server.py; Swift unit tests for the watchdog state machine |
| 15 | Session roles and settings; index after first backfill (§3.6) | test_postgres.py builds a model table, backfills, builds the index, searches |
| 16 | Remaining suffix tables read the registry; scanner and walker agree (§2.3) | test_scanner.py counts == walker counts over the fixture corpus |
| 17 | Delegation: scanned PDF pages and embedded images to image/* (§2.2) |
fixture PDF with one scanned page; budget setting documented and in the schema |
| 18 | tsv configuration per chunk kind; drop documents.lang (§3.5) |
test_postgres.py keyword search over an identifier |
| 19 | Producers: Messages through the one pipeline (§2.4) | test_messages.py, test_fake_messages.py |
| 20 | messages rebuilt with chunks.message_id, author, time and thread key; mail quoted-text split; back-fill; rag_list_authors per-person counts (§3.8) |
test_postgres.py back-fill over fixture threads and mail; test_messages.py one row per message with the resolved author; test_mail_extract.py a reply’s quoted chunks carry no message_id and a three-mail thread orders by thread_key |
| 21 | document_versions written by replace_document, retention setting, rag_get_document(version=) (§3.9) |
test_ingest_gateway.py two replaces give two versions and head unchanged chunks keep ids; test_postgres.py migration copies the head |
| 22 | Comments as messages rows with kind/anchor; .docx/.pdf/.xlsx extractors report them; tapbacks as reactions (§3.9, §2.2) |
fixture docx with two comments and a tracked change; test_messages.py a tapback is one row and no chunk |
| 23 | document_shares from mail and Messages attachments and kMDItemWhereFroms; attribution evidence share:… (§3.9) |
test_mail_extract.py an attachment matches a fixture document by hash; test_attribution.py a received share sets received with evidence |
| 24 | Constraints (§3.7); ListCatalog and Swift config through GetSetting (§4) |
test_postgres.py; Swift tests |
| 25 | HelperConfiguration as one struct; server stops chdir (§1.5) |
Swift tests; test_grpc_server.py |
Groups: {1, 2, 3} → {4, 5, 6, 7} → {8, 9, 10} → {11, 13} → {12, 14, 15, 16} → {17, 18, 19} →
{20, 21} → {22, 23} → {24, 25}. Step 12 follows 11 because it updates the outcome row that 11
starts keeping for indexed files; today success deletes that row, so landing 12 first would make
record_seen update nothing. Nothing here changes the egress guard; every step that touches embed/, enrich/ or
net/ keeps test_egress_block.py and test_embed_egress.py green, and the deletion of the
worker embed path (§1.2) removes one caller from CALLERS rather than adding one.