How it works
How documents become searchable
Between “file in a drive” and “cited in an answer” every document passes the same pipeline, fully automatic:
- Sync (sweep): each source is compared against the knowledge base — new, changed and removed items are detected, and only what actually changed is re-processed.
- Conversion: Word, PDF, Excel, PowerPoint and the rest are converted to clean text; scanned PDFs are OCR'd automatically.
- Preparation: each document is readied for citation along its own structure — headings, clauses, tables — so an answer can point at the exact passage, not just the file.
- Indexing: everything becomes searchable, whether a question comes at it by idea or by an exact term like an invoice number.
- Understanding: the engine works out each document's date, its kind (contract, invoice, report…), the people and companies in it, and where it stands as evidence when documents conflict. This is what powers timeline and “everything about X” questions.
Reading the Documents table
| Status | active = searchable · failed = couldn't be read (the red reason says why) · deleted = removed, kept as a tombstone so syncs don't resurrect it. |
|---|---|
| Passages | How many searchable passages the document produced. 0 usually means empty or unreadable content. |
| Access | Who may read it — groups, named people, or “everyone” (all signed-in users). |
A source that can't be reached (expired credentials, an outage) shows an error and its documents stay in place — a temporary problem never empties the knowledge base. Failures of single documents are listed on the source and under Needs attention.
Was this page helpful?
Last updated 29 Aug 2026