Accepted Phase 6 source-preserving document knowledge and grounded retrieval.
Supported sources are extractable PDF, DOCX, Markdown, and UTF-8 text. The service retains original bytes, SHA-256 metadata, normalized sections, parser/version metadata, versions, chunks, and model-tagged 768-dimensional Gemini embeddings. PDF citations retain real pages; DOCX citations use actual paragraph/section locations rather than fabricated pages.Re-ingestion creates later document versions. Index rebuild recreates derived chunks/vectors from normalized content. Invalidation excludes a document from current retrieval but retains source and historical version data. Corrupt input is retained with a visible failure state. An owner correction may cite a verified chunk without altering the source.OCR/scanned PDFs, spreadsheets, crawling, advanced multimodal parsing, graph storage, and external parser frameworks remain deferred.