{
  "id": "glossary/ingestion/documents/extracted-text",
  "mentions": [
    {
      "context": "…ptioned by the first line of its extracted text, and leading to that text; once …",
      "file": {
        "href": "../../../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 5,
      "surface": "extracted text",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… its words are recorded from its extracted text, but it enters the model only th…",
      "file": {
        "href": "../document/index.html",
        "label": "ingestion/documents/document.md"
      },
      "href": "../document/index.html#L6",
      "kind": "written",
      "line": 6,
      "passages": 3,
      "surface": "extracted text",
      "title": "Document",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…low behind the tabs \"Document\", \"Extracted text\" and \"Related notes\", with \"Down…",
      "file": {
        "href": "../../../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "Extracted text",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…oned by their first line and the extracted text of every slide; a mention read i…",
      "file": {
        "href": "../deck/index.html",
        "label": "ingestion/documents/deck.md"
      },
      "href": "../deck/index.html#L7",
      "kind": "written",
      "line": 7,
      "passages": 2,
      "surface": "extracted text",
      "title": "Deck",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…ck identifier after it; then the extracted text position by position under \"Extr…",
      "file": {
        "href": "../../../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 2,
      "surface": "extracted text",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…its pages, slides or cues, their extracted text in disclosure blocks, and the vi…",
      "file": {
        "href": "../../../../specs/screens/pages/entity-page/index.html",
        "label": "screens/pages/entity-page.md"
      },
      "href": "../../../../specs/screens/pages/entity-page/index.html#L39",
      "kind": "recognised",
      "line": 39,
      "passages": 2,
      "surface": "extracted text",
      "title": "Entity page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…hown is marked. Under the rail, \"Extracted text\" gives the text of every positio…",
      "file": {
        "href": "../../../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 5,
      "surface": "Extracted text",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…s on slide three. A document without extracted text says so; the download stays.",
      "file": {
        "href": "../../../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 5,
      "surface": "extracted text",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… the document\" that searches the extracted text of the positions, jumps to the f…",
      "file": {
        "href": "../../../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../../../specs/viewers/document-viewer/index.html#L14",
      "kind": "recognised",
      "line": 14,
      "passages": 5,
      "surface": "extracted text",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…pen a position of the rail → its extracted text, and its page once the viewer ru…",
      "file": {
        "href": "../../../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../../../specs/viewers/document-viewer/index.html#L27",
      "kind": "recognised",
      "line": 27,
      "passages": 5,
      "surface": "extracted text",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… its words are recorded from its extracted text, but it enters the model only th…",
      "file": {
        "href": "../document/index.html",
        "label": "ingestion/documents/document.md"
      },
      "href": "../document/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 3,
      "surface": "extracted text",
      "title": "Document",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…the file for download, shows its extracted text page by page and, when the conve…",
      "file": {
        "href": "../document/index.html",
        "label": "ingestion/documents/document.md"
      },
      "href": "../document/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 3,
      "surface": "extracted text",
      "title": "Document",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…zoom and its find field over the extracted text, and the strip follows the page …",
      "file": {
        "href": "../../../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "extracted text",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…riginal stays downloadable\". The extracted text view gives the text of every pag…",
      "file": {
        "href": "../../../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "extracted text",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…oned by their first line and the extracted text of every slide; a mention read i…",
      "file": {
        "href": "../deck/index.html",
        "label": "ingestion/documents/deck.md"
      },
      "href": "../deck/index.html#L7",
      "kind": "recognised",
      "line": 7,
      "passages": 2,
      "surface": "extracted text",
      "title": "Deck",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…text position by position under \"Extracted text — 24 pages · indexed and searcha…",
      "file": {
        "href": "../../../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 2,
      "surface": "Extracted text",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…en one is known, the viewer, the extracted text and the download opening on dema…",
      "file": {
        "href": "../../../../specs/screens/pages/entity-page/index.html",
        "label": "screens/pages/entity-page.md"
      },
      "href": "../../../../specs/screens/pages/entity-page/index.html#L41",
      "kind": "recognised",
      "line": 41,
      "passages": 2,
      "surface": "extracted text",
      "title": "Entity page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…e equal to a heading, or similar extracted text; the commit and the folder proxi…",
      "file": {
        "href": "../../../../briefs/meetings/2026/2026-07-09-twin-resources-review/twin-resources-review/index.html",
        "label": "meetings/2026/2026-07-09-twin-resources-review/twin-resources-review.md"
      },
      "href": "../../../../briefs/meetings/2026/2026-07-09-twin-resources-review/twin-resources-review/index.html#L13",
      "kind": "recognised",
      "line": 13,
      "surface": "extracted text",
      "title": "Twin resources review",
      "type": "meeting",
      "typeLabel": "Meeting"
    },
    {
      "context": "… site: the download link and the extracted text remain. A committed preview shar…",
      "file": {
        "href": "../preview/index.html",
        "label": "ingestion/documents/preview.md"
      },
      "href": "../preview/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted text",
      "title": "Preview",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…ther, the file for download, its extracted text and its viewer. An office docume…",
      "file": {
        "href": "../../sources/representation/index.html",
        "label": "ingestion/sources/representation.md"
      },
      "href": "../../sources/representation/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted text",
      "title": "Representation",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…e forms of one document works on extracted text, never on binary content, and ne…",
      "file": {
        "href": "../../../../specs/decisions/inference/minhash-for-twin-resources/index.html",
        "label": "decisions/inference/minhash-for-twin-resources.md"
      },
      "href": "../../../../specs/decisions/inference/minhash-for-twin-resources/index.html#L8",
      "kind": "recognised",
      "line": 8,
      "surface": "extracted text",
      "title": "MinHash for twin resources",
      "type": "decision",
      "typeLabel": "Decision"
    },
    {
      "context": "…ts kind and, for a document, its extracted text position by position, and the pa…",
      "file": {
        "href": "../../../../specs/objects/ingestion/representation/index.html",
        "label": "objects/ingestion/representation.md"
      },
      "href": "../../../../specs/objects/ingestion/representation/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted text",
      "title": "Representation",
      "type": "business_object",
      "typeLabel": "Business object"
    },
    {
      "context": "…g converter. The document stays downloadable, without preview or extracted text.",
      "file": {
        "href": "../../../../specs/rules/documents/conversion-failed/index.html",
        "label": "rules/documents/conversion-failed.rule.md"
      },
      "href": "../../../../specs/rules/documents/conversion-failed/index.html#L7",
      "kind": "recognised",
      "line": 7,
      "surface": "extracted text",
      "title": "Conversion failed",
      "type": "rule",
      "typeLabel": "Business rule"
    },
    {
      "context": "…oon as the tab is shown, and the extracted text, once. The tabs follow the tabli…",
      "file": {
        "href": "../../../../specs/screens/pages/meeting-page/index.html",
        "label": "screens/pages/meeting-page.md"
      },
      "href": "../../../../specs/screens/pages/meeting-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "surface": "extracted text",
      "title": "Meeting page",
      "type": "screen",
      "typeLabel": "Screen"
    }
  ]
}
