{
  "id": "keywords/extracted",
  "mentions": [
    {
      "context": "…ument, named slide 3 or by its timecode, captioned by the first line of its extracted text, and leading to that text; once the viewer runs, the entry also shows…",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 6,
      "surface": "extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…not be generated.\" with, when its text was read all the same, \"The text was extracted all the same: it is indexed, cited on the other pages, and readable below.…",
      "file": {
        "href": "../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 5,
      "surface": "extracted",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "Three views follow behind the tabs \"Document\", \"Extracted text\" and \"Related notes\", with \"Download the original\" at the end of the bar, the original file place…",
      "file": {
        "href": "../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "Extracted",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…DF, a transcript. A document is indexed and its words are recorded from its extracted text, but it enters the model only through what it documents. Its page off…",
      "file": {
        "href": "../../glossary/ingestion/documents/document/index.html",
        "label": "ingestion/documents/document.md"
      },
      "href": "../../glossary/ingestion/documents/document/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 2,
      "surface": "extracted",
      "title": "Document",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…esource carries its path, commit and date, and, for documents, the metadata extracted by a reader, the PDF produced by a converter and the text extracted from t…",
      "file": {
        "href": "../../specs/objects/ingestion/resource/index.html",
        "label": "objects/ingestion/resource.md"
      },
      "href": "../../specs/objects/ingestion/resource/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 2,
      "surface": "extracted",
      "title": "Resource",
      "type": "business_object",
      "typeLabel": "Business object"
    },
    {
      "context": "…load link of each file, its PDF, a rail of its pages, slides or cues, their extracted text in disclosure blocks, and the viewer that opens the PDF on demand. Th…",
      "file": {
        "href": "../../specs/screens/pages/entity-page/index.html",
        "label": "screens/pages/entity-page.md"
      },
      "href": "../../specs/screens/pages/entity-page/index.html#L39",
      "kind": "recognised",
      "line": 39,
      "passages": 2,
      "surface": "extracted",
      "title": "Entity page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…that page in it and the entry of the page shown is marked. Under the rail, \"Extracted text\" gives the text of every position in a disclosure block, the first on…",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 6,
      "surface": "Extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… the way the mentions panel cites it, so that a mention read on slide three lands on slide three. A document without extracted text says so; the download stays.",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L12",
      "kind": "recognised",
      "line": 12,
      "passages": 6,
      "surface": "extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… \"+\", over fixed steps, and the field \"⌕ in the document\" that searches the extracted text of the positions, jumps to the first matching page and says how many …",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L14",
      "kind": "recognised",
      "line": 14,
      "passages": 6,
      "surface": "extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…rter produces thumbnails. The text shown and searched is the text the build extracted from the PDF, cut at build.extracted_text_max_chars in all; the download g…",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 6,
      "surface": "extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "Open a position of the rail → its extracted text, and its page once the viewer runs",
      "file": {
        "href": "../../specs/viewers/document-viewer/index.html",
        "label": "viewers/document-viewer.md"
      },
      "href": "../../specs/viewers/document-viewer/index.html#L27",
      "kind": "recognised",
      "line": 27,
      "passages": 6,
      "surface": "extracted",
      "title": "Document viewer",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…cited on the other pages, and readable below.\", else \"Its text could not be extracted either: the original stays downloadable, and the note that describes it st…",
      "file": {
        "href": "../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 5,
      "surface": "extracted",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…failed: timed out after 120 s\", and the check identifier after it; then the extracted text position by position under \"Extracted text — 24 pages · indexed and s…",
      "file": {
        "href": "../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 5,
      "surface": "extracted",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…ck identifier after it; then the extracted text position by position under \"Extracted text — 24 pages · indexed and searchable\". The panel adds \"State of the re…",
      "file": {
        "href": "../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 5,
      "surface": "Extracted",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…vailable, .pdf preview failed, the word written and not only coloured, text extracted or missing, with the note that the same finding stands in the publication …",
      "file": {
        "href": "../../specs/screens/pages/edge-cases/index.html",
        "label": "screens/pages/edge-cases.md"
      },
      "href": "../../specs/screens/pages/edge-cases/index.html#L19",
      "kind": "recognised",
      "line": 19,
      "passages": 5,
      "surface": "extracted",
      "title": "Edge cases",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "… as its script runs, with its counter, its zoom and its find field over the extracted text, and the strip follows the page shown. Without JavaScript the browser…",
      "file": {
        "href": "../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "extracted",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…lication, cached by fingerprint\" and \"The original stays downloadable\". The extracted text view gives the text of every page in a disclosure block, anchored the…",
      "file": {
        "href": "../../specs/screens/pages/document-page/index.html",
        "label": "screens/pages/document-page.md"
      },
      "href": "../../specs/screens/pages/document-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "passages": 3,
      "surface": "extracted",
      "title": "Document page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…through what it documents. Its page offers the file for download, shows its extracted text page by page and, when the conversion produced a PDF, opens it in a v…",
      "file": {
        "href": "../../glossary/ingestion/documents/document/index.html",
        "label": "ingestion/documents/document.md"
      },
      "href": "../../glossary/ingestion/documents/document/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 2,
      "surface": "extracted",
      "title": "Document",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…etadata extracted by a reader, the PDF produced by a converter and the text extracted from that PDF, page by page, or the cues of a transcript with their timeco…",
      "file": {
        "href": "../../specs/objects/ingestion/resource/index.html",
        "label": "objects/ingestion/resource.md"
      },
      "href": "../../specs/objects/ingestion/resource/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "passages": 2,
      "surface": "extracted",
      "title": "Resource",
      "type": "business_object",
      "typeLabel": "Business object"
    },
    {
      "context": "…e file name, its kind and its page count when one is known, the viewer, the extracted text and the download opening on demand. A document without a note leads, …",
      "file": {
        "href": "../../specs/screens/pages/entity-page/index.html",
        "label": "screens/pages/entity-page.md"
      },
      "href": "../../specs/screens/pages/entity-page/index.html#L41",
      "kind": "recognised",
      "line": 41,
      "passages": 2,
      "surface": "extracted",
      "title": "Entity page",
      "type": "screen",
      "typeLabel": "Screen"
    },
    {
      "context": "…eclaration in the frontmatter, a shared base name, a title equal to a heading, or similar extracted text; the commit and the folder proximity only reinforce it.",
      "file": {
        "href": "../../briefs/meetings/2026/2026-07-09-twin-resources-review/twin-resources-review/index.html",
        "label": "meetings/2026/2026-07-09-twin-resources-review/twin-resources-review.md"
      },
      "href": "../../briefs/meetings/2026/2026-07-09-twin-resources-review/twin-resources-review/index.html#L13",
      "kind": "recognised",
      "line": 13,
      "surface": "extracted",
      "title": "Twin resources review",
      "type": "meeting",
      "typeLabel": "Meeting"
    },
    {
      "context": "…ce document into a PDF at build, by a converter plugin, so that its text is extracted from that PDF and nothing else, and the conversion: block of the configura…",
      "file": {
        "href": "../../glossary/ingestion/documents/conversion/index.html",
        "label": "ingestion/documents/conversion.md"
      },
      "href": "../../glossary/ingestion/documents/conversion/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted",
      "title": "Conversion",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…tted next to it, a rail of its slides captioned by their first line and the extracted text of every slide; a mention read in a deck cites its slide. A deck is o…",
      "file": {
        "href": "../../glossary/ingestion/documents/deck/index.html",
        "label": "ingestion/documents/deck.md"
      },
      "href": "../../glossary/ingestion/documents/deck/index.html#L7",
      "kind": "recognised",
      "line": 7,
      "surface": "extracted",
      "title": "Deck",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…ts text any other way, so that one extraction path serves every format. The extracted text feeds the search index, the recognition of occurrences and the compar…",
      "file": {
        "href": "../../glossary/ingestion/documents/extracted-text/index.html",
        "label": "ingestion/documents/extracted-text.md"
      },
      "href": "../../glossary/ingestion/documents/extracted-text/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted",
      "title": "Extracted text",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "… source keeps the PDFs out of the published site: the download link and the extracted text remain. A committed preview shares the base name of its original and …",
      "file": {
        "href": "../../glossary/ingestion/documents/preview/index.html",
        "label": "ingestion/documents/preview.md"
      },
      "href": "../../glossary/ingestion/documents/preview/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted",
      "title": "Preview",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "…al, and its page shows them one after the other, the file for download, its extracted text and its viewer. An office document without markdown next to it is rep…",
      "file": {
        "href": "../../glossary/ingestion/sources/representation/index.html",
        "label": "ingestion/sources/representation.md"
      },
      "href": "../../glossary/ingestion/sources/representation/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted",
      "title": "Representation",
      "type": "term",
      "typeLabel": "Term"
    },
    {
      "context": "The text similarity that helps recognise the forms of one document works on extracted text, never on binary content, and never compares every pair. Each text go…",
      "file": {
        "href": "../../specs/decisions/inference/minhash-for-twin-resources/index.html",
        "label": "decisions/inference/minhash-for-twin-resources.md"
      },
      "href": "../../specs/decisions/inference/minhash-for-twin-resources/index.html#L8",
      "kind": "recognised",
      "line": 8,
      "surface": "extracted",
      "title": "MinHash for twin resources",
      "type": "decision",
      "typeLabel": "Decision"
    },
    {
      "context": "…entity carry several, each with its path, its kind and, for a document, its extracted text position by position, and the page shows them one after the other. An…",
      "file": {
        "href": "../../specs/objects/ingestion/representation/index.html",
        "label": "objects/ingestion/representation.md"
      },
      "href": "../../specs/objects/ingestion/representation/index.html#L6",
      "kind": "recognised",
      "line": 6,
      "surface": "extracted",
      "title": "Representation",
      "type": "business_object",
      "typeLabel": "Business object"
    },
    {
      "context": "…fice document could not be converted to PDF: timeout, size, corruption or missing converter. The document stays downloadable, without preview or extracted text.",
      "file": {
        "href": "../../specs/rules/documents/conversion-failed/index.html",
        "label": "rules/documents/conversion-failed.rule.md"
      },
      "href": "../../specs/rules/documents/conversion-failed/index.html#L7",
      "kind": "recognised",
      "line": 7,
      "surface": "extracted",
      "title": "Conversion failed",
      "type": "rule",
      "typeLabel": "Business rule"
    },
    {
      "context": "…, or merged with other documents only. Its words are recorded from the text extracted from its PDF or its cues, but nobody wrote about it. The to-do page lists …",
      "file": {
        "href": "../../specs/rules/documents/document-without-markdown/index.html",
        "label": "rules/documents/document-without-markdown.rule.md"
      },
      "href": "../../specs/rules/documents/document-without-markdown/index.html#L7",
      "kind": "recognised",
      "line": 7,
      "surface": "extracted",
      "title": "Document without markdown",
      "type": "rule",
      "typeLabel": "Business rule"
    },
    {
      "context": "…preview in the document viewer, opened as soon as the tab is shown, and the extracted text, once. The tabs follow the tablist pattern the accessibility rule ask…",
      "file": {
        "href": "../../specs/screens/pages/meeting-page/index.html",
        "label": "screens/pages/meeting-page.md"
      },
      "href": "../../specs/screens/pages/meeting-page/index.html#L16",
      "kind": "recognised",
      "line": 16,
      "surface": "extracted",
      "title": "Meeting page",
      "type": "screen",
      "typeLabel": "Screen"
    }
  ]
}
