The words of a document as the build reads them: the text of every page of the PDF the converter produced, or of the PDF given as a source, and the cues of a transcript with their timecodes. An office document never gives its text any other way, so that one extraction path serves every format. The extracted text feeds the search index, the recognition of occurrences and the comparison of twin resources, each word keeping the page, slide or cue it was read on, so that a mention cites slide 3 where a note cites a line. The page of the document shows it position by position (3 passages, no note), cut at the configured size (12 passages, no note); the download (30 passages, no note) gives the rest.
Extracted text
Properties3
- Application
- Concordance command line
- Domain
- Ingestion
- Status
- valid
3 declared keys. The rest of the file is free text.
See the neighbourhood map6 pages6Neighbourhood mapExtracted text
Neighbourhood map Extracted text
Distance1 hop
58 neighbours in total, more than the map shows.
The 6 neighbourstextual equivalent
- MentionTermspecializes5
- OccurrenceTermspecializes2
- DeckTermgeneralizes9
- DocumentTermspecializes12
- NoteTermspecializes10
- Twin resourcesTermspecializes7
Six neighbours at most, always named. Beyond that the map teaches nothing: the list takes over.