slide 1
Transcript publication framing
slide 2
Why publish transcripts: the spoken corpus counts as much as the written one.
slide 3
What a transcript is: cues with a timecode, a speaker and a text, read by the reader plugin.
slide 4
Pseudonymisation before indexing: no real name enters the search index.
slide 5
The dictionary of pseudonyms is never written into the published site.
slide 6
Roles kept instead of pseudonyms when the configuration asks for it.
slide 7
Publication is off by default: publish_transcripts must be requested explicitly.
slide 8
Detected personal data outside the dictionary raises a finding for review.
slide 9
Timecodes are anchors: a cue is citable from any page.
slide 10
Cues are grouped by consecutive speaker in the rendered page.
slide 11
The decisions extracted from a meeting point at their cue.
slide 12
Twin resources: a deck, its notes and its transcript form one page.
slide 13
Grouping criteria: folder, date and textual overlap.
slide 14
One entry in the search index for the three files.
slide 15
The original file stays downloadable, whatever the conversion did.
slide 16
Conversion at publication, cached by fingerprint.
slide 17
The extracted text feeds recognition and similarity alike.
slide 18
The position of a passage, page or slide, is kept for the citation.
slide 19
What the reader sees without JavaScript: the first page, the text, the notes.
slide 20
The viewer loads on demand and never in the initial bundle.
slide 21
Staleness: a transcript older than the threshold is reported like a note.
slide 22
Open questions: keep_roles per source, the size of the cache.
slide 23
Next steps: the privacy framing note, the review process.
slide 24
Questions