Tessera

Span-level, multi-label language identification.

Document-level language ID gives one language per page. Real pages are not one language: the navigation is English, the body might be Ojibwe, the comments might be three other things. Filtering a corpus on the document label throws that text away, and low-resource languages are the ones that cannot spare it.

On the benchmark corpus, document-level LID recovers 0.0% of the low-resource spans it would otherwise discard or misfile. Tessera recovers 86% (90% on the rarest tier). Tessera runs behind GlotLID / OpenLID, not instead of them.

Loading the model (17 MB, runs entirely in your browser)…
Analyse some text to see what a single-label filter would cost.

Spans

Detail

rangelenrolelanguagestier scoretext