norgitov/ trends
Technology · people · ideas
Back to discovery/arXiv2 hours ago

Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch

Researchers present Wieszcz-XIX, a 3.1-billion-word corpus of Polish from 1800 to 1918, compiled from 294,369 documents. The corpus was created using an automated pipeline for filtering and auditing, including removal of duplicates and exclusion of post-1918 content. Language models with 47M to 349M parameters were trained on this corpus.

Open original
SIGNAL FROM THE SOURCE
9
October 2026
Tracking sinceOctober 9, 2026
Momentum—More observations needed
Discussion—No comment count provided
PublishedOctober 9, 2026Szymon Kocur
BEHIND THE NUMBERS

How interest changes

The publication is the signal

This official source does not publish popularity metrics. The story is refreshed from its RSS feed.

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

The corpus may be useful for historical linguistics research and developing period-specific language models. It is important to consider potential biases present in the data.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic