norgitov/ trends
Technology · people · ideas
Back to discovery/arXiv1 hour ago

Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

The study challenges the common practice of removing stopwords in legal text analysis, showing that it can distort the accuracy of legal classification tasks. By testing various stopword lists on Supreme Court opinions, the research finds that removing stopwords often performs worse than keeping them, suggesting that this preprocessing step may hinder the recovery of important legal signals.

Open original
SIGNAL FROM THE SOURCE
18
September 2026
Tracking sinceSeptember 18, 2026
MomentumMore observations needed
DiscussionNo comment count provided
PublishedSeptember 18, 2026Gregory M. Dickinson
BEHIND THE NUMBERS

How interest changes

The publication is the signal

This official source does not publish popularity metrics. The story is refreshed from its RSS feed.

Real observations only. History before source connection is not reconstructed.

WHY IT MAY MATTER

This research is relevant for legal scholars using text-as-data methods, as it highlights the potential negative impact of removing stopwords on the accuracy of legal classification tasks.

A useful discovery?
KEEP EXPLORING

Connected ideas

Explore topic