Skip to content

Text Data Cleaning for NLP and LLMs

Cleaning and normalising text data: encoding issues, boilerplate, markup, language detection and quality filters.

Editorial team 1 min read

Raw text is messy. Cleaning it improves search, analytics and model training.

Encoding

Fix mis-encoded characters (mojibake), normalise Unicode forms, and standardise quotes and whitespace.

Remove Boilerplate

Strip navigation menus, cookie banners, footers, signatures and repeated headers from scraped or exported text.

Markup

Convert HTML to text or Markdown, keeping meaningful structure such as headings, lists and tables.

Language Detection

Identify languages to route or filter documents.

Quality Filters

For training data, filter out very short or very long documents, gibberish, spam, repeated text and low-information pages. Heuristics and classifiers both help.

Sensitive Content

Detect and redact personal data and secrets; filter harmful content where appropriate.

Keep Originals

Store raw data alongside cleaned versions so cleaning rules can be changed later.

Don't Over-Clean

Lowercasing, removing punctuation or stripping numbers can destroy meaning that modern models use. Match cleaning to the downstream use.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026