Skip to content

Datasets

glitchdata/subdomain-wordlist

♡ Like 0 + Adopt 0

The Subdomain Wordlist is a curated, high-quality dataset designed to power accurate and efficient subdomain enumeration for security assessments, penetration testing, asset discovery, and attack-surface management. Built from millions of real-world DNS records, OSINT sources, enterprise naming conventions, and observed patterns across global infrastructures, this wordlist provides unparalleled coverage for discovering hidden, legacy, and undocumented subdomains.

Key Features

Extensive Coverage: Contains thousands to millions of subdomain name candidates covering corporate, cloud, SaaS, dev, staging, internal, and region-specific patterns. Real-World Data: Generated from live DNS scans, public corpuses, and common naming schemas used in enterprise networks. Optimized for Recon Tools: Fully compatible with tools such as subfinder, amass, assetfinder, dnsx, ffuf, and gobuster. High Signal-to-Noise: Cleaned, normalized, deduplicated, and weighted to minimize false positives and maximize discovery. Regularly Updated: Continuously refreshed to reflect evolving infrastructure patterns, cloud services, devops environments, and security research findings. Context-Aware Lists: Includes specialized subsets for Cloud platforms (AWS, Azure, GCP) CI/CD and development environments (git, dev, staging, QA, sandbox) Security and IT systems (vpn, portal, wsus, oauth, sso) Industry-specific structures (banking, telecom, logistics, retail)

Use Cases

External attack-surface mapping Subdomain brute-forcing Enterprise asset discovery Red-team and pentest reconnaissance Bug bounty reconnaissance Shadow IT identification Monitoring for newly emerging infrastructure

Benefits

Dramatically improves detection of hidden or unknown subdomains Reduces time spent on manual enumeration Increases recon efficiency and discovery accuracy Ideal for both offensive and defensive cybersecurity operations

Formats Provided

Plain text wordlist (.txt) Ranked wordlist by frequency/use

Where the data comes from now

Where to get it

glitchdata catalogues this dataset but does not host the files here. It is published in the glitchdata shop (US$50), where the download covers every format listed above.

Last updated on the source site: 2025-11-22.

Related datasets

Browse datasets →

Predict whether income exceeds $50K a year from 1994 US Census attributes. 48,842 people, 14 attributes.

CSV · Updated 4d ago · ↓ 18

Physicochemical tests and sensory quality scores for 6,497 Portuguese vinho verde wines.

CSV · Updated 4d ago · ↓ 10

Sepal and petal measurements for 150 iris flowers across three species — the classic first classification dataset.

CSV · Updated 4d ago · ↓ 8

Monthly and annual global land–ocean temperature anomalies from 1880 to the present, relative to the 1951–1980 average.

CSV · Updated 4d ago · ↓ 6

Bill, flipper and body-mass measurements for 344 Adélie, Chinstrap and Gentoo penguins from Antarctica.

CSV · Updated 4d ago · ↓ 4

Every yellow taxi trip in New York City in January 2024: 2,964,624 trips with pick-up and drop-off times, zones, distances and fares.

Parquet · Updated 4d ago · ↓ 4