The crawler behind Common Crawl, an open, freely downloadable web archive used as training data by many AI labs and researchers.

CCBot is not run by any single AI company -- it builds the shared Common Crawl dataset, which numerous language models beyond the well-known named AI-company bots draw on for training. A page crawled by CCBot may end up influencing far more models than the named bots alone suggest.

Related terms