training Allen Institute for AI respects robots.txt

Ai2Bot-Dolma

Allen Institute for AI uses this crawler to collect content for AI model training. Collects content for the open Dolma training dataset.

What is Ai2Bot-Dolma?

Collects content for the open Dolma training dataset. Training-data crawlers collect publicly available pages to build future AI models. Allowing one means your content can shape how models write and reason about your topics, and your site may be quoted from model memory. Blocking one keeps future training runs away, though anything already collected typically stays collected.

If you want AI engines to know your brand and quote your pages from memory, this is the crawler that gets you there.

Does Ai2Bot-Dolma respect robots.txt?

Yes. Ai2Bot-Dolma honors robots.txt directives, including wildcard rules and Allow lines. A Disallow rule for this crawler is respected — but remember robots.txt is a voluntary protocol across the whole industry, not a wall.

How to block Ai2Bot-Dolma

Add this group to your robots.txt to deny Ai2Bot-Dolma across your whole site:

User-agent: Ai2Bot-Dolma
Disallow: /

If you monetize content behind a paywall or care about control, allowing broad training crawlers means giving that material away with no attribution loop.

How to allow Ai2Bot-Dolma

Ai2Bot-Dolma is allowed by default on every site. It only gets blocked if a matching rule exists — most often a global User-agent: * group with Disallow: /. To exempt it while blocking other bots, put an empty Disallow group for it before the global block:

User-agent: Ai2Bot-Dolma
Disallow:

User-agent: *
Disallow: /

Is Ai2Bot-Dolma currently allowed on your site?

Paste your robots.txt into the free auditor and see verdicts for all 61 AI crawlers — including the specific paths that matter.

Related AI crawlers