Two different businesses, two different answers

If your content is the product — a subscription publication, a paid research library — then models reproducing it without sending traffic is a genuine commercial threat, and restricting access is a defensible position.

For everyone else the maths is inverted

If your content exists to make people aware of a service you sell, being read by the systems that now answer buying questions is the entire objective. Blocking them removes you from the answers while your competitors stay in them.

Blocking AI crawlers protects content that is the product. It hides content that is the marketing.

Absence doesn’t mean silence

Blocking your own site does not stop models describing you — they use everything written about you elsewhere. It removes your version and leaves theirs, which is exactly the situation where wrong descriptions become hard to correct.

Distinguish training from retrieval

Some crawlers gather training data; others fetch pages to answer a live question with a citation. Blocking the second category is the costly one, because that is the traffic-and-attribution path — worth checking which you are actually excluding.

The middle position

Allow access to your marketing and educational content and restrict genuinely proprietary material, in the same way you would decide what a retrieval layer may read internally. Blanket rules on either side are rarely right.

Whatever you choose, verify it

Directives are frequently applied more broadly than intended, and an accidental block is invisible until visibility falls. Check what is reachable after any change, the same pre-launch discipline as a redesign.