---
title: "I did not block the bots: I took them off the origin"
canonical: https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin
language: en
author: Fernando Vilches (https://entia.systems/#founder)
publisher: ENTIA Systems (https://entia.systems/#organization)
date_modified: 2026-09-24
alternate_language: https://entia.systems/fernando-vilches/notes/bots-at-the-edge-not-the-origin
json_ld: https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin.json
html: https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin
---

# I did not block the bots. *I took them off the origin.*

Two crawling incidents left the platform failing and forced an emergency migration. The answer was not to close the door: it was to turn the edge into an economic boundary.

## Crawlers were both the audience and the threat

ENTIA has always lived with an uncomfortable tension: crawlers are both desired consumers and a potential source of uncontrolled load. Blocking them would have protected the infrastructure, but it would also have cut off the very channel I was trying to study and serve.

> The answer was not to stop them reading. *It was to let them read without touching the origin.*

## The first warning

Googlebot began concurrently crawling a sitemap of 48,000 identity URLs. The query layer and compute concurrency of the time were not sized for that pattern. The pressure reached the health check and caused a cascade of 502 errors.

The record preserves 483 5xx responses over roughly three hours, 242 of them associated with Googlebot. The first response was conventional: caching, concurrency tuning and a crawler circuit breaker.

## The failure that changed the infrastructure

A week later the problem was no longer only availability. On 22 April egress from crawling identity routes spiked, and in the early hours of the 23rd the platform of the time cut billing after the configured threshold was exceeded.

Unable to reactivate it immediately, I decided to migrate. On 24 April ENTIA was running again, with Cloudflare in front and a compatibility layer for legacy data. According to the incident record, the emergency migration was completed in about 36 hours.

## Migrating did not remove the problem

The 26 April session close already recorded 229,000 Meta-ExternalAgent visits, 45,000 GPTBot visits and 18,000 ClaudeBot visits, and set moving Entia Home to pre-rendered storage and aggressive Cloudflare caching as a priority.

That changed how I thought about infrastructure. A bot reading more could not mean ENTIA paying for more servers. The cost of a repeated read had to approach zero, and the origin had to leave the normal machine-consumption path.

> Bots reading ENTIA is fine. *Bots driving up origin consumption or the infrastructure bill is not.*

## The edge as a boundary

By 12 May one of the first operational versions of that principle was documented. Routes that previously returned `cf-cache-status: DYNAMIC` moved under an edge-cache rule: the first request was a MISS, the second already a HIT with a long TTL. Identity routes also started being served as static content.

The design kept evolving. In June came *born in edge* tooling. In July, Entia Home prewarming through the edge, later extended to every generation change. The idea is simple: if many machines are going to read an Entia Home, ENTIA performs the first read that fills the cache so the following requests never reach the origin.

For many websites an aggressive crawler is just traffic to throttle or block. Not for ENTIA: those systems are one of the audiences of the infrastructure, and their consumption is itself a signal I want to observe.

The architecture ended up separating two goals that looked incompatible: **keep the corpus readable by machines and make that readability unable to scale origin cost**. It also opened another line of work: if consumption happens at the edge, it can be measured at the edge. That is where the telemetry on which operators retrieve which routes, how often and from which networks comes from.

## Notes in this series

- [From Master Hair Academy to ENTIA](https://entia.systems/en/fernando-vilches/notes/from-master-hair-academy-to-entia) · [JSON-LD](https://entia.systems/en/fernando-vilches/notes/from-master-hair-academy-to-entia.json) · [Markdown](https://entia.systems/en/fernando-vilches/notes/from-master-hair-academy-to-entia.md)
- [Before you can be visible, a machine has to know who you are](https://entia.systems/en/fernando-vilches/notes/identity-before-visibility) · [JSON-LD](https://entia.systems/en/fernando-vilches/notes/identity-before-visibility.json) · [Markdown](https://entia.systems/en/fernando-vilches/notes/identity-before-visibility.md)
- [I did not block the bots: I took them off the origin](https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin) · [JSON-LD](https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin.json) · [Markdown](https://entia.systems/en/fernando-vilches/notes/bots-at-the-edge-not-the-origin.md)
- [Profile: Fernando Vilches](https://entia.systems/en/fernando-vilches) · [JSON-LD](https://entia.systems/en/fernando-vilches.json) · [Markdown](https://entia.systems/en/fernando-vilches.md)
