ENTIA coverage · 10 published registry countries · 51 Entia Home launch countries · 38 Europe + 13 LatAm · 57-country generator matrix
Canonical person entity · public technical record

Fernando
Vilches

I am the founder and architect of ENTIA. I work on how to represent, verify and serve business identity so AI systems and agents can recognize a company, resolve conflicts between sources and retrieve a canonical representation backed by evidence.

ENTIA · entia.systems Madrid / Tallinn PERSON ID · /#founder
01 / CANONICALHTML
02 / STRUCTUREDJSON-LD
03 / EVIDENCEBOE · BORME · VIES
04 / AGENTSMCP
05 / OBSERVEEDGE TELEMETRY
OPERATORS THAT CONSUME OR CITE THE CORPUSIDENTITY → RETRIEVAL → MEASUREMENT
FV
PERSON ENTITY
#founder
OpenAI
Anthropic
Google
Perplexity
xAI
DeepSeek
Microsoft
Meta
Mistral AI
Apple
Amazon
ByteDance
Cohere
CCBot

Where ENTIA came from

Pre-ENTIA · masterhair.academy

I was not wrong about the demand. I was missing the trust layer.

Before ENTIA I built Master Hair Academy, a digital academy for balayage, barbering and hairstyling courses. I wanted a scalable, self-service business, and I saw generative search as a surface where a new brand could compete before the inertia of traditional SEO set in.

I started measuring what different LLMs returned, changing JSON-LD, landing pages and URL structures, and comparing thousands of answers. Citation was too variable to work on reproducibly, so I moved the problem upstream: to indexing and to how the entity is represented.

The academy got discovery and traffic. But customers reached checkout and then looked elsewhere for a sign that the mentors were really tied to the course, and that corroboration did not exist. The demand was there; the trust was not.

Today, without me investing in it, enquiries still reach the academy. That is the proof the approach was right: the market exists. What did not exist was a way for a machine, and then a person, to check who was behind it. That is ENTIA.

Read the technical origin →
BEHIND THE SCREEN

I design the architecture. Coding agents write it.

All of ENTIA is built on my architecture and executed by coding agents. It is not an AI experiment: it is how the company runs.

It took me a long time to understand how to really work with AI and coding agents: where they decide, where I decide, and which executable rules are needed so they never make the same mistake twice. Once I understood it, I stopped buying tools. Today I have everything I need, and it is proprietary and native to ENTIA.

22months building
14.3 Btokens across Claude and Codex
(average 650 M / month)
11.3 Mcompanies from official registries
in 10 countries
760+HTTP API endpoints
12public MCP tools
20+deployed services
(containers and Workers)
110scheduled jobs
in production
830+test files

No CRM. No third-party dashboards. No subscriptions.

What other companies rent by subscription, I built myself inside ENTIA, on ENTIA’s own infrastructure and data.

Paid CRM

Own CRM and customer intelligence

Every customer, their product, their payment and their behaviour in a single internal view.

BI and dashboard suite

Mission Control

More than 60 data readers: revenue, machine consumption, indexing, funnel and platform health.

Email marketing platform

Own mail server

Sending, DKIM signing, suppression, bounces, reputation and tracking with no intermediaries.

Web analytics

Edge telemetry

Every request from bots, crawlers and agents classified and measured where it happens.

External monitoring

Availability sentinel

Continuous endpoint monitoring with alerts through a single channel.

Cloud CI

Rule kernel and merge gate

Executable rules and tests that stop a change before it reaches production.

TECHNICAL SUMMARY

What ENTIA is on the inside

Verified business identity infrastructure for machines, anchored to official registries. Edge on Cloudflare (Workers, R2, KV), origin on Hetzner (FastAPI in containers), data on R2 queried with DuckDB, a public MCP server in TypeScript and automated billing with Stripe. Crawlers are served from the edge: they consume the corpus without scaling the origin.

It has not been easy. Two crawling outages, an emergency migration and leaving two clouds behind have shaped this architecture.

COUNTRIES WITH ENTIA HOME
SpainFranceUnited KingdomSwitzerlandCzechiaNorwayFinlandSwedenEstoniaIreland

What I am working on now

A central part of my current work is understanding how AI labs and providers organize their bot and crawler families, what uses they declare for each one, and what can be measured externally in a reproducible way. The goal is to build a methodology that rigorously separates observed access, declared purpose and internal use that cannot be inferred from telemetry alone. In parallel I am studying whether robots.txt, a convention born for crawl exclusion, has become an overloaded control point for governing use purposes, rights and machine access in the AI era.

Business identity for machines

Canonical representations of companies, anchored to official sources and served in formats that automated systems can retrieve.

ENTIA · definition and architecture →

Agent infrastructure

MCP and HTTP interfaces that let AI agents query business identity and evidence in structured form.

ENTIA MCP →

Provenance and resolution

How to separate claim, source, date and evidence strength when several sources describe the same entity.

Methodology →

AI-lab crawlers, telemetry and TDM Reservation V3.1

I study how to distinguish crawler families associated with training, search, real-time retrieval, grounding or other declared uses; how to measure request, delivery, path, volume, frequency and persistence; and how to bind each access to the rights representation actually served at that moment. V3.1 is the first-contact evidentiary representation anchored to policy V3: it fixes canonical immutable bytes in robots.txt so a specific request can be tied to the exact signal served, without turning access into automatic proof of training, remote parsing, understanding, compliance or downstream use.

Public telemetry →
TDM history · V1 → V3.1 →
ACTIVE RESEARCH · GOVERNANCE ARCHITECTURE

robots.txt as an accidental AI control plane

A mechanism created in 1994 for crawl exclusion is absorbing decisions about search, retrieval, grounding, training and TDM reservations. I am studying where it stops being sufficient and what architecture should sit above it.

The problem is not robots.txt itself. The problem is assigning it a governance function its original trust model was never designed to support.
01 · ORIGIN 1994: crawl exclusion

Martijn Koster proposed the mechanism on 25 February 1994; consensus on the original standard followed on 30 June. Its purpose was simple: indicate which areas a crawler should not traverse. RFC 9309 formalized the protocol in 2022 and still makes clear that it is neither access authorization nor a security measure.

02 · OVERLOAD 2026: one signal, too many purposes

OpenAI separates search from training; Anthropic distinguishes crawling, search and user-initiated retrieval; Google-Extended acts as a control token for training and grounding. Not all AI depends on robots.txt, but a critical part of access and use governance ends up leaning on it.

03 · SYSTEMIC RISK Integrity, scale and licensing

I am studying what happens if the signal changes, is manipulated, remains cached or is interpreted differently; and what happens if major content producers express TDM reservations at scale. Agent identity, versioning, temporality, evidence, purpose, exceptions and licensing need a stronger layer.

V3.1 · EXPERIMENTAL DIRECTION

The architecture I am exploring separates layers: robots.txt for discovery and first contact; a structured, versioned and verifiable policy as authority; enforcement outside robots.txt; and temporal evidence of which policy each request received. ENTIA V3.1 keeps the file as a first-contact evidentiary representation while anchoring it to canonical bytes, a hash, a manifest, a policy registry and temporal evidence.

VERIFIABLE MILESTONES

The work is measured by what it forced me to solve

These are not mentions or reputation badges. They are points where an experiment, real load, external measurement or institution forced a hypothesis to become architecture or method.

M01 · CANONICAL IDENTITY2024–OCT 2025

From a landing page to a multi-country canonical identity

Close to eight months of iteration across URLs, JSON-LD, field structure and retrieval behaviour in different models. The question moved from “how do I obtain a citation?” to “what representation makes different systems converge on the same entity?”. Then came weighting: not every item of evidence could carry the same weight, which became the Entia Home Risk Score.

CANONICAL URLJSON-LDRISK SCOREINDEXED BY OCT 2025
M02 · ORIGIN FAILURE16–24 APR 2026

The crawlers proved the architecture was not ready yet

Googlebot concurrently traversed a sitemap containing 48,000 identity URLs and pressure on the then-active query and compute layer produced a cascade of errors. A week later another crawling spike forced an emergency migration. The answer was not to block bots; it was to make serving them unable to scale the origin.

48,000 URLS483 × 5XXORIGIN PRESSUREEMERGENCY MIGRATION
M03 · EDGE-FIRSTMAY–JUL 2026

Bots can read all of ENTIA without driving up origin consumption or the infrastructure bill

The response became edge-first: aggressive caching, static serving, born-in-edge and prewarming. The 26 April close already recorded 229,000 Meta-ExternalAgent accesses, 45,000 GPTBot accesses and 18,000 ClaudeBot accesses. The edge became an economic and observability boundary rather than merely a latency optimization.

229K META45K GPTBOT18K CLAUDEBOTPREWARMBORN-IN-EDGE
M04 · CITATION REGIME SHIFT20–21 AUG 2026

Bing stopped behaving like it had in the preceding months

Bing AI Performance began recording 100–200 citations per day versus a typical 10–30/day in July, while the three-month cumulative count moved from 3.5K to 3.7K. This does not establish causality; it does document an observable regime change after months of machine consumption.

3.5K → 3.7K100–200 / DAY202 PEAK DAYBING / COPILOT

Questions I am working on

What should a company’s canonical identity be when the consumer is a machine rather than a person?
How should an agent establish that the company it found is the correct legal entity?
What should happen when two reliable sources disagree about the same attribute?
How should provenance survive when an answer is assembled at inference time?
How should rights, permissions and reservations be expressed so they can be understood machine to machine?
How can evidence distinguish a training crawler from search, grounding or real-time retrieval when the same operator maintains several bot families?
What can access telemetry actually establish — request, delivery, path, volume, frequency and persistence — and what must remain unknown about a lab's internal use?
How can each request be bound to the exact rights version and representation served at that moment without inferring remote parsing, understanding or compliance?
Can robots.txt remain a sufficient control point when the same operator separates crawling, search, grounding, user-directed retrieval and training?
How should the policy a crawler received be authenticated, versioned and preserved as evidence without turning an advisory text file into an authorization system?
How should an accidental or malicious change to robots.txt be detected and attributed before compliant crawlers consume it as policy input?
What rights-resolution and licensing infrastructure would the web need if major content producers expressed TDM reservations at scale?

Technical record

Before ENTIA

The first experiments are carried out on another domain, before the project is anchored on entia.systems. For close to eight months I iterate on URL structures and structured data: some are retrieved by certain LLMs while others do not reach the same representation. My objective becomes one stable canonical URL that can work across different systems.

entia.systems

Registering and using entia.systems marks the point where the earlier experiments are anchored in a dedicated infrastructure. From that point the canonical URL stops being only a retrieval experiment and becomes the basis of what will later be called Entia Home.

Afterward

Once the canonical URL is stabilized, the work shifts to weighting identity fields. The most difficult part is building an algorithmic Risk Score that assigns different weights to signals and fields and determines the composition of an Entia Home.

Oct · 2025

By October 2025, Entia Home pages were already indexed. The canonical identity layer and its weighting system therefore predate the public 2026 work on MCP, telemetry and Cognitive Resistance.

Cognitive Resistance and Risk Score. At ENTIA, Fernando Vilches develops the Cognitive Resistance framework: the search, interpretation, reconciliation, inference and verification friction a machine must resolve before it can use a claim about an entity with an explicit evidence basis. The ENTIA Risk Score uses the canonical direction 0 RC = minimum resistance; 100 RC = maximum resistance. The methodology and knowledge cluster document the framework and its evidentiary boundaries.

ENTIA seed proposition. «Un LLM siempre elegirá recomendar la entidad con menor Resistencia Cognitiva.» — “An LLM will always choose to recommend the entity with the lowest Cognitive Resistance.” Canonical axiom and source.

16 Apr · 2026

First documented scale incident: Googlebot concurrently crawls a sitemap containing 48,000 identity URLs. Pressure on the then-active query and compute layer produces a 502 cascade. The incident record counts 483 5xx responses over roughly three hours.

23–24 Apr · 2026

A second crawling spike drives egress up and the then-active platform cuts billing after the threshold is exceeded. ENTIA performs an emergency migration and is operational again on 24 April.

May–Jul · 2026

The architectural response evolves into an edge-first principle: let crawlers consume the corpus without allowing repeated reads to scale the origin. Aggressive caching, static serving, born-in-edge tooling and Entia Home prewarming are documented.

20–21 Aug · 2026

Bing AI Performance monitoring records 3.5K→3.7K citations over three months and daily levels of 100–200 citations, compared with a typical 10–30/day in July. The highest day available at the time was 202 citations.

28 Jul · 2026

ENTIA presents the case for its machine-readable rights reservation to the AI Office and DG CNECT · Copyright, together with the broader problem of governing automated access and use through machine-readable signals.

24 Aug · 2026

The European Commission formally invites a written contribution for consideration. The contribution is submitted that same day and updated on 2 September. This is a historical fact in the record: an invitation to contribute to an open process, not approval or institutional endorsement. Full case →

Jul – Sep · 2026

Joint study E082 on the access → eligibility → retrieval → citation → influence chain, with instruments kept separate by construction and the five-layer framework I proposed, attributed to my name. An ongoing methodological collaboration, not a published result. Study record (restricted access) →

19 Sep · 2026

I put the TDM Reservation V3 and the V3.1 first-contact evidentiary representation into production. robots.txt stops being an isolated signal: the representation is anchored to canonical bytes, a hash, a manifest and a policy registry so a request can be tied to the exact signal served at that moment. The experiment turns a historical limitation of robots.txt into a verifiable architecture question.

2026

Publication of methodology, MCP and machine-access telemetry turns retrieval, provenance and corpus consumption into observable variables.

This record does not backdate articles. When earlier work is documented here, the date of the original artifact will be distinguished from the date it was added to this archive.

Notes

28 August 2026 · Note 001

From Master Hair Academy to ENTIA

How an experiment in generative-search distribution exposed discovery, indexing, identity and trust as different layers of the same problem.

Read the note →
28 August 2026 · Note 002

Before visibility, a machine has to know who you are

The traditional web mixes identity, marketing and content. An agent needs something different: a stable representation of the entity, traceable sources and a way to resolve contradictions before it can trust what it retrieves.

Read the note →
28 August 2026 · Note 003

Do not block the bots: keep them off the origin

How two crawling incidents turned the edge into a technical and economic boundary: crawlers can read the whole corpus without driving up origin consumption or the infrastructure bill.

Read the note →

Machine representations

One entity, several coherent representations. The person identifier remains stable.

SERIES · TECHNICAL NOTES

Keep reading

NOTE 001From Master Hair Academy to ENTIAThe demand was there. The trust was not.READ →NOTE 002Before you can be visible, a machine has to know who you areIdentity, evidence and canonicity before visibility.READ →NOTE 003I did not block the bots: I took them off the originTwo crawling incidents and the edge-first architecture.READ →PROFILEFernando VilchesFounder and architect of ENTIA. Technical record, milestones and behind the screen.READ →
THIS PAGE FOR MACHINESHTMLJSON-LDMARKDOWNESPAÑOL