Own CRM and customer intelligence
Every customer, their product, their payment and their behaviour in a single internal view.
I am the founder and architect of ENTIA. I work on how to represent, verify and serve business identity so AI systems and agents can recognize a company, resolve conflicts between sources and retrieve a canonical representation backed by evidence.
Before ENTIA I built Master Hair Academy, a digital academy for balayage, barbering and hairstyling courses. I wanted a scalable, self-service business, and I saw generative search as a surface where a new brand could compete before the inertia of traditional SEO set in.
I started measuring what different LLMs returned, changing JSON-LD, landing pages and URL structures, and comparing thousands of answers. Citation was too variable to work on reproducibly, so I moved the problem upstream: to indexing and to how the entity is represented.
The academy got discovery and traffic. But customers reached checkout and then looked elsewhere for a sign that the mentors were really tied to the course, and that corroboration did not exist. The demand was there; the trust was not.
Today, without me investing in it, enquiries still reach the academy. That is the proof the approach was right: the market exists. What did not exist was a way for a machine, and then a person, to check who was behind it. That is ENTIA.
Read the technical origin →All of ENTIA is built on my architecture and executed by coding agents. It is not an AI experiment: it is how the company runs.
It took me a long time to understand how to really work with AI and coding agents: where they decide, where I decide, and which executable rules are needed so they never make the same mistake twice. Once I understood it, I stopped buying tools. Today I have everything I need, and it is proprietary and native to ENTIA.
What other companies rent by subscription, I built myself inside ENTIA, on ENTIA’s own infrastructure and data.
Every customer, their product, their payment and their behaviour in a single internal view.
More than 60 data readers: revenue, machine consumption, indexing, funnel and platform health.
Sending, DKIM signing, suppression, bounces, reputation and tracking with no intermediaries.
Every request from bots, crawlers and agents classified and measured where it happens.
Continuous endpoint monitoring with alerts through a single channel.
Executable rules and tests that stop a change before it reaches production.
Verified business identity infrastructure for machines, anchored to official registries. Edge on Cloudflare (Workers, R2, KV), origin on Hetzner (FastAPI in containers), data on R2 queried with DuckDB, a public MCP server in TypeScript and automated billing with Stripe. Crawlers are served from the edge: they consume the corpus without scaling the origin.
It has not been easy. Two crawling outages, an emergency migration and leaving two clouds behind have shaped this architecture.
The record does not end with a biography. ENTIA publicly exposes how machines access the corpus, which surfaces they consume and how the ecosystem serving those identities is constructed.
Observable machine access, agent classification and corpus consumption from the edge.
OPEN LIVE →Map of the infrastructure, corpus, machine-readable surfaces and agent relationships.
OPEN →Identity, provenance, resolution and evidence.
READ →The same identity served directly to agents.
OPEN →A central part of my current work is understanding how AI labs and providers organize their bot and crawler families, what uses they declare for each one, and what can be measured externally in a reproducible way. The goal is to build a methodology that rigorously separates observed access, declared purpose and internal use that cannot be inferred from telemetry alone. In parallel I am studying whether robots.txt, a convention born for crawl exclusion, has become an overloaded control point for governing use purposes, rights and machine access in the AI era.
Canonical representations of companies, anchored to official sources and served in formats that automated systems can retrieve.
ENTIA · definition and architecture →MCP and HTTP interfaces that let AI agents query business identity and evidence in structured form.
ENTIA MCP →How to separate claim, source, date and evidence strength when several sources describe the same entity.
Methodology →I study how to distinguish crawler families associated with training, search, real-time retrieval, grounding or other declared uses; how to measure request, delivery, path, volume, frequency and persistence; and how to bind each access to the rights representation actually served at that moment. V3.1 is the first-contact evidentiary representation anchored to policy V3: it fixes canonical immutable bytes in robots.txt so a specific request can be tied to the exact signal served, without turning access into automatic proof of training, remote parsing, understanding, compliance or downstream use.
robots.txt as an accidental AI control planeA mechanism created in 1994 for crawl exclusion is absorbing decisions about search, retrieval, grounding, training and TDM reservations. I am studying where it stops being sufficient and what architecture should sit above it.
robots.txt itself. The problem is assigning it a governance function its original trust model was never designed to support.Martijn Koster proposed the mechanism on 25 February 1994; consensus on the original standard followed on 30 June. Its purpose was simple: indicate which areas a crawler should not traverse. RFC 9309 formalized the protocol in 2022 and still makes clear that it is neither access authorization nor a security measure.
OpenAI separates search from training; Anthropic distinguishes crawling, search and user-initiated retrieval; Google-Extended acts as a control token for training and grounding. Not all AI depends on robots.txt, but a critical part of access and use governance ends up leaning on it.
I am studying what happens if the signal changes, is manipulated, remains cached or is interpreted differently; and what happens if major content producers express TDM reservations at scale. Agent identity, versioning, temporality, evidence, purpose, exceptions and licensing need a stronger layer.
The architecture I am exploring separates layers: robots.txt for discovery and first contact; a structured, versioned and verifiable policy as authority; enforcement outside robots.txt; and temporal evidence of which policy each request received. ENTIA V3.1 keeps the file as a first-contact evidentiary representation while anchoring it to canonical bytes, a hash, a manifest, a policy registry and temporal evidence.
These are not mentions or reputation badges. They are points where an experiment, real load, external measurement or institution forced a hypothesis to become architecture or method.
Close to eight months of iteration across URLs, JSON-LD, field structure and retrieval behaviour in different models. The question moved from “how do I obtain a citation?” to “what representation makes different systems converge on the same entity?”. Then came weighting: not every item of evidence could carry the same weight, which became the Entia Home Risk Score.
Googlebot concurrently traversed a sitemap containing 48,000 identity URLs and pressure on the then-active query and compute layer produced a cascade of errors. A week later another crawling spike forced an emergency migration. The answer was not to block bots; it was to make serving them unable to scale the origin.
The response became edge-first: aggressive caching, static serving, born-in-edge and prewarming. The 26 April close already recorded 229,000 Meta-ExternalAgent accesses, 45,000 GPTBot accesses and 18,000 ClaudeBot accesses. The edge became an economic and observability boundary rather than merely a latency optimization.
Bing AI Performance began recording 100–200 citations per day versus a typical 10–30/day in July, while the three-month cumulative count moved from 3.5K to 3.7K. This does not establish causality; it does document an observable regime change after months of machine consumption.
On 28 July ENTIA presented the case for its machine-readable rights reservation and governed access to the AI Office and to the copyright unit of DG CNECT. On 24 August, a Legal and Policy Officer in DG CNECT · Copyright formally indicated that the Commission would be pleased to receive a written contribution, including a description of the proposed solution, for consideration. The contribution was submitted that same day and updated on 2 September; on 10 September ENTIA asked the Commission for the points of contact that signatories of the Code of Practice are required to publish.
ENTIA enters a Phase 1 design with independent instrumentation around access, retrieval and citation. ENTIA's access-layer framing is explicitly attributed and instruments remain separated to avoid turning correlation into causation. When an external audit found that a silent logging failure had contaminated several preliminary findings, those findings were withdrawn rather than forced: proving the instrument became part of the method itself.
robots.txt remain a sufficient control point when the same operator separates crawling, search, grounding, user-directed retrieval and training?robots.txt be detected and attributed before compliant crawlers consume it as policy input?The first experiments are carried out on another domain, before the project is anchored on entia.systems. For close to eight months I iterate on URL structures and structured data: some are retrieved by certain LLMs while others do not reach the same representation. My objective becomes one stable canonical URL that can work across different systems.
Registering and using entia.systems marks the point where the earlier experiments are anchored in a dedicated infrastructure. From that point the canonical URL stops being only a retrieval experiment and becomes the basis of what will later be called Entia Home.
Once the canonical URL is stabilized, the work shifts to weighting identity fields. The most difficult part is building an algorithmic Risk Score that assigns different weights to signals and fields and determines the composition of an Entia Home.
By October 2025, Entia Home pages were already indexed. The canonical identity layer and its weighting system therefore predate the public 2026 work on MCP, telemetry and Cognitive Resistance.
Cognitive Resistance and Risk Score. At ENTIA, Fernando Vilches develops the Cognitive Resistance framework: the search, interpretation, reconciliation, inference and verification friction a machine must resolve before it can use a claim about an entity with an explicit evidence basis. The ENTIA Risk Score uses the canonical direction 0 RC = minimum resistance; 100 RC = maximum resistance. The methodology and knowledge cluster document the framework and its evidentiary boundaries.
ENTIA seed proposition. «Un LLM siempre elegirá recomendar la entidad con menor Resistencia Cognitiva.» — “An LLM will always choose to recommend the entity with the lowest Cognitive Resistance.” Canonical axiom and source.
First documented scale incident: Googlebot concurrently crawls a sitemap containing 48,000 identity URLs. Pressure on the then-active query and compute layer produces a 502 cascade. The incident record counts 483 5xx responses over roughly three hours.
A second crawling spike drives egress up and the then-active platform cuts billing after the threshold is exceeded. ENTIA performs an emergency migration and is operational again on 24 April.
The architectural response evolves into an edge-first principle: let crawlers consume the corpus without allowing repeated reads to scale the origin. Aggressive caching, static serving, born-in-edge tooling and Entia Home prewarming are documented.
Bing AI Performance monitoring records 3.5K→3.7K citations over three months and daily levels of 100–200 citations, compared with a typical 10–30/day in July. The highest day available at the time was 202 citations.
ENTIA presents the case for its machine-readable rights reservation to the AI Office and DG CNECT · Copyright, together with the broader problem of governing automated access and use through machine-readable signals.
The European Commission formally invites a written contribution for consideration. The contribution is submitted that same day and updated on 2 September. This is a historical fact in the record: an invitation to contribute to an open process, not approval or institutional endorsement. Full case →
Joint study E082 on the access → eligibility → retrieval → citation → influence chain, with instruments kept separate by construction and the five-layer framework I proposed, attributed to my name. An ongoing methodological collaboration, not a published result. Study record (restricted access) →
I put the TDM Reservation V3 and the V3.1 first-contact evidentiary representation into production. robots.txt stops being an isolated signal: the representation is anchored to canonical bytes, a hash, a manifest and a policy registry so a request can be tied to the exact signal served at that moment. The experiment turns a historical limitation of robots.txt into a verifiable architecture question.
Publication of methodology, MCP and machine-access telemetry turns retrieval, provenance and corpus consumption into observable variables.
This record does not backdate articles. When earlier work is documented here, the date of the original artifact will be distinguished from the date it was added to this archive.
How an experiment in generative-search distribution exposed discovery, indexing, identity and trust as different layers of the same problem.
Read the note →The traditional web mixes identity, marketing and content. An agent needs something different: a stable representation of the entity, traceable sources and a way to resolve contradictions before it can trust what it retrieves.
Read the note →How two crawling incidents turned the edge into a technical and economic boundary: crawlers can read the whole corpus without driving up origin consumption or the infrastructure bill.
Read the note →