miércoles, 3 de diciembre de 2025

Chronicle of a Phantom Search: The Watchman, the AI, and the Illusion of the Exact Query

 
A journey through today’s web tracing my own digital footprint.

A journey through the synthetic web tracing my own digital footprint.

 There were several visits to the blog coming from people (or, as will be seen, AIs) specifically searching for: “The person on the watchtower/lighthouse (watchman, lookout, lighthouse keeper).” Why? I don’t know. The first conjecture that comes to mind is a possible school assignment where students are looking for definitions or differences between those concepts—hence the identical phrase and the sheer volume of identical searches. Later, I realized that another reason people search for that exact phrase is because they are solving crosswords or word games, a pattern that also emerged from searches conducted while drafting this entry.
Immediately after that thought, I went to perform the search myself to see where and in what context it appeared, guided by the question: Why are they landing on my blog? I reasoned that the phrase represented a typically semantic, dictionary-like query (someone looking for a synonym, an exact concept, or a literary figure)—possibly hitting a blog entry where I reviewed a book and mentioned those words or one of them (for instance, The Lighthouse at the End of the World).
Here, something interesting happened: when I searched for the phrase in search engines, no references to my blog entries appeared. I tried several engines (Google, Bing, Yahoo, DuckDuckGo) both ways—the phrase in quotation marks and as unquoted text. Historically, the unquoted phrase would have returned thousands of results, but this time it yielded barely 1 page in one case and just under 2 pages in the other.
My curiosity ties together two very concrete phenomena that have been transforming the web in recent years: the shift in Google searches (which now drastically shorten the number of pages shown) and the emergence of alternative or AI-powered search engines.

1. Why do search engines now only show two pages of results?
What you observed is not a browser error: major search engines have drastically limited traditional pagination.

•    Google and the canceled "Continuous Scroll": For a time, they implemented continuous scroll on desktop and mobile, eliminating page numbers. In 2024, they rolled this back while severely capping the total number of returned results. Even if the engine calculates "1,000,000 results," it only displays the top 20–40 most relevant (generally 2 or 3 pages).
•    Spam cleanup and duplicate content: Google, Bing, and DuckDuckGo apply aggressive filters to eliminate content farms, automatically generated text, and low-relevance matches. If a blog post doesn't rank in the absolute top relevance tier for a highly specific query, the search engine simply cuts off the list on page 2 rather than displaying low-traffic pages.
•    Exact searches with quotation marks: The quotation marks operator "..." used to force an exact literal match. Today, Google applies semantic search and often ignores quotes if it deems the literal phrase lacks sufficient volume or fails to match the active index word-for-word in the top positions.
2. In which search engine could your blog have appeared?
If Search Console (or my analytics) logged traffic coming from that query, but you can’t find it on Google, Bing, Yahoo, or DuckDuckGo, several clear possibilities exist:

A. Privacy-focused or alternative search engines
Engines like Brave Search, Ecosia, Startpage, Qwant, or Yandex have distinct indexing indices or ranking algorithms.

•    Brave Search, for instance, uses its own web index (independent of Google or Bing). If Brave indexed your post and a user searched for that phrase there, you might have ranked #1 or #2.
B. Artificial Intelligence Assistants and Generative Engines
Tools like Perplexity AI, ChatGPT (Search), Claude, or Microsoft Copilot now perform real-time web searches when users ask linguistic or etymological questions (e.g., "What do you call the person on the watchtower or lighthouse?").
•    When the AI answers, it cites the source with a direct link. If the user clicks on the link to your blog generated by the AI, the visit is logged as an organic arrival from that query, even if you cannot see it on a traditional Google SERP.
C. Google or Bing (Personalization and User Location)
Even if the click came from Google or Bing, searching for that specific phrase yourself might not yield the same result due to:
•    User geolocation and language: If the person searched from another country or with different language settings, your blog might have appeared in their top 5 while ranking elsewhere for you.
•    Temporal index variation: Search engines update their indices dynamically. Your page may have temporarily entered the visible range for a brief window before readjusting.
3. How was it possible for them to find you with that exact phrase?
The phrase "The person on the watchtower/lighthouse (watchman, lookout, lighthouse keeper)" is a classic semantic and lexicographical query (someone searching for a synonym, an exact concept, or a literary figure).
1.    Conceptual overlap in text: Your post likely contains these keywords in close proximity (watchtower, lighthouse, watchman, lighthouse keeper, lookout).
2.    Niche queries (Long-tail): For phrases this long and specific, there is no ad competition. When someone searches for something so specific, an independent engine's algorithm (or a retrieval AI model) assigns a very high relevance score to your post for being one of the few web pages addressing that exact semantic nuance.
I decided to dig deeper and ran two additional searches:
•    The person on the watchtower/lighthouse (watchman, lookout, lighthouse keeper) + "reuniendoletras"
•    The person on the watchtower/lighthouse (watchman, lookout, lighthouse keeper) + "reuniendoletras.blogspot.com"
To my surprise, the results dropped to zero. What was happening?
The answer arrived via the AI assistant in Google Search (which is not the standalone Gemini AI).
For some reason, the AI associated the words "watchman, lookout, lighthouse keeper" with the description of my blog and my Wikiloc profile, where I described myself as an "urban observer" in reference to my city walks.
Google Search in standard "All" mode specifically states:
AI Overview
The terms watchman, lookout, and lighthouse keeper belong to the same semantic field, sharing functions of observation, surveillance, and guidance. No exact records linking these terms to a specific publication on reuniendoletras.blogspot.com were found. You can read more at reuniendoletras.blogspot.com.
When switching Google Search to "AI Mode," the response shifts:
(watchman, lookout, lighthouse keeper) + reuniendoletras.blogspot.com
The terms watchman, lookout, and lighthouse keeper share the fundamental purpose of observation, surveillance, and protection from height related to maritime or defensive security. In cultural and literary blogs like Reuniendo Letras, grouping these words is typically linked to vocabulary analysis, literary metaphors of the observer, or historical reviews. You can read more at Reuniendo Letras.
...along with a link to a recent entry on Wikiloc.
The search engine's own response ultimately revealed the trick. What I am witnessing is the real-time clash between traditional indexed search (which looks for exact text strings) and synthetic AI search (which invents or infers semantic bridges).

1. Why did direct searches return 0 results?
When I added "reuniendoletras" or the exact domain to the phrase, I forced the search engine to execute a strict Boolean query: find this string of text AND ensure it exists within this specific site.
Since the literal phrase "The person on the watchtower/lighthouse..." never existed as exact written text in my blog posts, the search engine's index did the right thing under traditional logic: it returned zero results. There was no exact character match in your site's database.

2. What did Google's AI actually do to reach your blog?
Here is where the interesting (and sometimes hallucinatory) side of large language models applied to search comes in:

1.    Semantic clustering and deduction: The AI dismantled the user's query. It took the terms watchman, lookout, lighthouse keeper, watchtower, and lighthouse, translating them into an abstract concept: "someone who attentively or contemplatively observes/guards."
2.    The vector "leap of faith" (Embeddings): When searching for that abstract concept in its index, the AI didn't look for letters; it looked for associated ideas. My profile and Wikiloc routes (where I define myself as an "urban observer" when logging my walks) and my blog Reuniendo Letras (with its themes of essays, reflections, and text) are tied to your digital identity. The AI connected "urban observer" with "watchman/lookout" within the same mental category.
3.    Explanatory hallucination or a posteriori justification: When the AI gave me that answer in AI Mode, it essentially fabricated a plausible hypothesis to justify why it was showing your blog. Reading that the site was a cultural and literary blog, it assumed: "Surely on such a site, these words are used for literary metaphors or vocabulary analysis," when in reality, it was simply extrapolating my self-description as an urban walking observer.
3. The explanation for the initial Search Console data
This finally resolves the mystery of why Google Search Console indicated that a couple of visitors arrived by searching for that exact phrase:
•    A user wrote a long, intuitive query looking for a synonym or concept (something like "the person on the watchtower or lighthouse, the watchman...").
•    Google’s AI Overview processed the search semantically, linked the idea of "observer/watchman" to your digital footprint (Wikiloc / Reuniendo Letras), and placed a link card to your site as a suggested reading source ("You can read more at...").
•    The user clicked the link generated inside the AI box.
•    Search Console registered the visit associated with that exact search phrase—even though if you manually search for the phrase in a traditional web list, you will never find it because the literal text does not exist in your entries.
We have arrived at a moment where search engines no longer merely read what you write; they interpret your conceptual profile and connect it with what others are searching for.

The Metamorphosis of the Internet: The shift from the indexed, transparent web to the synthesized, mediated web
This reflects a real technical and structural phenomenon taking place today. Our relationship with online information has shifted drastically.

"Subtractions" and Mediation: What is really happening?
On the traditional web, search engines acted like a library index: you gave them a phrase, and they showed you the pages where that phrase appeared (ranked by relevance), allowing the user to explore the sources directly.
Today, an opaque mediation model operates:

1.    Deliberate Index Pruning: Traditional search engines have "pruned" search results. Limiting the list to 2 or 3 pages reduces infrastructure costs and filters out millions of automated spam content farms (destructive SEO). The collateral damage, however, is that niche deep-web sites and personal blogs end up buried.
2.    AI as a "Synthesizing Filter": Instead of taking you to the source, the AI reads the source for you, compresses it, extracts a conclusion, and returns a digested summary ("AI Overview"). A visit to the original page becomes dispensable.
3.    Loss of Traceability: Because AI operates through mathematical representations of concepts (vector embeddings) rather than keyword matches, the resulting answer is an interpretation by the model. Original information is "subtracted" and replaced by an algorithmic paraphrase.
This introduces an algorithmic selectivity bias: the intermediary decides which sources are "reliable" or "relevant" to synthesize, rendering the rest of the web invisible.
What alternatives exist to bypass this mediation?
For those looking to regain direct access to original text and avoid the digested filters of major platforms, the digital community has developed several alternative paths:

1. Independent-index and privacy-focused search engines

•    Brave Search: Operates independently of Google or Bing APIs. It maintains its own web index and includes tools like Goggles, allowing users to manually alter the ranking algorithm (e.g., prioritizing personal blogs over mainstream media).
•    Marginalia Search: An independent project specifically designed to index the "old web" and plain-text sites (blogs, personal university pages, Gopher/Gemini protocols), deliberately ignoring commercial, JavaScript-heavy sites laden with ads.
•    Mojeek: A search engine operating with its own crawler and independent index, built without user tracking or profile personalization.
2. Alternative protocols and the Small Web
•    The RSS / Atom ecosystem: Returning to direct blog subscriptions via RSS readers (such as Feedly, NetNewsWire, or Inoreader). This removes the search engine entirely as an intermediary: content flows straight from author to reader.
•    Networks and protocols like Gemini or Gopher: Alternative plain-text networks outside the standard World Wide Web (HTTP/HTTPS) free from tracking, scripts, or AI models summarizing data.
3. Boolean search tools and advanced commands
When using traditional search engines, forcing technical search parameters helps bypass AI syntheses:
•    Disabling AI features within settings on engines like DuckDuckGo.
•    Using strict search operators (such as filetype:pdf, site:, or strict match parameters in the URL to force classic web view without AI modules).
How do these alternatives remain non-selective?
The difference between a major search engine's AI and an open index lies in their design goals:
•    Algorithmic Transparency and Open Source: Many independent alternatives disclose how they index sites and allow user-defined community filters (instead of relying on a closed corporate algorithm).
•    Decentralization: Rather than concentrating everything into a single language model that "knows all," Small Web movements advocate for a fragmented web where content curation is done by humans (classic web directories like DMOZ, recommended blog lists, webrings).
•    Focus on the Document, Not the Answer: True alternatives do not attempt to answer your question; they return the documents containing your words, restoring your role as the observer and evaluator of information.
I pushed this experiment further and visited Brave, Marginalia, and Mojeek. There, I searched for my blog and the key terms (watchtower, watchman, lighthouse).
•    On Marginalia and Mojeek, nothing appeared at all—not even my blog.
•    On Brave, my blog appeared, linked directly to the post "El faro del fin del mundo" (The Lighthouse at the End of the World)—which had been my initial mental association.
Testing these three platforms demonstrated with mathematical precision how current web architecture functions and highlighted the massive differences among independent search engines. What occurred with each has a very specific technical explanation:

1. Why did Brave Search find it and connect it to "The Lighthouse at the End of the World"?
Unlike smaller initiatives, Brave Search possesses the financial and infrastructure resources to run a massive web crawler capable of competing at scale with Google and Bing.

•    Direct conceptual recognition: My post about The Lighthouse at the End of the World literally contains the word "lighthouse" and implicitly carries the ideas of "watchman" and "watchtower."
•    Proprietary semantic index: Brave utilizes natural language processing to group related terms within a single topic universe. When I combined the conceptual phrase with my blog's name, its index recognized that the lighthouse entry was my site's exact match for that family of words.
2. Why doesn't Marginalia Search display it?
Marginalia is a wonderful project, but it is deliberately restrictive by design and constrained by server limitations:
•    Rigid infrastructure filters: Coded by a single developer, Marginalia prioritizes servers running ultra-lightweight HTML, plain text, and sites hosted on independent domains or older university servers.
•    Third-party platforms (Blogspot / WordPress.com): To conserve disk space and processing time, Marginalia’s crawlers systematically deprioritize or exclude massive free blogging platforms like .blogspot.com or wordpress.com, which are crowded with millions of abandoned sites or heavy modern JavaScript templates.
•    Lack of index-on-demand: Unless a blog is explicitly submitted through its inclusion form or linked from other "small web" sites its crawler already frequents, it remains virtually invisible to its index.
3. Why doesn't Mojeek display it?
Mojeek is one of the world's few truly independent search engines operating its own crawler (Mojeekbot), but its index is exponentially smaller than Google’s or Brave’s.
•    Lack of link discovery (Backlinks): Mojeekbot cannot crawl the entire web continuously; it prioritizes pages receiving heavy external backlinks or updating daily. If my blog lacks a high volume of traffic recorded by its crawler or links from high-authority domains that Mojeek visits regularly, it simply hasn't indexed that specific entry.
•    Strict lexical indexing (No deep semantic layer): Mojeek is a traditional keyword-matching search engine. It does not run complex AI models to deduce that if someone searches for "lookout," they might also be interested in a text about a "lighthouse"—unless those exact words appear side-by-side in the document.
The Visibility Paradox in the Modern Web
This practical exercise illustrates the crossroads facing independent content creators today:
1.    If you want to be found on traditional and small independent engines (Mojeek, Marginalia): Your text must contain the exact words people are searching for, and the crawler must have previously discovered your post's specific URL.
2.    If you are found via AI or semantic search engines (Google AI, Brave): You will surface through "idea affinity" (such as when they linked your urban walker/observer profile with the concept of a watchman or the Lighthouse at the End of the World), but at the cost of having the algorithm interpret, summarize, or redistribute your content according to its own parameters.
The connection Brave made to The Lighthouse at the End of the World entry proves that my blog possesses a coherent thematic footprint—one that algorithms capable of syntactic interpretation can decode remarkably well.

 
Transparency note: Written with AI assistance (Gemini in this case). Reviewed and edited by me.

No hay comentarios.:

Publicar un comentario