Information Gain Is Founded on Entropy

The greater the reduction in disorder (entropy reduction), the greater the information gain.

Information Entropy − Residual Entropy = Information Gain

H(X) − H(X|Y) = I(X;Y)

An LLM is itself a high-information-entropy string generator: it knows every possible string, and any answer can come out. All the engineering that gets an LLM moving and able to accomplish anything is entropy reduction: prompt injection is entropy reduction, context design is entropy reduction, agent planning is entropy reduction, harness building is entropy reduction; only the scope of influence differs. The purpose of entropy-reduction engineering is to remove uncertainty, and extra erroneous noise does not make an LLM more precise; it only makes it more confidently wrong.

Overshoot · Build base · Break out

The breakthroughs in AI entropy engineering are infrastructure forged from repeated wall-hits.

Agent hit a wall; Context backfilled the foundation.

一個渺小、穿著樸素希臘長袍的人形機器人，手提一盞小燈，在一座龐大宏偉的石柱長廊中自信地往深處走去，廊柱朝遠方無止盡延伸。渺小的機器人是 Llama 3B 這種垃圾級小模型，手中的小燈只照亮自己腳下，是它有限的世界知識。但牡步伐自信，因為真正在引路的是周圍的石柱秩序，不是手裡的燈。柱列朝深處延伸，對應 master → category → post 的漸進式披露。模型小不要緊，秩序夠清楚的時候，每一 hop 都收斂成一道選擇題。這幅畫的主角不是機器人，是廊柱本身，能力強弱不是關鍵，結構正確才是。

A Small Language Model (SLM) Actually Understood an Entire Website

The small model Llama 3.2 3B is a language model with only 3B parameters, about as small as they come. Ask it a question and it can only answer from its 3B of training data. It does not know what your site says, does not know which articles you published, and knows nothing about the content you have built up recently. Using it to run website Q&A should have been a fantasy.

A humanoid robot in a Greek tunic climbs a pre-carved stone spiral staircase inside a grand old library, moving toward the light above, with crumpled and ignored scraps of paper scattered on the floor. The spiral staircase is WordPress's existing categories and hierarchy; the carving was already there, not cut by this traveler. The robot climbing on foot matches RAG Sitemap retrieving directly along a ready-made path. The crumpled scraps on the floor are the reverse work of vectorization, tearing organized content back into fragments and reassembling them with cosine similarity. The orderly shelves are the low-entropy sediment humans lay down article by article, category by category, while running a site. The light above is the direction of the answer: the structure itself leads the way, and the model only has to understand and choose.

Why RAG Doesn't Need a Vector Database

A vector database is not a requirement for RAG; it is only one way to feed data to an AI. When data is inherently messy and lacks clear boundaries, vectorization helps a model guess semantic relevance from large amounts of text, and that has its value. But when content already has order, the question is no longer how to force relevance out of chaos, but how to let the AI see the most important interpretive clues first. Effective RAG does not have to slice the full text, compress it into vectors, and then guess the answer; it can instead organize content into a path the AI understands layer by layer, lowering contextual uncertainty first and then expanding the detail.

A group of people circle a central light, each receiving and cupping a flame of their own, light spreading from one place into many separate palms. The central light is the cloud API of the past decade, where every inference had to come back and pay the bill. The light passed to each pair of hands corresponds to the trajectory of the NPU, chip is model, and the Chrome Prompt API, with inference moved back onto the visitor's own device. Each flame is close in size, meaning the edge small model is already capable enough to carry a site's navigation task. The posture of hands cupping a flame is privacy and non-disclosure; privacy holds naturally under this architecture. The distances between people are even: this is not a new center replacing the old one but the center dissolving entirely.

The End Goal: Moving Compute onto the User's Device

"AI-on-Chip" means that when every device has a small AI model carved into a chip, the model is no longer software that must be loaded but a compute chip always on standby. The LLM inference an application needs can run locally on the visitor's device, bringing the site owner's AI compute cost to zero. This is the end goal of RAG Chatbot.

AI Entropy-Reduction

AI Entropy-Reduction Engineering｜The LLM Is the Entropy Source; Using It Is Entropy Reduction

Why an LLM Can Be Seen as Entropy

Why Using It Is Entropy Reduction

Information Gain Is Founded on Entropy

From Prompt to Harness Engineering

The Four Leaps of Entropy-Reduction Engineering

Overshoot · Build base · Break out

Related articles

A Small Language Model (SLM) Actually Understood an Entire Website

Why RAG Doesn't Need a Vector Database

The End Goal: Moving Compute onto the User's Device