RAG: Teaching AI to Look Things Up Before It Speaks
From TF–IDF and cosine similarity to biological knowledge retrieval — and why this could become surprisingly important for drug discovery
You know that colleague who seems to know everything? You ask them about a gene, a pathway, or some obscure paper from 2017, and somehow they have an answer.
Now imagine that colleague has read millions of papers. Sounds great — except there is one problem. Sometimes they confidently remember something that isn’t actually true.
That, in an admittedly simplified analogy, is one of the problems we have with large language models (LLMs). LLMs can encode an extraordinary amount of information in their parameters. But their knowledge is imperfect, frozen at particular points in training, and they can generate statements that sound completely plausible while being factually wrong. They also cannot magically know the contents of your unpublished RNA-seq dataset, your internal compound-screening reports, or the paper that appeared yesterday.
So what if, instead of asking the model to remember everything, we let it look things up first? That is the basic idea behind Retrieval-Augmented Generation, or RAG. And the interesting part is that underneath one of the most talked-about ideas in modern AI sits some wonderfully classical mathematics.
So, what exactly is RAG?
Think of an LLM as a scientist sitting an exam. Without RAG, it is essentially a closed-book exam. You ask:
“What evidence links microglial activation to synaptic dysfunction in Alzheimer’s disease?”
The model has to answer largely from patterns encoded during training. With RAG, we turn it into an open-book exam. Before answering, the system searches a collection — perhaps PubMed abstracts, internal reports, experimental protocols, clinical-trial documents, or your own laboratory data — retrieves passages that appear relevant, and places them in front of the model. The LLM then generates its answer using both your question and those retrieved passages.
RAG has two stages: retrieval, which identifies relevant passages, followed by generation, where the LLM answers while conditioned on those passages. Simple idea. Big consequences.
But RAG didn’t suddenly appear with ChatGPT
Humans have been trying to make computers retrieve answers for almost as long as computers have existed. In 1961, the BASEBALL question-answering system could answer questions about baseball statistics stored in a structured database. A few years later, systems were already retrieving candidate sentences based partly on weighted overlap between the words in a question and words in possible answers.
Then information retrieval became its own major field. Search engines essentially asked:
Given millions of documents and a few words typed by a user, which documents should appear first?
That question gave us ideas such as the vector-space model, TF–IDF, inverted indexes, BM25 and eventually dense neural retrieval. By the 1990s, TREC provided standardized evaluation of information-retrieval systems. Neural retrieval and machine reading followed, and by the late 2010s retrieval was increasingly being combined with neural language models. The term Retrieval-Augmented Generation became widely established with the 2020 RAG work of Lewis and colleagues.
So RAG isn’t really one completely new trick. It is more like a marriage between two mature technologies: information retrieval + generative language models. And to understand the first half of that marriage, we need to talk about TF–IDF.
TF–IDF sounds horrible. It really isn’t.
Imagine I give you 10,000 neuroscience papers. Now you search: “HTT CAG repeat expansion”. Which words should matter most?
- The word “the” might occur hundreds of thousands of times. Completely useless.
- “Neuron” might occur in thousands of papers. Useful, but not very specific.
- “HTT” may occur in a much smaller subset.
- And the combination of HTT + CAG + repeat expansion suddenly becomes extremely informative.
TF–IDF is essentially a mathematical way of capturing that intuition. It asks two questions: How much does this term matter inside this document? and How unusual is this term across all my documents?
Part 1: Term Frequency - TF
Suppose one paper mentions HTT once and another mentions it 20 times. Intuitively, HTT is probably more central to the second paper. That is term frequency. A simple version would just count occurrences, but retrieval systems often damp repeated counts logarithmically, for terms that occur in the document:
Why use a logarithm? Because mentioning HTT 100 times doesn’t make a paper 100 times more relevant than mentioning it once.
- 1 occurrence → TF = 1
- 10 occurrences → TF = 2
- 100 occurrences → TF = 3
So repetition matters — just not infinitely.
Part 2: IDF — rewarding the unusual words
Now comes the clever bit. Imagine HTT occurs in only 200 of your 100,000 papers. That makes it informative. But suppose the word “cell” occurs in 90,000. Not nearly as discriminating. Inverse Document Frequency captures this:
where \(N\) is the total number of documents and \(df(t)\) is the number of documents containing term \(t\). Rare terms receive larger weights; common terms receive smaller ones. And if a term appears in every document, then:
In other words: if everybody is saying it, it isn’t very useful for figuring out which document you want. That is the heart of IDF. Finally:
That’s it. A term becomes important when it is frequent enough locally but distinctive globally.
Okay — but how does a computer retrieve a paper?
This is my favourite part, because retrieval suddenly becomes geometry. Imagine three tiny, artificial “papers”:
- Paper A: HTT CAG expansion
- Paper B: HTT expression
- Paper C: APP amyloid
And we search: Query: HTT CAG. TF–IDF converts each document into a vector. Instead of thinking of Paper A as a sentence, imagine it as a vector over the vocabulary
with a TF–IDF weight sitting in each position. Most positions are zero. Paper A points strongly toward HTT + CAG + expansion. Paper C points toward APP + amyloid. Our query points toward HTT + CAG. Now imagine those vectors as arrows in space. The retrieval system asks:
Which document-arrow points in roughly the same direction as my query-arrow?
That is where cosine similarity enters:
You don’t really need to fear the equation. The numerator measures how much the two vectors overlap. The denominator normalizes their lengths. And the result tells us how closely their directions align. Recall that the dot product of two vectors can be written in two equivalent ways:
If two vectors point in similar directions, their cosine similarity is high. (If the angle between them is obtuse, the dot product is negative: for example, \(3(-1)+0.5(3)=-1.5\).) So our HTT CAG query would naturally sit much closer to “HTT CAG expansion” than to “APP amyloid”. TF–IDF document and query vectors are compared using cosine similarity, and documents with higher similarity are ranked higher. And suddenly a search engine starts to look less mysterious.
Sparse retrieval: the librarian who remembers exact words
TF–IDF is what we call sparse retrieval. Why sparse? Because imagine a vocabulary containing 100,000 words. A particular paper may contain only a few thousand of them. So its vector looks roughly like:
Mostly zeros. Sparse retrieval is extremely useful when exact terminology matters, and biology is absolutely full of exact terminology: TP53. BRCA1. rs429358. IL6. EGFR T790M. CDKN2A. If I search for EGFR T790M, I probably want a system to care enormously about those exact characters.
But sparse retrieval has an annoying weakness. Imagine I search:
“brain immune cells involved in Alzheimer’s disease”
while the paper says “microglial activation in AD”. A human immediately understands the connection. Pure keyword matching may struggle because the words don’t overlap sufficiently. This is called the vocabulary mismatch problem.
Dense retrieval: searching by meaning instead of spelling
Dense retrieval takes a very different approach. Instead of representing a document by thousands of mostly-zero word coordinates, a neural encoder converts the whole passage into a dense embedding. Conceptually:
and:
The numbers themselves don’t correspond neatly to individual words. Instead, the model has learned a representation in which semantically related texts can occupy nearby regions of embedding space. Now the system can potentially retrieve something relevant even when the wording is different. In a common bi-encoder architecture:
and relevance can be scored using:
The useful engineering trick is that document embeddings can be calculated before anyone asks a question and stored. When a query arrives, only the query needs to be embedded and compared with the stored vectors. So I like thinking about it this way:
- Sparse search asks: “Did you use the words I’m looking for?”
- Dense search asks: “Are you talking about the thing I’m looking for?”
Neither is perfect. Which is exactly why combining them can be powerful.
Hybrid retrieval: because biology needs both
Suppose I ask:
“What evidence links APOE4 to microglial lipid dysregulation in Alzheimer’s disease?”
I care about exact entities: APOE4. But I also care about broader concepts: lipid metabolism, lipid droplets, microglial states, cholesterol handling, neuroinflammation. Sparse retrieval is excellent at protecting the exact biology. Dense retrieval can help find semantically related passages that use different language. A practical system can therefore do something like:
The general strategy is that inexpensive retrieval generates candidates first, followed by a more computationally expensive neural model that reranks only the top results. For scientific applications, that balance matters. We don’t want semantic cleverness to make the system forget that APOE, APOE4, and some vaguely related lipid gene are not interchangeable biological entities.
A tiny DNA example makes this even more intuitive
Let’s leave papers for a moment. Suppose we have CAGCAGCAG. If we slide a three-base window across it:
CAG occurs three times. So, conceptually, we can treat DNA k-mers a little like words. A sequence becomes a collection of short motifs. We can count them, weight them, and compare sequence representations. This analogy is useful because biology is full of “languages”:
- DNA has nucleotides.
- Proteins have amino acids.
- Single-cell data contains patterns of gene expression.
- Chemical structures have molecular substructures.
- Scientific papers contain natural language.
Of course, the analogy has limits. TF–IDF over k-mers is not sequence alignment, does not preserve all positional information, and certainly does not prove biological function. But the underlying computational idea is beautiful:
Turn complicated objects into representations that allow useful similarity to be measured.
That idea connects classical information retrieval to much of modern machine learning.
But how do we know whether a RAG system is actually good?
This part matters enormously. A RAG system can produce an incredibly convincing answer from completely irrelevant evidence. So evaluating only the final prose is not enough. We need to evaluate retrieval and generation separately. For retrieval, two old metrics remain wonderfully useful.
Precision
Of everything I retrieved, how much was actually relevant?
Recall
Of all the relevant information that existed, how much did I actually find?
These definitions extend to ranked retrieval, where we also care about whether relevant documents appear near the top. For a scientific RAG system, though, I would go further. I’d ask:
- Did it retrieve the right paper?
- Did it retrieve the right passage from that paper?
- Did the answer actually use that passage correctly?
- Does every major scientific claim have evidence?
- Did the model exaggerate what the experiment demonstrated?
That last question is especially important. Imagine a paper reports “Gene X was associated with neuronal degeneration.” And the RAG system writes “Gene X causes neuronal degeneration.” Perfect retrieval. Bad science. RAG doesn’t eliminate scientific reasoning errors. It simply gives us a much better chance of seeing where the answer came from.
This is where RAG gets really interesting for computational biology
Biology doesn’t exactly suffer from a lack of data. We have the opposite problem. Papers. Preprints. Clinical trials. RNA-seq. Single-cell atlases. Proteomics. CRISPR screens. Imaging. Protein structures. Perturbation datasets. Drug-response databases. The difficult part is increasingly connecting them.
Imagine finishing a single-cell RNA-seq experiment and discovering a disease-associated neuronal population with Gene A ↑, Gene B ↑, Gene C ↓. Instead of manually spending days searching papers and databases, you could ask:
“What mechanisms could explain this transcriptional state?”
A well-designed scientific RAG system might retrieve evidence connecting those genes to pathways, perturbation studies, disease models, compounds and previous experiments. But importantly, it could return the evidence with the hypothesis. That changes the interaction from “AI says Gene X matters” to “AI suggests Gene X may matter — here are the experiments and papers that led it there.” For science, that distinction is huge.
And drug discovery might be one of the most interesting use cases
Imagine a researcher investigating a disease-associated cell state. The question isn’t simply “Which drugs treat this disease?” It might be:
“Which perturbations could reverse this molecular signature, through mechanisms supported by evidence in this cell type?”
Now your retrieval layer could potentially connect:
The LLM’s job isn’t to magically discover the drug. Its job becomes closer to a scientific navigator across heterogeneous evidence. That could help researchers identify connections they might otherwise miss, surface contradictory findings, and generate hypotheses worth testing. And that’s an important distinction:
A few reality checks
This technology is exciting, but there are some traps worth keeping firmly in view.
- A RAG system is only as useful as what it can retrieve.
- If your database contains bad science, RAG can retrieve bad science.
- If the relevant paper isn’t indexed, the model cannot retrieve it.
- If your chunks split an important experimental result away from its context, retrieval can become misleading.
- If the embedding model doesn’t understand biological terminology well, semantically important evidence may never appear.
And perhaps most importantly:
A paper can be relevant and still be wrong. Two papers can be relevant and contradict one another. A mouse experiment can be relevant but not establish what happens in humans. A tumour-cell-line result may not generalize to primary tissue. Those distinctions are exactly where scientific RAG needs to become more sophisticated than ordinary document search.