---
title: "RAG Complexity Is a Bet Against the Model"
published: 2026-09-30T08:00:17.000-04:00
updated: 2026-09-30T08:00:17.000-04:00
excerpt: "One Postgres table with two indexes beat every published RAG pipeline on MuSiQue. Every architectural addition is a bet the model stays weak."
tags: AI, RAG
authors: Matvey Arye
---

> **TimescaleDB is now Tiger Data.**

A single Postgres table with two indexes outscored every published RAG pipeline on MuSiQue and ranked third on BRIGHT against systems with fine-tuned retrievers. While the numbers are great, they're really just to prove that more is less when it comes to these systems.

Every architectural addition in a RAG pipeline is a bet that the model will keep needing that help. A planner because today's model can't decompose multi-hop questions. A knowledge graph because it confuses entities. A reranker because the top-10 ordering is wrong. Each solves a real failure mode against a specific model. Each becomes maintenance debt the moment that model improves, and improvement now arrives every few months.

That bet gets worse every model generation. Query decomposition, entity disambiguation, knowing when to search again, reformulating after a missed search. Work that used to need scaffolding now moves back into the model. The weakness vanishes in a release. The complexity built around it takes months to ship and years to maintain. A long-lived tax for a short-lived problem.

**A thin stack rides the model frontier; a complex pipeline has to be rebuilt to keep up.**

Headlines on the two hardest open RAG benchmarks, run on the same minimal stack:

| Benchmark | Our result | Reference |
| --- | --- | --- |
| MuSiQue (500q) | 0.418 EM / 0.564 Acc | PAR-RAG: 0.33 EM / 0.43 Acc |
| BRIGHT (12 domains, mean nDCG@10) | 0.556 | 2nd–3rd rank tier on the public leaderboard (verified September 2026); only result there without a fine-tuned retriever |

On MuSiQue, the loop reverted every "improvement" we tried against the baseline. Five consecutive regressions. On BRIGHT, the same loop kept two cheap, removable additions and rejected the rest. The loop can't predict which additions will age well across model generations. It can reject the ones that don't help today, which is most of them.

## The stack fits on two lines

One Postgres table per corpus, with BM25 and [HNSW](https://arxiv.org/abs/1603.09320) indexes. An [MCP](https://modelcontextprotocol.io/) tool server giving Claude direct access to hybrid search.

That's it. No knowledge graphs. No hierarchical indexing. No retrieval planners. No fine-tuned retrievers.

Both benchmarks ran on [Tiger Cloud](https://www.tigerdata.com/cloud) (built by our team). Two features of the platform were important for the methodology. [pg\_textsearch](https://github.com/timescale/pg_textsearch) gives us [native BM25 indexes](https://www.tigerdata.com/blog/introducing-pg_textsearch-true-bm25-ranking-hybrid-retrieval-postgres) alongside HNSW, so hybrid retrieval is two indexes on one table instead of two systems. And [Fluid Storage](https://www.tigerdata.com/blog/fluid-storage-forkable-ephemeral-durable-infrastructure-age-of-agents) forks a database in seconds at any corpus size, which turns out to be a precondition for the methodology rather than a convenience. More on that below.

The entire schema is a single table:

```SQL
CREATE TABLE corpus (
  id         uuid PRIMARY KEY,
  content    text NOT NULL,
  embedding  halfvec(1536),
  created_at timestamptz NOT NULL
               DEFAULT now()
);
```

Two indexes: HNSW for semantic similarity, BM25 for keyword matching.

The MCP tool server gives Claude one search tool with three modes. Semantic for natural-language queries, embedded with text-embedding-3-small and matched via HNSW cosine similarity. Fulltext for keyword matching with BM25. Grep for case-insensitive regex when entity matching needs to be literal.

When the model passes semantic and fulltext together (it does ~93% of the time on MuSiQue), results fuse with [Reciprocal Rank Fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf). No reranker, no second-pass model, no query planner. Claude iteratively searches via tool calls. The model decides what to search for, which modes to use, how many results to request, and when to stop. There is no pre-programmed retrieval strategy.

For MuSiQue we used Haiku throughout. For BRIGHT we used Opus on most domains because the reasoning-intensive corpora benefited from a larger model. Same MCP tool, same Postgres schema, same retrieval logic.

## Staying thin needs a loop that says no

Staying thin sounds easy and is hard in practice. Every failure case in the eval looks like an argument for adding something. A planner. A graph. A reranker. Some of those additions help. Most don't. You can't tell which without testing, and the default outcome of not testing is that you ship all of them.

The autoresearch loop is the enforcement mechanism, and most of it is ordinary eval hygiene, loosely inspired by Karpathy's autoresearch framing. Start from failure analysis. Name the pattern before proposing a change. State the hypothesis. Implement the smallest change that isolates one variable. Log everything, including the reverts, so future sessions don't re-run dead ideas. Two commitments inside it are less obvious, and they mattered far more than iteration speed.

**Paired statistics on per-query deltas, not aggregate scores.** We run paired t-tests and sign tests on the same queries before and after. A mean nDCG@10 that moves 0.01 tells you almost nothing about whether a change helped. When the two tests disagree, the disagreement is the finding: it usually means the change helped a handful of queries a lot and hurt many slightly, which is a different decision than a uniform small win.

**Three attempts per hypothesis, then it's dead.** If three variants of an idea all fail, the hypothesis is declared non-viable and nobody revisits it. Without a stopping rule, a plausible-sounding idea absorbs unbounded effort, and at some point the sunk cost starts arguing for shipping a version of it.

### Cheap reverts are what make thin stacks defensible

A revert-heavy methodology has a hard prerequisite: reverting has to cost nothing. Most of these experiments mutated database state. Re-embedding 139,416 paragraphs with entity-enriched content. Adding a sketch column across 62k robotics docs. Changing an index type. If the path back from a failed experiment is a dump-and-restore, one of two things happens. You don't run the experiment, or you run it, watch it regress, and then find a reason to keep it anyway.

So every DB-mutating change ran on a fresh fork. Wins got promoted. Regressions got paused, and DATABASE\_URL pointed back at the previous service. Nothing to unwind by hand, no half-migrated corpus to reason about. Fluid Storage forks by copy-on-write, so fork time is seconds to minutes regardless of corpus size, instead of the many hours a dump-and-restore takes at this scale.

That number is the reason the loop produced ten reverts instead of ten reluctant keeps. The thesis says delete what doesn't earn its place. You will only actually delete things when deleting is cheaper than defending them.

## On MuSiQue, every improvement regressed

[MuSiQue](https://github.com/StonyBrookNLP/musique) is a multi-hop reading comprehension benchmark with 2,417 dev questions spanning 2, 3, and 4-hop reasoning chains. Each question is constructed by composing single-hop questions where each hop genuinely depends on the previous. The corpus is 139,416 Wikipedia paragraphs.

A representative 4-hop:

"What is the capital of the county that shares a border with the county that contains the birthplace of Erik Jensen?"

Chain: Erik Jensen → born in Appleton → Outagamie County → borders Brown County → capital is Green Bay

We evaluated on 500 random questions matching [PAR-RAG](https://arxiv.org/abs/2504.16787)'s methodology for an apples-to-apples comparison:

| Hops | F1 | EM | Accuracy | Recall | n |
| --- | --- | --- | --- | --- | --- |
| 2-hop | 0.617 | 0.485 | 0.636 | 0.835 | 239 |
| 3-hop | 0.539 | 0.400 | 0.558 | 0.848 | 165 |
| 4-hop | 0.351 | 0.281 | 0.396 | 0.703 | 96 |
| Overall | 0.540 | 0.418 | 0.564 | 0.814 | 500 |

PAR-RAG (Table 3, revised January 2026) benchmarks several RAG approaches on MuSiQue using Qwen-Plus:

| System | EM | Acc | Notes |
| --- | --- | --- | --- |
| Standard RAG | 0.08 | 0.08 | Baseline |
| RAPTOR | 0.06 | 0.12 | Hierarchical retrieval |
| IRCoT | 0.31 | 0.35 | Iterative retrieval |
| HippoRAG w/ IRCoT | 0.30 | 0.42 | + knowledge graph |
| ReAct | 0.15 | 0.36 | Agent-based reasoning |
| Self-Ask | 0.13 | 0.24 | Iterative decomposition |
| PAR-RAG | 0.33 | 0.43 | Plan-driven decomposition |
| Ours (Postgres + Haiku) | 0.418 | 0.564 | Single table, hybrid search, MCP tools |

A dramatically simpler architecture outscores every system in the table. Each row was built against an older model's limitations. A stronger model on a thin stack captures most of what those pipelines were designed to provide, at a fraction of the maintenance footprint.

Every architectural change we tried against it regressed. Five reverts in a row:

| Experiment | Impact on F1 |
| --- | --- |
| Auto-hybrid search (force both BM25+semantic) | -0.127 |
| Increase results per search (10→20) | -0.092 |
| Add search hints to tool descriptions | -0.142 |
| Entity-enriched content in embeddings | -0.045 |
| Sub-question decomposition prompts | -0.019 |

This is the thesis in miniature. Each one looked reasonable. Each one regressed. Without the loop's discipline we would have shipped every one and locked ourselves into pipeline complexity to maintain. The most valuable output on MuSiQue isn't a kept change. It's five reverted ones.

A note on the accuracy number: the 0.564 headline is depressed by errors in MuSiQue's gold answers, mostly entity-name collisions ("Cleveland, Ohio" resolves to Cleveland, North Carolina in the expected chain). A failure audit walked from the start of the eval to assemble a 100-question gold-verified sample (excluding 6 questions where we judged gold to contain errors). Accuracy on this sample is 0.600. We report 0.564 as the headline because it's the apples-to-apples comparison.

Retrieval on MuSiQue is mostly handled. The model uses semantic and fulltext together on 93% of queries, adjusts candidate limits when initial searches come back too narrow, and falls back to grep for exact entity matching. Retrieval recall sits above 80% at every hop depth (0.92 / 0.84 / 0.82). Accuracy falls more steeply (0.74 / 0.54 / 0.47). Even on 4-hop, the system finds the supporting paragraphs most of the time. It just can't always compose the answer. The remaining gap is a reasoning problem.

What happens when retrieval itself is the bottleneck?

## On BRIGHT, the additions that survived were the removable kind

Sometimes retrieval really is the wall, and no model on the other side of the tool call can reason its way past a document it never saw. That is BRIGHT. It is also where the thin-stack thesis gets its sharpest test, because the honest answer is that we added things. Two classes of them. Both are text you can delete.

[BRIGHT](https://arxiv.org/abs/2407.12883) is designed specifically to break retrieval systems that rely on surface overlap. Twelve separate corpora, each with its own queries, each scored by nDCG@10. The queries come from real venues like Stack Exchange, AoPS, Reddit, and GitHub. The gold-labeling rule is "what an expert answer would actually cite," which produces a very specific failure mode. The query uses everyday or stuck-user phrasing while the gold document uses formal, canonical vocabulary.

| Domain | Query phrasing | Gold document |
| --- | --- | --- |
| economics | "Samsung's contribution to South Korea's GDP" | ASC 606 revenue recognition accounting standard |
| biology | "Why do I only breathe out of one nostril?" | Wikipedia "Nasal cycle" |
| stackoverflow | "Is there a melt command in Snowflake?" | Snowflake SQL UNPIVOT reference |
| robotics | "Subscriber in hardware interface" | ros2_control TopicBasedSystem API reference |
| aops | "Mary baking 10 cookies of 3 shapes, distribute diversely" | ProofWiki "Pigeonhole Principle" theorem |

Look at the economics row. "Samsung's contribution to South Korea's GDP" and ASC 606 share no surface vocabulary, so a general-purpose embedding model puts them in different neighborhoods. You have to know that GDP contribution is value-added derived from revenue, and that revenue recognition is governed by ASC 606, before the standard becomes a plausible citation.

The minimal baseline gets ~0.45 mean nDCG@10 here, capped on the retrieval side. Even with raw queries through BM25 plus semantic plus RRF at top-100, retrieval recall on robotics maxes out around 0.34. The ceiling is corpus-side and it is real.

### Per-domain prompts were the biggest lever

In any real production RAG system you're building for one corpus, and you'd naturally write a prompt tuned to what that corpus looks like. The agent benefits from knowing what kind of document it's searching across. The unusual thing in this benchmark setup isn't writing per-domain prompts. It's that BRIGHT bundles 12 unrelated corpora into one evaluation and implicitly invites a generic prompt that has to handle all of them at once.

For each domain, we looked at queries with zero retrieval recall and named the gold archetype. Five distinct shapes emerged:

| Gold archetype | Domains |
| --- | --- |
| Foundational language/runtime docs (avoid framework helpers) | pony, leetcode |
| Wikipedia article on the underlying concept | biology, psychology, earth_science, sustainable_living |
| Canonical academic source (NBER/IMF/textbook chapter) | economics |
| Official API reference (avoid tutorial or framework wrapper) | stackoverflow, robotics |
| Named formal theorem the story-wrapped problem reduces to | aops, theoremqa_theorems |

Each domain got a prompt that told Claude what gold looked like in that corpus, with concrete examples drawn from the actual gold IDs. The prompt also told Claude what to avoid searching for: framework keywords from training, story-specific surface entities, error-message phrasing. Same model, same retrieval infrastructure. The prompt was the only variable.

The wins were largest where Opus's training had been steering it toward the wrong gold documents. On pony, nDCG@10 went from 0.329 to 0.576 (+0.247, p < 0.001). On leetcode, 0.370 to 0.522 (+0.152, p < 0.001). Four more domains moved between +0.072 and +0.137, all significant at p ≤ 0.008. A universal one-size-fits-all prompt underperformed the per-domain prompts by 5 to 25 nDCG@10 points on every domain we tested, which rules out "any prompt change helps."

A methodology caveat: BRIGHT ships no train/dev/test split, and we wrote each archetype prompt by inspecting zero-recall queries on the same set we then re-scored. Strictly speaking, test-set tuning. Two reasons we think the wins are genuine retrieval improvement rather than fitting. The prompts name corpus-level archetypes ("Wikipedia article on the concept", "official API reference") rather than per-query gold. And the gains concentrate in retrieval recall, which is a corpus-level vocabulary-bridge effect rather than a per-query ranking one.

### Concept sketches broke the retrieval ceiling on robotics

On robotics, even with the right prompt and the right model, 59% of queries had at least one gold doc that never appeared in any tool-call result. The agent could not rank what it could not see, and the vocabulary gap was in the corpus, not the model.

The fix was to give every document a second, differently-worded body. We generated an 80–120-word concept sketch per doc, designed to bridge both vocabularies at once. A sketch on the image\_proc rectify documentation contains the canonical API names (rectify, image\_rect, camera\_info), the user-symptom phrasing a stuck developer would actually type ("distorted camera images", "weird warping in rviz"), and adjacent packages (cv\_bridge, image\_view).

Sketches live in their own column with their own BM25 and HNSW indexes. The MCP server detects the sketch column at startup and switches from 2-way RRF (content BM25 + content semantic) to 4-way (adding sketch BM25 + sketch semantic). A query in user-symptom vocabulary now hits the sketch indexes even when it misses the formal-vocabulary content indexes. The doc surfaces, and the agent reasons on the actual content.

Two diagnostic metrics matter alongside nDCG@10. **Retrieval recall** is the fraction of gold docs the agent saw in any tool-call result. **Ranking recall** is the fraction that made it into the final top-10. An experiment can move one without the other.

| Metric | No sketches | With sketches | Δ |
| --- | --- | --- | --- |
| nDCG@10 | 0.458 | 0.512 | +0.054 |
| Retrieval recall | 0.527 | 0.594 | +0.067 (sign test p = 0.014) |
| Ranking recall | 0.480 | 0.520 | +0.040 |

A similar sketch enrichment on aops produced +0.053 retrieval recall (p = 0.017). Both wins are retrieval-recall driven, which is exactly the bottleneck the sketches were designed to break. They cost time and tokens: we tagged the robotics corpus at ~62k docs over ~6 hours using Haiku, with the recipe designed to be resumable across rate-limit windows. That cost is one-time per corpus.

Both additions are disposable in the sense that matters for the thesis. A prompt is a string. A sketch is a column. When the next model bridges user-symptom vocabulary to formal API vocabulary on its own, you delete the column and the 4-way RRF falls back to 2-way. Nothing about the architecture has to change, because none of it was architecture.

### What the loop ruled out

The loop rejected more than it kept. On stackoverflow, three consecutive prompt experiments (grep-first for canonical API names, requiring a written answer before searching, requiring one alongside the ranked IDs) all measurably changed Claude's tool-call behavior and none moved nDCG@10 outside noise. On theoremqa\_theorems, Opus plus a specialized prompt was statistically indistinguishable from the Sonnet baseline on all three metrics (p > 0.14); math-domain gold-binding already aligns with Opus's training. And a minimal system prompt, tested on the theory that Claude Code's default was the bottleneck, slightly regressed retrieval recall (p = 0.058), which means the default prompt is providing real scaffolding.

The pattern across all of these is that prompt iteration cannot break a corpus-side ceiling. Where retrieval recall is capped by vocabulary the corpus does not contain, the only paths forward are sketches or an MCP-side change. Ten reverts across the two benchmarks, five on each, every one of which would have been a speculative addition without the discipline to kill it.

### Where this lands on the leaderboard

Best per-domain results across all 12 domains, mean nDCG@10 of 0.556:

| Domain | nDCG@10 | Retrieval recall | Ranking recall | Config |
| --- | --- | --- | --- | --- |
| biology | 0.803 | 0.825 | 0.831 | opus + Wikipedia-concept prompt |
| psychology | 0.654 | 0.731 | 0.636 | opus + Wikipedia-concept prompt |
| theoremqa_questions | 0.614 | 0.773 | 0.711 | sonnet max + math prompt |
| pony | 0.576 | 0.581 | 0.283 | opus + foundational-docs prompt |
| sustainable_living | 0.560 | 0.711 | 0.593 | opus + Wikipedia-concept prompt |
| earth_science | 0.551 | 0.639 | 0.557 | opus + Wikipedia-concept prompt |
| leetcode | 0.522 | 0.575 | 0.527 | opus + foundational-docs prompt |
| robotics | 0.512 | 0.595 | 0.527 | opus + specialized + concept sketches |
| theoremqa_theorems | 0.507 | 0.761 | 0.669 | opus + specialized |
| economics | 0.483 | 0.660 | 0.474 | opus + canonical-source prompt |
| stackoverflow | 0.476 | 0.629 | 0.554 | opus + foundational-docs prompt |
| aops | 0.369 | 0.652 | 0.446 | sonnet xhigh + concept sketches |
| Mean | 0.556 | 0.678 | 0.567 |  |

Against the public [BRIGHT leaderboard](https://brightbenchmark.github.io/) (Short Document, verified September 2026):

| Rank | System | nDCG@10 | Approach |
| --- | --- | --- | --- |
| 1 | Mira-Reasoning-Retrieval (Forward AI Labs) | 0.669 | Specialized retriever |
| 2 | INF-X-Retriever (INF) | 0.634 | Specialized retriever |
|  | Ours: Postgres + hybrid + autoresearch loop | 0.556 | Off-the-shelf retrievers + agent |
| 3 | RakanEmbed4B (RakanLabs) | 0.524 | Specialized retriever |
| 4 | NeMo Retriever's Agentic Retrieval (NVIDIA) | 0.509 | Agentic |
| 5 | DIVER-v3-GroupRank (Ant Group / SYSU) | 0.468 | Specialized retriever + reranker |
| 6 | BGE-Reasoner-0928 (BAAI) | 0.464 | Reasoning-tuned retriever |

Every system at or above our score fine-tunes an embedding model on reasoning-intensive data. We use OpenAI's general-purpose text-embedding-3-small and pg\_textsearch BM25, both off the shelf. The only specialization is a prompt and, on two domains, a text column. Against the one other agentic system on the board, NVIDIA's NeMo Retriever at 0.509, we're +0.047 with the same general approach and a different methodology for getting the agent to perform.

The leaderboard's top tier is dominated by training-based approaches, and we sit in the middle of it with a system that does no task-specific training.

## The thin stack costs more per query, and that's fine

The honest case for the thesis has to own this. Thin at build time, expensive at inference time.

Per query, measured via `claude -p` invocations through the same MCP server and prompts as the headline eval:

| Setup | $/query | wall sec | mean tool calls |
| --- | --- | --- | --- |
| BRIGHT, Opus max (10 of 12 domains) | ~$0.65 | ~30 | 10–12 |
| BRIGHT, Sonnet xhigh (aops, theoremqa_questions) | ~$0.34 | ~70 | 17 |
| MuSiQue, Haiku | ~$0.09 | ~77 | 11 |

A specialized retriever (Mira-class, BGE-class) runs ~$0.0001 per query at <1s latency. Roughly 6,500× cheaper and 30× faster than our agent loop. The retriever carries a training cost the agent loop doesn't: our estimate is $10K to $100K to fine-tune a ~500M-param reasoning-aware embedding model, amortized over 100K to 10M queries before the next retrain.

With $50K training amortized over N queries, the retriever's total cost equals our $0.65/query when N ≈ 77K. Below that, the agent loop is cheaper on total cost. Above it, the retriever wins. A 1,384-query benchmark sits well below. A 100K-queries/month customer support system sits well above.

The crossover moves each model generation, because holding capability fixed gets cheaper. [Epoch AI](https://epoch.ai/data-insights/llm-inference-price-trends) tracked the price of hitting a fixed benchmark milestone across six benchmarks over three years and found declines of 9× to 900× per year, with GPT-4-level performance on PhD-level science questions falling 40× per year. We project with 10× per 12 to 18 months, at the conservative floor of that range, because the fastest declines Epoch measured are recent and may not persist. Haiku 4.5 today handles work that needed Sonnet 3.5 a year ago at a fraction of the per-token cost.

A specialized retriever is already near the floor of inference economics and its capability is fixed at training time, so it does not ride this curve without a retrain. At 10× per 18 months, the crossover sits at ~77K queries today, ~500K in 18 months, and ~2.5M in three years. Each generation, more workloads land on the agent-loop side. Pick a faster decay from Epoch's range and the crossover arrives sooner, which is why we used the slow end.

Whether per-token prices keep compressing at all is the prediction here. The cost is real today. The trajectory is the case for treating it as a near-term tax rather than a structural disadvantage.

## RAG's bitter lesson

This isn't a leaderboard victory lap. It's a posture for building RAG systems in a regime where the model is improving faster than your pipeline can.

A thin stack rides the model frontier. A complex pipeline has to be rebuilt to keep up. That's Sutton's [bitter lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html) applied to retrieval, and it takes three concrete shapes here.

The thin stack is already competitive. A capable model plus hybrid search plus RRF plus an agent loop matches or beats most of the complex pipelines on both benchmarks today. Most architectural complexity in the literature was solving for yesterday's model.

The thin stack carries less debt forward. When the next model lands, our stack is one Postgres table and an MCP tool. The new model plugs into the same primitives and is immediately better at using them. A knowledge graph, a hierarchical index, a fine-tuned retriever: each gets re-justified against the new model's baseline, and often torn down.

What structure remains is either portable or removable. Hybrid search, RRF, and the tool-using agent loop are portable: they aren't bets on the current model's weaknesses, they're primitives a stronger model uses better. The two optimizations the loop adopted on BRIGHT are the other safe kind. A prompt is text. A sketch is a column. When they stop earning their keep, you delete them. Same loop, opposite verdicts on the two benchmarks, same underlying rule: stay thin by default.

The thin stack already being competitive on quality in 2026 is the finding. Winning on quality, maintenance, and cost over the next few model generations is the prediction. We're betting on it, and the bet is cheap. If it's right, every pipeline built for yesterday's model becomes someone's maintenance burden. If it's wrong, the cost of being wrong is a Postgres table and a prompt.

The full code, experiment log, and per-domain methodology notes are at [https://github.com/timescale/lab-rag-retrieval](https://github.com/timescale/lab-rag-retrieval). If you want to run the same stack, start with [pg\_textsearch](https://github.com/timescale/pg_textsearch) and the [hybrid search docs](https://docs.tigerdata.com/).