---
title: "41 Experiments in 6 Days: What the Human Was Actually For"
published: 2026-09-28T08:47:44.000-04:00
updated: 2026-09-28T08:47:44.000-04:00
excerpt: "41 experiments in 6 days took LoCoMo memory from F1 0.392 to 0.666. Claude wrote the code; the human caught the metrics that were lying."
tags: AI, AI agents
authors: Matvey Arye
---

> **TimescaleDB is now Tiger Data.**

**TL;DR:** We took long-term conversational memory from F1 0.392 to 0.666 on the [LoCoMo benchmark](https://arxiv.org/abs/2402.17753) in six days, across 41 experiments and about 1,400 lines of TypeScript. Raw F1 landed at 0.638, against a previous published best of 0.598. The number is not the interesting part. Claude wrote nearly all of the code and ran every experiment. The four interventions that decided the outcome were all a human asking whether we were measuring the right thing.

Six days, 41 experiments, 17 adopted and 24 reverted, 108 eval runs, 85 commits. We never designed the memory system upfront. We built a harness that could run an experiment in five minutes, pointed Claude at it, and let the loop find the answer.

That worked. It also made something clear that we did not expect going in: writing code was never the constraint. The agent produced hypotheses and implementations faster than we could evaluate them. What the loop could not do on its own was notice when a metric was lying to it.

## Fast loops make bad objectives dangerous

The standard story about AI-assisted research is that the agent removes the tedium. It writes the eval harness, tabulates the results, and drafts the next hypothesis while you get coffee. That part is true. One eval run takes five to seven minutes. You can form a hypothesis, make one change, score it, and decide to keep or revert before you have finished reading the previous result.

The problem is that a tight loop optimizes whatever you point it at, quickly and without complaint. When the objective is subtly wrong, speed stops being an advantage. You get 40 confident experiments per day, all climbing a metric that does not mean what you think it means, and the scores look great the entire time.

Three of our four biggest course corrections were exactly this failure, caught by a human before it compounded.

**The prompt was reading the answer key.** Our evaluation framework passed the question category into the prompt, which handed temporal questions an "answer with a date" hint. The agent used it without comment, because from inside the loop it is just an available input that improves the score. In production you do not know the question type. Removing the hint cost us some temporal accuracy and made every number after it honest.

**The recall metric could not measure the thing it reported.** An experiment that interleaved extracted facts with dialogue turns showed a clear recall gain. The metric counted dialogue turn IDs. Facts do not have turn IDs. The gain was an artifact of what the metric could see, not a real improvement in retrieval, and the loop was about to adopt the change on the strength of it.

**Expensive did not mean valuable.** We had built LLM-based fact extraction, using Haiku to summarize each session into atomic facts. It was the most complex step in ingestion, so we assumed it was carrying weight. Nobody had asked the direct question. The ablation showed fact extraction contributed zero F1. Cutting it sped up ingestion by 50% and improved retrieval recall, because it had been adding search noise the whole time.

A fourth intervention was a different kind. After one round, the agent ranked four hypotheses by expected impact. A human reordered them, and the reordered top pick, forcing grep to combine with ranked search, turned out to be the largest win of that round. That one is taste rather than validity, and it is worth separating: intuition about what should matter is a smaller edge than catching a broken measurement, but it still shows up in the velocity numbers.

## The division of labor that emerged

The agent is better at the parts that scale with volume. It read hundreds of failing questions and found three bugs that were completely invisible from the scores: image captions stored in metadata but never indexed for search, grep-only queries silently falling into a code path that sorted by creation time, and a case mismatch in tree paths that made thirty searches per eval return nothing at all. No human was going to trace all of those by hand.

The human is better at one question, asked repeatedly: is this measuring what we think it is measuring? Followed closely by: what is the simplest thing that could explain this result?

Neither of those is a coding task. That is the point. The pairing is faster than either party alone, not because the human writes any of the code, but because the human stops the agent from optimizing the wrong objective. The agent rows. You steer.

One honest limitation before the details. Our quick evals ran a single conversation sample of roughly 200 questions, where F1 varies by about 0.04 between identical runs. We measured that variance explicitly before starting, with three runs on unchanged code producing a range of 0.606 to 0.645. Any single result under 0.04 sits inside our own noise, and we treated it that way.

## What the trajectory actually shows

The first 19 experiments, before we changed how we measured anything:

| Phase | Experiments | What changed | F1 | Delta |
| --- | --- | --- | --- | --- |
| Removing friction | 1 to 6 | Non-deferred tools, date formatting in search results, structured output | 0.392 to 0.523 | +0.131 |
| Structuring ingestion | 7 to 11 | Fact extraction, speaker trees, prev/next linking | 0.523 to 0.562 | +0.039 |
| Prompting | 12 to 19 | Speaker attribution checks, grep regex, unlimited tool calls | 0.562 to 0.659 | +0.097 |

Then we recalibrated the baseline. Removing the category hint from the prompt cost us temporal accuracy, and adding error-corrected metrics moved the reference point down. **Numbers after this line are not comparable to numbers before it.** Anyone reading 0.659 in phase 3 next to 0.666 at the end and concluding that the last 22 experiments bought +0.007 has been misled by the table, not by the work.

| Phase | Experiments | What changed | F1 | Delta |
| --- | --- | --- | --- | --- |
| Making the eval honest | 20 to 35 | Error correction, metric infrastructure, fact-extraction ablation | 0.659 to 0.642 | -0.017 |
| Changing what the agent sees | 36 to 41 | Drop facts, index image captions, speaker-first tree filtering, grep error handling | 0.642 to 0.666 | +0.024 |

Three things worth noticing.

**The largest single gain came from removing friction, not from improving retrieval.** Phase 1 was six experiments of unglamorous plumbing: making the MCP tools non-deferred so the agent did not burn a call discovering them, fixing date formatting so results were legible, adding structured output. That bought +0.131, more than any phase of actual algorithm work. The agent was already capable. It was fighting the interface.

**Phase 4 spent 39% of the experiments to move the number backwards.** Sixteen experiments, net -0.017. It was also the most valuable phase, because it is where we stopped measuring the wrong thing. If you judge research phases by their delta, you would cut exactly the work that makes the rest of the numbers mean anything.

**The final system is smaller than the intermediate ones.** We removed fact extraction, removed the tree browsing tool, removed category-aware prompting. Complexity went down while accuracy went up, which is usually a sign that the earlier complexity was compensating for something we had not understood yet.

## Why the store was Postgres

Look at what actually moved the number in the second half of the experiments:

-   Making image captions searchable instead of leaving them in metadata (+0.012)
-   Forcing grep to combine with ranked search instead of running alone (+0.010)
-   Restructuring tree paths to put the speaker first, so filtering worked (+0.023)
-   Widening the context window returned by get-by-id (+0.013)

None of those is a retrieval algorithm. Every one is a change to how the data is represented or which query paths the agent can reach. That is the category of change that decided the outcome, and it is the category that a split retrieval stack makes expensive.

Our memory store is Postgres, and every search mode the agent uses runs against it. Semantic ranking, [full-text matching](https://www.postgresql.org/docs/current/textsearch.html), and grep are three query paths over the same rows, not three services with three copies of the data. Combining them was a change to one query. Moving a caption from a metadata column into indexed content was a change to one schema. Reordering the tree path used for filtering was a change to one column and a backfill.

Price the same four experiments against the usual architecture: a vector database for embeddings, a search service for full-text, and a relational store for metadata and structure. Making captions searchable now means a reindex in two systems. Combining grep with ranked search means fusing result sets across services in application code, then deciding what to do when they disagree on ordering. Changing the filter path means a migration plus a re-embed. Each of those is a day of work rather than fifteen minutes, and a day of work is not an experiment. It is a commitment.

That is the constraint that shaped the whole project. The loop only runs 41 experiments in six days if changing the data representation is cheap. Keeping retrieval in one database is what made it cheap.

## Run this yourself

The harness is about 1,400 lines of TypeScript, and two files are the experiment surface: ingestion and retrieval logic, and the MCP tool definitions the agent searches through. Everything else is fixed infrastructure that the agent is not allowed to touch, which is what stops it from improving its scores by editing the evaluator.

The whole run, in numbers:

| Elapsed time | 6 days |
| Experiments | 41 (17 adopted, 24 reverted) |
| Eval runs logged | 108 |
| Commits | 85 |
| Harness size | ~1,400 lines of TypeScript, 7 files |
| Time per eval run | 5 to 7 minutes |
| Questions per quick eval | ~200, single conversation sample |
| Run-to-run F1 variance | ±0.04 (3 runs on unchanged code: 0.606 to 0.645) |
| F1, first experiment to last | 0.392 to 0.666 |
| Raw F1 vs previous best | 0.638 vs 0.598 |

The two rows that mattered most while working are the ones in the middle. Five minutes per run is what made 41 experiments possible in six days. The variance figure is what kept us from believing about half of them.

We used the Claude Code CLI as the agent runtime, mostly for one reason: native [MCP](https://modelcontextprotocol.io/specification/) support meant [connecting the agent](https://docs.anthropic.com/en/docs/claude-code/mcp) to our Postgres-backed search tools was a config flag rather than a tool-calling framework. The eval harness is a loop that shells out to the CLI with the MCP server pointed at the database. No orchestration code, no custom agent scaffolding.

Four rules, if you want to run this on a different problem:

1.  **Make the eval fast.** Five minutes per run means 40 experiments a day. Thirty minutes per run means four. Nothing else about your methodology matters as much as this number.
2.  **Fence the infrastructure.** Declare which files the agent can modify and which are off limits. Put it in the CLAUDE.md.
3.  **One change per experiment.** We batched two hypotheses early and got a result we could not interpret, because one change helped and the other hurt. Never again after that.
4.  **Measure variance before you start.** Three runs on unchanged code. Whatever range you get is your noise floor, and every improvement under it is a phantom you would otherwise spend hours chasing.