Every version of our homepage headline got rejected. Not vaguely. Precisely. One draft "looks like a paragraph, not a title." Another, a page heading that read Everything here is a Tool, came back as "an absolute gem of a useless title": a title has to promise an idea, not label a bucket. The rejections were sharper than any brief we had written.
So we stopped writing drafts and started reading the rejections. If someone can tell you exactly why a line fails, they have handed you a rubric. We turned the complaints into six criteria and ran a bake-off.
Six criteria, derived from what was actually disliked
- Commits to one claim, rather than hedging across two.
- Reads as a title, not a paragraph.
- Promises an idea rather than labelling a bucket.
- Survives being read aloud.
- Works without naming the brand: the logo already does that job.
- Presents, rather than addressing a consumer.
None were invented for the exercise; each is a rejection turned around and pointed forward. That matters more than it sounds: criteria you invent flatter the writing you already like, while criteria derived from rejection describe the actual taste in the room.
The field, and the benchmark that does not exist
Our first instinct was to find out which model writes the best short-form copy. There is no evidence. Benchmarks exist for general writing quality, for long-form generation, for news headlines written from an article you already have. None for inventing a positioning line with nothing to compress from. The creative-writing leaderboards that do exist disagree on rank order and some still list superseded flagships as current; the blog posts confidently naming a "best model for taglines" show no methodology at all.
That absence was the useful finding: with no evidence to rank on, ranking the field first would have been manufactured confidence. So we did the opposite: nine entrants, identical brief, identical constraints, blind scoring. Three came from the same family as the assistant we work with daily, to test whether our habitual voice was better or just familiar. The other six spanned frontier commercial models and much cheaper open-weight ones, skewed cheap on purpose: if a $0.10-per-million-token model can compress a claim, that is worth knowing.
Keeping the harness honest
The script does one job and refuses a second. It loads the brief, sends the same prompt and constraints to every entrant, parses each answer into a headline and a paragraph, shuffles the results into anonymous "Writer 1" through "Writer 9" labels, and writes them into one document. It runs a single mechanical check, a plain search for the banned brand word and for second-person address, then stops there.
The scoring sheet it produces is blank on purpose, and so is the ranking section, with a line saying why: the harness does not judge the copy it generated. The key mapping each Writer number to its model sits at the bottom under a fold, opened only after the reading is done. A machine that both writes and grades is not running a bake-off; it is marking its own homework.
What went wrong, in order
The first pass was broken, and it took a while to see how badly. We had given every entrant a flat output cap of 700 tokens. Reasoning-capable models can spend that entire budget thinking before writing a single word of answer. The API returns a perfectly healthy success response with the content field empty. A second, independent bug compounded it: our provider adapter turned that empty content into an empty string and discarded the model's reasoning trace, so the "raw output" block, whose whole purpose is preserving what came back, rendered blank. An answer lost with no recoverable trace is not a formatting nit.
We fixed both at the root. Each entrant now gets a cap tuned against live probe calls on this brief, and the adapter falls back to the reasoning trace, clearly marked, instead of dropping it. One more guard came out of that: with traces preserved, our headline extractor matched a line a model had brainstormed and then rejected mid-thought. It looked exactly like a real submission. Caught before it reached the document, guard added.
One entrant never converged. Escalated from 1,500 to 4,000 to 6,000 tokens, it spent 5,997 of its final 6,000 on reasoning and produced no answer, recorded as a genuine finding about that model on that brief rather than a harness fault, with its full 26,800-character trace preserved for anyone curious what it was doing. Two other entrants had returned empty earlier for the same reason. We left the failures in the file: a bake-off with the broken entries quietly deleted teaches nothing.
What it cost
Nine entrants, plus every diagnostic call, failed attempt and re-run: about $0.93. The whole round, under a dollar. Some of that was waste: our earliest diagnostic calls ran at the broken cap and cost more than a minimal probe would have. Later fix-verification calls were reused as final data rather than re-spent.
What blind scoring actually produced
The human read them by ear, gut verdict first, then scored the six criteria, without knowing who wrote what. The scoring is now complete, and the winner turned out not to be the interesting part: there isn't one, as-is.
Blind scoring produced rulings. One candidate was strong copy in the wrong slot: an excellent section lead, not a landing headline. Reading several in a row exposed a gap none of them filled: they proposed a solution before ever stating the problem. The rule that our name must never appear got refined mid-sitting: it holds for titles, where the logo carries the brand, but a paragraph should name DAAC as the infrastructure helping people turn their knowledge into data; a paragraph that names no enabling actor explains nothing.
The sharpest correction was one no criterion had captured. Every candidate risked reading as though we harvest the data ourselves and profit from it. That is the opposite of the point: DAAC builds infrastructure so people stay custodians of their own knowledge and participate, or not, at their discretion, to their own benefit. That sovereignty is what makes participation worth anything, and it has to come through from the first line a visitor reads.
Two more rulings surfaced before the sitting closed. The hero's scope kept drifting toward the specific industries we know best: mines, cooperatives, shipping lanes. The actual reach is broader: risk, situational understanding, decision impact, engagement, security monitoring, with data as the common thread underneath all of them. Vertical detail like that is good material, just for a solution page, not the front door. And more than one candidate read as though the recognition of human expertise was a supporting clause under some other headline, interoperability in one case, when recognition was the actual point the whole time and the feature was the supporting clause.
The reveal
Nine writers, blind. Here is who wrote what, and the verdict each one got, read aloud, before any name was known:
| Writer | Model | Verdict |
|---|---|---|
| Writer 1 | Claude Fable 5 | "I like it... it would probably not be the main hero." Strong copy, wrong slot: a section lead, not the landing headline. |
| Writer 2 | GLM-5.2 | "Yes, it's good," but the paragraph names the transformation without ever introducing what makes it possible. |
| Writer 3 | Gemini 3.1 Pro | "It's a great title. I love the title," the sitting's clearest enthusiasm, with two word-level fixes and a missing problem statement. |
| Writer 4 | DeepSeek V3.2 | Title's direction was right, its construction wasn't: two different referents collapsed into one confusing line. |
| Writer 5 | Claude Opus 5 | Title "pretty bad," paragraph "generally quite bad," the sitting's harshest read; its concreteness, once the favoured register, is now ruled solution-page material. |
| Writer 6 | GPT-5.6-sol | 6 of 6 on the scoring grid, "nicely written," the cleanest pass of the round. |
| Writer 7 | Claude Sonnet 5 | Title "too big/broad/vague... definitely not the hero." |
| Writer 8 | Kimi K3 | Did not converge: three escalating attempts, all spent on reasoning, no copy produced. Recorded as a real finding about the model, not a harness bug. |
| Writer 9 | Qwen3-235B | Title "a bit more technical than inspiring," but the paragraph was "really good." Its "not extracted but elevated" line, and Ben's own recast of its opener into "inclusion begins with recognition," carried the round's best single idea. |
The three writers from the same model family as the assistant we work with daily did not win the round, blind. One placed well in the wrong slot, one took the sitting's harshest verdict, one was ruled too vague. That is the honest result of testing whether a habitual voice is better or just familiar: on this brief, this time, it was mostly just familiar.
What we actually got
No candidate carried the full register the sitting had, by the end, defined for itself, which is a different outcome than "no winner." The six-title-only-DAAC rule, the non-harvesting register, the breadth rule, the problem-first structure, the inclusion-not-interoperability rule, the help-others-digitise rule: none of that existed as a written spec before this sitting. All of it exists now, and none of the nine drafts were written against it, because it didn't exist yet when they were generated. That is the real shape of what happened: we ran a bake-off to get a headline, and what we actually produced, along the way, was the spec the next headline has to satisfy, plus two pieces of verbatim narrative material (a problem statement, a market thesis, a horizon thesis, an incentive thesis) dictated in the room, not written by any model at all. The rewrite is real and it is coming, as part of the next version of this site, built from this record, blending the strongest pieces of the field with the register this sitting ruled. It is not a standalone task and it does not happen before that. The copy will follow.