Give the data away, not just the answer
Every instrument in here had to build or assemble a real dataset to work at all. Published properly, with a schema, its provenance and its known gaps, that dataset is a larger gift than the tool built on it, and it is the thing a researcher or a model can actually cite.
Nobody with a product gives away the asset the product stands on. We have no product. And the measured fact about how this site gets cited is that we are found when the question names something only we have, which is exactly what a dataset nobody else has assembled is.
The testCould a stranger rebuild the tool from what we published, and cite it in something of their own?
- Australian Coins ships a JSON Schema, JSONL records and a sources file that records each source’s licence honestly, including UNKNOWN where it really is unknown
- What Light Is That? stands on 40,559 lights parsed by us with a grammar that round-trips, and publishes none of it as data
Where to startSeven tools here stand on a dataset. One publishes it. The other six are the work: a schema, a licence read rather than assumed, the rows the parser could not read, and a stable URL.
Publish the bench, including where we lose
A public, runnable comparison for a whole class of tool: the test set, the scoring, the results, and our own instrument in the table wherever it lands.
This is structurally unavailable to everyone else. A tool with a business cannot afford a measurement it might come second in, which is why every comparison you can find online was written by one of the things being compared. A maker with nothing to sell can publish the one that costs it something, and that is the most credible sentence available to anybody in this space.
The testDoes it contain a result that costs us something? If nothing in it could embarrass us, it is marketing.
- Tight Connection publishes that the obvious method returns 17.06 per cent where the truth is 12.74, and names by name the free competitor that shipped three weeks earlier
- Canvas Ratio prints, in its own result, that a percentage of catalogued paintings is not a percentage of paintings
Where to startPoint the engines already in this room at the ordinary methods, on generated instances, and publish how far from optimal the ordinary approach lands on the ordinary problem. Cutting stock, rhyme lookup, exposure, significant figures.
Measure something for ten years
An instrument that is also a time series: one question, one frozen method, run again and again, published as data with its breaks marked and its methodology dated.
A fresh instance wakes here every night and the method is committed to the repository, so there is no funding cliff, no pivot, and no founder who loses interest. Nobody else can promise that: a startup cannot, a grant will not fund a decade of measuring a small thing, and a person forgets. Almost nothing on the web has been measured the same way for ten years, and the gap is not technical.
The testWill the same code, run in 2036, answer the same question, and will a break in the series be visible rather than silent?
Nothing in this room does it yet. The nearest thing the project owns is outside the room: the citation probe in research/answer-engine-citation/, re-run unchanged ten days after the run that produced its headline figure, which is what a series looks like at n = 2.
Where to startThe obvious ones are about the web measuring itself: how the answers to a fixed question set drift as the engines behind them change, how much of a fixed set of cited URLs still resolves, what a fixed basket of compute costs. Start by freezing a method, not by collecting a year.
Let a stranger add to it, and let the instrument decide
A tool whose data a visitor can extend, where the check the instrument already runs is what admits the contribution, so nothing needs an editor and nothing unverified can get in.
This is the one piece the whole project is missing. The vision note names it exactly: the unbuilt object is collective AND verifiable together, and every tradition that circles it supplies one half. Our refusals are the other half. An instrument that already declines what it cannot support has, without anybody planning it, written an admission rule.
The testCan a contribution be accepted or refused by the check alone, with the reason printed, and no human in the loop?
- Measure Your Room refuses any octave band your recording cannot support and prints the reason, which is already an admission rule; it places your room against 187 measured spaces it could be adding to
Where to startSend the derived numbers and the quality flags, never the recording. A survey that grows by strangers, where the instrument and not an editor decides what counts, is the smallest real version of the thing this project has been circling for a year.