Back to Blog Overview

Attribution at Machine Speed: Why Scientific Records Need to Be Verifiable

As discovery accelerates, attribution becomes the harder problem. With Molecule Labs, a verifiable and traceable record is created automatically.

7 min read
Ella McCarthy-Page
Attribution at Machine Speed: Why Scientific Records Need to Be Verifiable

How a Millennium Problem Has Upended the AI World

The Navier-Stokes equations describe how fluids move, and the question of whether their solutions always remain well-behaved has stayed open for well over a century. It is one of the seven Millennium Prize Problems, each carrying a million-dollar reward from the Clay Mathematics Institute, and settling one of them defines a career. This week OpenAI announced that an internal AI system had resolved a version of the problem in which a smooth external force acts on the fluid, while stating that it does not intend to claim the prize for the result. Within days the announcement had become a public dispute over who reached what first. We think an immutable record (hey blockchain!) could’ve resolved the issue.

What Happened

Tristan Buckmaster, a mathematician at NYU, and his collaborator Levent Alpöge had been working toward closely related results for months, using several AI systems as they went, among them Codex, OpenAI's coding assistant. Their drafts went into those Codex sessions throughout the project, which placed the working archive of an unpublished result inside a product built by the company they would shortly find themselves in dispute with.

Realizing that OpenAI had solved the problem they were feeding into Codex, they reached out. According to the statement Buckmaster published the questions that mattered went largely unanswered. He asked when OpenAI had first begun prompting on the problem and was told, after some delay, that it had been within the previous few days, once word of his and Alpöge's work had reached the company. He asked whether the model had been trained on the Codex sessions holding their drafts, and on that point received no answer. He was shown a prompt rather than the proof, and by his own account he has still not seen the proof itself.

Why Neither Side Could Prove It

Neither party was in a position to settle the matter by opening its files, which is the ordinary condition of research rather than a failure of either side. Unpublished methods and intermediate results are among the most valuable material a scientist holds, and handing them to a counterparty in order to win an argument about timing is a poor trade. Verification in the only form available here would have meant disclosing the very thing at stake, so the argument moved into public instead, where it is now being conducted through statements, press coverage and inference.

Attribution at Machine Speed

Scientific priority has always been established through the informal machinery of submission dates, dated notebooks and the collective memory of a field, and that machinery held up reasonably well for as long as discovery moved at the speed of the people doing it. It holds up considerably less well now that a well-resourced organisation can direct thousands of AI agents at a problem within days of hearing that it might be close to solved, and it offers almost nothing to a researcher whose working archive happens to sit inside the infrastructure of the party on the other side of the argument. As discovery accelerates, attribution becomes the harder problem, and the systems science uses to record who did what were designed for a slower era.

What a Hash Proves

A cryptographic record of a file is a modest thing that answers precisely this question. Hashing a document produces a short, unique fingerprint of its contents, and writing that fingerprint to a public blockchain fixes the moment at which the document existed in that exact form along with who held it. It also dissolves the standoff over disclosure, because a fingerprint reveals nothing whatsoever about the contents of the file it points to, which means a researcher can prove they held a specific result on a specific date without discolsing what that result was.

How a Lab Records It

Molecule Labs was built to produce that record as a matter of course. Every file uploaded to a Lab receives a content identifier, a cryptographic hash derived from the contents of the file itself, so that changing a single byte produces an entirely different identifier. Each version carries that hash alongside a timestamp, the author's decentralised identifier linked to their wallet, and a pointer back to the version that preceded it, and because the history is append-only, no version can be quietly revised after the fact. The Lab's onchain account records the identifier and its metadata, producing a tamper-evident link between the researcher and the file that anyone can verify independently, without needing to trust Molecule or any other single provider. Files marked confidential are encrypted and stay closed, while their existence, their timing and their authorship remain publicly verifiable.

Why Mathematics Went First

Mathematics is the canary in the coalmine. It is the field where the models are furthest along, and it carries the least friction between having an idea and having a finished result, since a proof requires no reagents and no eighteen-month experimental timeline. A claim can be produced, checked and contested inside a single week, which is why one of the first public arguments over machine-assisted priority has happened in mathematics rather than anywhere else.

Biology Is Next

Biology sits further back on that curve, with human researchers still necessary in the loop in ways they are no longer strictly necessary for a machine-verifiable proof, though the direction of travel is not in question - better models, laboratory robotics and simulation will close the distance. However, we purport that the stakes are larger. The same dispute inside drug discovery decides patents, licensing revenue and ownership of the company built around the result.

Whatever the outcome for the parties involved in this case, the more pressing question for every other research team is what they would be able to show if the same thing happened to them next month. A Lab answers it with a timestamped, independently checkable history that nobody in the argument controls, and it does so without asking a researcher to reveal a single line of unpublished work.

Create your own Lab.

About the author

Ella McCarthy-Page

Ella McCarthy-Page

BSc(Hons) in Biochemistry and Materials Science. A communicator working at the intersection of biotechnology and web3.