The CIA and related secrecy issues are part and parcel of the difficulty in communicating the importance of Hume’s Guillotine prize funding:
We are in a profoundly adversarial environment where conflicts of interest over causal narratives at the macro social scale go beyond mere ethical laziness.
People have a hard time conceiving of any scientific academic field as being as viciously adversarial as they must in order to understand why I insist on no withheld data for validation and lossless compression on data sets that are agreed upon by adversaries. This is quite aside from the technical issues involved in educating the public that Jay Forrester confronted. Hell even Forrester didn’t conceive of how serious one must be regarding information criteria for model selection. However I don’t think that Forrester would have been confused by reliance on data even though he makes a point of how little information is represented in data sets compared to wetware. The idea that you have to do cultural archeology to get at the structures underlaying policies is correct. The question is how do you go about funding that archeology? I think I could have convinced him with little more than 5 minutes conversation.
When you become this profoundly serious about ethics, people think you are paranoid.
Just one dimension of the dilemma facing anyone that attempts to advance the scientific method in this manner is just how ridiculous is the idea of “validation datasets” when the stakes are macrosocial narrative control. It’s tiresome reiterating this over and over but perhaps LLMs can help relieving some of my stress:
Exactly. A hidden validation set solves the wrong threat model.
It assumes the main adversary is an overfitter who lacks access to the test answers. But in a high-stakes macro-social modeling prize, the hidden dataset itself becomes a secret-bearing instrument of authority. Then the contest is no longer just about modeling. It is about access control, insider risk, leakage, oracle attacks, selective disclosure, and suspicion.
The machine-learning world already has a weaker version of this problem. Adaptive data-analysis work shows that repeatedly querying a holdout set can invalidate ordinary statistical guarantees because later analyses are shaped by earlier results. Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth framed this as a central problem of adaptivity: classical validity assumes a fixed procedure chosen before seeing the data, while real analysts adapt after prior interactions with the data. (Google Research)
And that is before we even get to espionage. Leaderboard-oracle attacks show that even partial score feedback can leak information about hidden labels. Whitehill demonstrated an attack on log-loss feedback that could infer test-set labels and reported achieving a high Kaggle leaderboard rank without training a classifier for the actual task; the paper concludes that evaluation should be performed on test examples the oracle never examined.
But your point is stronger: in a macro-social causal prize, even a perfectly designed hidden test set creates a legitimacy trap.
The more people agree that the validation data are fair, the more valuable they become. The more valuable they become, the more rational it is to suspect that a winning model may have gained privileged access. Even a genuinely superior model is then socially contaminated by the possibility of leakage. The secret dataset becomes an epistemic crown jewel.
That means the usual hidden-test protocol produces three failures:
-
It creates a prize target for espionage.
Anyone with access to the validation set, its construction process, its metadata, or even partial feedback has possession of something strategically valuable.
-
It creates unverifiable legitimacy.
Outsiders cannot fully audit whether the winner generalized, memorized, leaked, inferred, or benefited from privileged curation knowledge.
-
It confuses secrecy with rigor.
The validation set is treated like a cryptographic secret, but scientific validation should not depend on uninspectable authority.
That is why your open-data/lossless-compression framing is not just a technical preference. It is a governance requirement.
The core move is:
[
\text{Do not hide the evidence. Price the explanation.}
]
With public, adversary-agreed datasets, everyone sees the same evidence. A contestant may “overfit” only by paying for the bits required to do so. If the model memorizes exceptions, encodes special cases, hard-codes historical episodes, or carries elaborate preprocessing, those choices appear in the executable description length. The anti-overfitting mechanism is no longer secrecy. It is accounting.
For Hume’s Guillotine, I would state the principle this way:
Hidden validation data are unsuitable for high-stakes macro-social causal modeling because they convert scientific evidence into a strategic secret. Validation must instead be conducted on adversary-agreed public corpora, with complete executable submissions scored by lossless compression, structural admissibility, behavioral reproduction, and policy-intervention semantics.
The strongest version is not “we trust no one.” It is:
The protocol should not require trust in anyone’s private access to decisive evidence.
That is why people misread the seriousness as paranoia. They are imagining ordinary academic sloppiness. You are describing an environment where causal narratives affect money, institutions, policy, reputation, and state power. Under that threat model, hidden validation sets are not neutral safeguards. They are classified assets waiting to happen.