I’m the author of a new preprint on repeated-query auditing of LLM brand recommendations, and the founder of Rankfor.AI.
The practical question: how many times should we repeat a prompt before comparing results?
The paper applies generalizability theory: estimate variance components from a pilot, then calculate the repeat count needed for a chosen reliability target.
Tested the reliability predictions on three independently collected corpora covering political-orientation questionnaires and benchmark stability. Across 39 prediction cells, 37 met the prespecified replication criterion and two were partial matches.
The fixed iteration thresholds did not transfer. Other preregistered tests, including parts of the drift diagnostics, also failed. Those results are reported in the paper.
An important limitation i see is that these external corpora do not contain brand recommendations. They test the statistical machinery outside our original application which is independent replication on repeated brand-recommendation data remains outstanding.
I’d particularly welcome criticism of the pilot-based variance estimates and the reliability validation design. Does anyone know an independently collected brand-recommendation dataset with repeated identical prompts?
Preprint: https://arxiv.org/abs/2609.04047
External validation materials: https://github.com/Rankfor/rankfor-open/tree/main/research/dice-roll-method/external-validation