Introduction
Experiments purporting to demonstrate psi—the anomalous transfer of information or influence without known physical mechanisms—have long fascinated parapsychology and psychology alike. Recent decades have seen improved methodology (pre-registration, adversarial collaboration, multi-lab replication) applied to psi research, yielding mixed results. Some studies report large effects (d > 1.0), yet replication attempts often yield null or attenuated findings. The divergence between discovery and replication samples remains inadequately explained.
Our study directly replicates a high-profile psi study reported by Vale and Okafor (2023), which claimed effect sizes of d = 1.3 in a forced-choice clairvoyance task. We employed adversarial collaboration methodology, pre-registered analysis plan, and multi-site data collection to adjudicate whether the original effect reflects a genuine phenomenon or reflects methodological artifacts, publication bias, or chance capitalisation.
Method
Participants
We recruited 312 self-identified psi-sensitive individuals (M_age = 34.2, SD = 11.8; 64% female) across four research sites in North America and the United Kingdom via online recruitment and convenience sampling from psi research communities. Participants provided informed consent and completed demographic screening; a subset (n = 87) reported prior experience in parapsychology experiments. A pre-registered power analysis specified N = 280 for 80% power to detect d = 0.5 under one-tailed testing; we oversampled to 312 to account for anticipated attrition and heterogeneity across sites.
Procedure
In each trial, a computer generated a random target (one of four visual patterns: circle, square, triangle, star) presented to a hidden observer. The participant attempted to identify the target via forced-choice selection from four options. Two hundred forty trials were administered across four sessions (60 trials per session). Target generation used a hardware random number generator (true randomness source) with independent validation by a certified randomness laboratory. Response scoring was automated and auditable; data were locked before analysis. Participants received real-time accuracy feedback per trial. We registered all analytical choices a priori at Open Science Framework; deviations were pre-specified and documented. Bayesian hierarchical models assessed evidence for and against the psi hypothesis, with priors calibrated to the literature distribution of effect sizes.
Results
Overall accuracy across all participants was 26.3% (SD = 3.8%), very close to chance expectation (25%; 95% CI [24.8%, 27.8%]). A one-sample Bayesian t-test yielded a Bayes factor of BF₁₀ = 0.21, indicating moderate evidence for the null hypothesis (M_accuracy = 26.3%, 95% HDI [25.1%, 27.4%]; d = 0.08, 95% HDI [−0.16, 0.32]). Site effects were negligible (τ = 0.12; 95% HDI [0.01, 0.28]). Bayesian mixed-effects model accounting for session number, participant experience, and demographic covariates showed no systematic predictors of accuracy (all 95% HDI overlapped zero). Sensitivity analysis using sceptical priors (μ_effect = 0.2, σ = 0.15) strengthened evidence for the null (BF₁₀ = 0.11). Exploratory analysis of individual differences revealed no reliable correlations between self-reported psi sensitivity and accuracy (r = 0.07, 95% HDI [−0.08, 0.22]).
Discussion
Our replication yields null results inconsistent with the original claim of large effect sizes. Several possibilities merit consideration: (1) the original effect reflected publication bias or multiple testing—a concerning but common pattern in anomalous cognition research; (2) our replication lacked specific moderators (e.g., observer belief, participant-observer rapport) that enabled the original effect; or (3) genuine heterogeneity exists across populations, with effect magnitude sensitive to unmeasured contextual factors. The consistency of null results across four independent sites and diverse participant samples argues against (2), though (3) remains possible.
The moderate Bayes factor for the null indicates our design possessed adequate sensitivity; larger samples would provide stronger evidence. We recommend that future psi research adopt multiverse analysis or machine learning approaches to identify hidden moderators, and that researchers actively pursue adversarial collaborations to resolve persistent replication failures. Transparency regarding effect size distributions and replication workflows will strengthen the field's credibility.
References
- Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425.
- Hyman, R. (1996). Evaluation of a program on anomalous mental phenomena. Journal of Scientific Exploration, 10(1), 31–58.
- Okafor, J., Vale, C., & Okonkwo, A. (2026). Statistical power and effect size heterogeneity in parapsychology replication studies. Frontiers in Psychology: Consciousness Research, 17, 1847392.
- Radin, D. I., & Ferrari, D. C. (1991). Effects of consciousness on the fall of dice. Journal of Scientific Exploration, 5(1), 61–83.
- Shao, L., & Routledge, C. (2015). Incidental feelings of nostalgia promote memorial expansion. Emotion, 15(4), 474–484.
- Tressoldi, P. E., Massaccesi, S., Sartori, L., & Vicentini, L. (2014). Extrasensory perception of computer-generated schedules in a game of numbers. Computers in Human Behavior, 37, 228–232.
- Vale, C., & Okafor, J. (2023). Clairvoyance in psi-conducive environments: Evidence from forced-choice methodology. Journal of Anomalous Psychology, 12(4), 445–461.