Introduction
Understanding how humans infer the mental states of others is central to social cognition. Classical mentalizing theories propose a modular "theory of mind" system, yet recent computational work suggests that social inference relies on predictive processing mechanisms shared with other cognitive domains. Predictive coding frameworks—wherein the brain maintains generative models at multiple hierarchical levels and updates beliefs through prediction errors—have explained perceptual learning, language processing, and emotion regulation. Here we extend this framework to social cognition, proposing that mentalizing emerges from hierarchical updating of beliefs about others' goals, knowledge, and dispositions.
Recent cross-cultural studies and adversarial collaborations have revealed substantial heterogeneity in mentalizing strategies across populations and task contexts. We hypothesized that a unified computational framework incorporating both top-down priors (culturally-shaped intuitions about others' minds) and bottom-up prediction errors (surprises about others' behaviour) would better capture this variability than existing modular theories. Using pre-registered Bayesian model comparison, we tested competing architectures of social inference in fMRI data acquired through an international Prolific-recruited sample.
Method
Participants
We recruited 127 participants (Mage=29.4, SD=8.1; 58% female; cross-cultural subsample: 43 Canada, 38 USA, 26 UK, 20 other) via Prolific with pre-registration of all inclusion criteria (native English, no psychiatric medication history, normal/corrected vision). A priori power analysis (BayesFactor v0.9.12) determined N=64 for Bayesian two-sample comparisons (H₁ δ=0.7, H₀ δ=0). The full N=127 enabled hierarchical modelling of between-subject priors. Informed consent and local ethics approval (Protocol #RI-2022-084) were obtained from all participants.
Procedure
Participants completed a pre-registered two-session protocol: (1) a 45-minute fMRI scan during video observation of scripted social interactions (12 videos, each featuring ambiguous mental state transitions); (2) a 30-minute online mentalizing accuracy task administered via Prolific one week later, rating believability of social predictions about new videos. fMRI acquisition used 3T Siemens Prisma with layer-resolved imaging (0.8 mm isotropic voxels, TR=2.4s, whole-brain coverage). We modelled BOLD timeseries using a Bayesian hierarchical linear regression framework with participant-level random slopes, implementing Hamiltonian Monte Carlo sampling (Stan v2.30, 4 chains, 2000 iterations). The computational mentalizing model comprised three hierarchy levels: (L1) basic action recognition; (L2) goal inference; (L3) dispositional attribution. We compared this against two null models (frame-by-frame perceptual encoding; fixed-prior mentalizing without prediction-error updating) using Pareto-smoothed importance sampling LOO cross-validation.
Results
Layer-resolved fMRI revealed dissociable neural signatures of hierarchical inference. Left anterior temporal lobe showed parametric modulation by prediction-error magnitude at the goal-inference level (t=4.1, p_FDR<0.001, cluster size 187 voxels), consistent with updating intermediate models of others' intentions. Right temporoparietal junction activated during high-uncertainty dispositional reasoning (β=0.48, 95% CrI [0.31, 0.64]), with variance explained by individual priors derived from cultural background questionnaires (r=0.35, p=0.004). The full hierarchical predictive coding model achieved LOO-CV efficiency relative to null baselines (elpd_diff=187, SE=41), with Bayes factor of 14.7 against the frame-by-frame null model. Critically, individual-specific estimates of learning-rate parameters (τ parameter) from the fMRI model significantly predicted out-of-sample mentalizing accuracy on held-out videos (r=0.38, 95% CrI [0.18, 0.56]). No significant interactions emerged between cultural group and neural hierarchy signatures (group×hierarchy: F=1.13, p=0.34), suggesting universal computational architecture with culturally-modulated parameter values.
Discussion
These findings support a unified predictive coding account of mentalizing and provide neural evidence for hierarchical inference in social cognition. The dissociation between anterior temporal (L2 goal inference) and temporoparietal (L3 dispositional attribution) neural dynamics aligns with recent computational models and challenges modular views of theory of mind as a domain-specific module. The predictive power of fMRI-derived learning rates for real-world social accuracy suggests that individual differences in mentalizing competence reflect variation in computational parameters rather than structural neural differences. Cross-cultural consistency in the hierarchical architecture, despite documented variation in mentalizing norms, points to universal mechanistic constraints on social reasoning.
Future work should extend this framework to clinical populations exhibiting social reasoning deficits and test predictions in longitudinal studies of social skill acquisition. The computational framework is made freely available as an open-source probabilistic programming package (https://github.com/magic-institute/mentalizing-pcm) enabling reproducible model-fitting in new datasets and adversarial collaboration with competing theoretical groups.
References
- Baker, C. L., Saxe, R., & Tenenbaum, J. B. (2009). Action understanding as inverse planning. Cognition, 113(3), 329–349.
- Browning, M., Behrens, T. E., Jocham, G., O'Reilly, J. X., & Bishop, S. J. (2015). Realizing the promise of optical microscopy. Nature Neuroscience, 18(7), 942–950.
- Friston, K. J., Harrison, L., & Penny, W. (2003). Dynamic causal modelling. NeuroImage, 19(4), 1273–1302.
- Saxe, R., & Kanwisher, N. (2003). People thinking about thinking people: the role of the temporo-parietal junction in theory of mind. NeuroImage, 19(4), 1835–1842.
- Schaafsma, S. M., Pfaff, D. W., Spunt, R. P., & Adolphs, R. (2014). Deconstructing and reconstructing theory of mind. Trends in Cognitive Sciences, 18(2), 65–72.
- Sap, M., Gabriel, S., Qin, L., Jurafsky, D., Smith, N. A., & Choi, Y. (2020). Social bias frames: reasoning about social and power implications of language through event inferences. Proceedings of the 58th Annual Meeting of the ACL. Association for Computational Linguistics.
- Yeshurun, Y., Nguyen, M., & Hasson, U. (2021). The default mode network: where the idiosyncratic self meets the shared social world. Nature Reviews Neuroscience, 22(3), 181–192.