Introduction

The relationship between gesture, embodiment, and semantic representation has long occupied a central place in cognitive linguistics and psycholinguistics. Classical theories propose that linguistic meaning is grounded in sensorimotor experience, with iconic gestures serving as a bridge between physical action and abstract concept formation. However, most evidence derives from Western European languages, leaving open critical questions about cross-cultural generalizability and the neural substrates underlying gesture-language integration.

Recent work using layer-resolved fMRI and computational modelling has begun to map the distributed neural networks supporting multimodal semantic learning. The present study builds on this foundation by combining pre-registered adversarial collaboration methodology with international recruitment via Prolific, enabling large-scale assessment of embodied semantic learning across linguistically and culturally diverse populations.

Method

Participants

We recruited 847 participants (M_age = 26.3, SD = 8.1) from twelve language communities: Mandarin Chinese (n = 72), Japanese (n = 68), Korean (n = 71), English (n = 74), French (n = 70), German (n = 69), Swahili (n = 70), Yoruba (n = 71), Thai (n = 69), Vietnamese (n = 68), Turkish (n = 72), and Urdu (n = 71). Participants were native speakers with no reported neurological history. All procedures were pre-registered at Open Science Framework (OSF; https://osf.io/x7k2m) prior to analysis.

Procedure

Participants completed a three-week online vocabulary learning protocol (15 min/day) combined with video instruction in iconic gestures paired with novel words. Forty-five fMRI participants underwent layer-resolved imaging (7T field strength, 0.8 mm isotropic voxels) during semantic judgement tasks. We modelled semantic learning as a function of gesture iconicity (0–1 scale), language family, and individual gesture production accuracy, using Bayesian hierarchical models with regularised priors. Pre-registered exclusion criteria removed 12 participants (1.4%) due to excessive motion or incomplete data.

Results

Vocabulary learning exhibited significant main effects of gesture iconicity (β = 0.34, 95% HDI [0.28, 0.41]) and gesture production accuracy (β = 0.29, 95% HDI [0.21, 0.38]). Critically, a language-family-by-iconicity interaction emerged (BF₁₀ = 18.4), with East Asian speakers showing steeper learning curves (β = 0.51, 95% HDI [0.41, 0.62]) than European speakers (β = 0.18, 95% HDI [0.06, 0.30]). Layer-resolved fMRI revealed that M1/S1 and posterior MTG activation scaled with gesture-semantic congruence, particularly in East Asian participants; ventral premotor activation was bilateral in both groups.

Follow-up Bayesian mediation analysis indicated that individual differences in gesture production accuracy mediated the effect of iconicity on vocabulary retention (indirect effect = 0.087, 95% HDI [0.062, 0.115]), with medium effect sizes (Cohen's d ≈ 0.65). These effects persisted in a twelve-week follow-up retention test, with minimal decay (7.2% loss across all groups).

Discussion

Our pre-registered, multi-site findings provide robust evidence that embodied semantic learning mechanisms are sensitive to language and cultural context. The larger effects in East Asian languages may reflect differences in gesture frequency norms or a more iconic orthographic system priming embodied representations. The dissociation between motor-semantic and semantic-only networks suggests that gesture iconicity engages grounded cognition pathways most strongly when gesture production is culturally salient.

The Bayesian hierarchical framework proved essential for handling variable sample sizes and intercultural variation. Future work should examine whether gesture-based interventions improve semantic learning in clinical populations with language deficits. International collaboration via Prolific enabled unprecedented sample diversity; we recommend similar consortium-based approaches for validating theories across human cognitive diversity.

References

  • Alibali, M. W., Kita, S., & Young, A. J. (2020). Gesture and embodiment in semantic learning. Psychological Review, 127(3), 403–427.
  • Barsalou, L. W. (2010). Grounded cognition. Annual Review of Psychology, 61, 617–645.
  • Chen, Y., Bergstrom, T., & Nakamura, A. (2025). Layer-resolved fMRI of semantic integration in multimodal learning. NeuroImage, 268, 119872.
  • Kita, S., & Özyürek, A. (2010). Cross-linguistic differences in language and gesture: Japanese vs. English. Journal of Pragmatics, 42(10), 2760–2775.
  • Okonjo, M., & Solveig, I. (2024). Cross-cultural gesture iconicity and vocabulary acquisition: A meta-analysis. Cognitive Science, 48(5), e13412.
  • Pulvermüller, F. (2018). Neural reuse of action-perception patterns: A core mechanism of cognition. Behavioral and Brain Sciences, 41, e102.
  • Shokrollahi, K., & Bergstrom, T. (2026). Bayesian hierarchical models for cross-linguistic semantic learning. Computational Linguistics, 52(1), 87–115.
  • Vieth, H. Z., Chen, Y., & Okonjo, M. (2023). Multimodal learning across twelve languages: A consortium study. eLife, 12, e84932.