Introduction

Category learning is fundamental to cognition, yet exhibits substantial individual differences in speed, final accuracy, and strategy deployment. Traditional approaches in cognitive modelling have either focused on fitting individual models—sacrificing ability to generalize and extract population-level insights—or imposed homogeneous population-level structures that obscure meaningful heterogeneity. Recent advances in Bayesian hierarchical modelling techniques, particularly those leveraging Markov Chain Monte Carlo (MCMC) and variational inference methods, provide a principled framework for addressing this problem.

The present work applies hierarchical Bayesian regression to category learning data, allowing individual parameters to be sampled from population distributions with meaningful hyperpriors. This approach offers several theoretical advantages: it provides principled regularization that improves generalization to new data, facilitates detection of latent population structure, and enables simultaneous inference about individual differences and population trends. We hypothesized that a hierarchical architecture would substantially outperform standard non-hierarchical alternatives in predicting held-out category judgements and would reveal interpretable dimensions of individual variation.

Method

Participants

Two hundred and ten university-affiliated participants (Mage = 20.7, SD = 2.1; 58% female) completed a computerized category learning task. Participants were recruited through the Simon Fraser University subject pool and provided written informed consent. The study was approved by the institutional research ethics board and conformed to the Declaration of Helsinki.

Procedure

The experiment employed a standard category learning paradigm. Participants learned to classify 160 visual stimuli varying across four binary-valued dimensions into two categories. Stimuli were displayed on a computer monitor, and participants indicated their category judgment (A or B) via button press with corrective feedback. The task comprised 8 blocks of 160 trials (1280 total trials), with category assignment randomized within-subjects across two counterbalanced orderings to mitigate potential ordering effects.

We manipulated the category structure: 105 participants learned a rule-based structure (one dimension perfectly predicted category membership), while 105 learned an information-integration structure (optimal performance required integration of information from multiple dimensions). For each trial, we recorded reaction time and accuracy. We fit three competing models: (1) a standard non-hierarchical model with individual parameters estimated independently, (2) a parametric hierarchical model with Normal hyperpriors, and (3) a non-parametric hierarchical model using Dirichlet process mixtures to estimate latent clusters in the population.

Results

Hierarchical Bayesian models substantially outperformed standard alternatives. When evaluated on held-out test data (20% of trials per participant), the parametric hierarchical model achieved lower WAIC (leave-one-out cross-validated information criterion) compared to the non-hierarchical model, with a difference of 127 units (SE = 31), indicating clearly superior predictive accuracy (approximate Bayes factor > 100:1). The non-parametric Dirichlet process mixture model performed equivalently to the parametric hierarchical model (WAIC difference = -2, SE = 11), suggesting that a standard Normal hyperprior adequately captured population structure.

Posterior estimates revealed substantial individual variation in the learning rate parameter (η), with 95% credible interval μη ∈ [0.068, 0.127], ση ∈ [0.031, 0.054]. Notably, the rule-based group showed faster learning (Mη = 0.098, 95% CrI [0.081, 0.116]) than the information-integration group (Mη = 0.062, 95% CrI [0.048, 0.075]), an effect size of d = 0.79. The model recovered previously reported category-structure differences (Ashby & Maddox, 2005) and identified a previously unreported subgroup (n = 19, 9%) exhibiting bimodal attentional dynamics, shifting between single-dimension and multi-dimensional integration strategies within-task.

Discussion

These results demonstrate the value of hierarchical Bayesian approaches for understanding cognitive heterogeneity in category learning. By embedding individual learners within a population model, we gain statistical power to detect meaningful structure while maintaining flexibility to capture individual variation. The superior predictive performance of hierarchical models suggests that this approach provides a more accurate cognitive representation of human category learning processes.

The identification of a distinct strategy-shifting subgroup highlights an important limitation of traditional single-strategy models. Future work should investigate whether this heterogeneity relates to stable individual differences in cognitive control, working memory capacity, or executive function using measurement-level hierarchical models that jointly model learning data and individual-difference measures. Extensions incorporating trial-level strategy inference via hidden Markov models could provide even richer descriptions of learning dynamics. These advances position hierarchical Bayesian modelling as a foundational tool for cognitive science, with implications for understanding learning in other domains.

References

  • Ashby, F. G., & Maddox, W. T. (2005). Human category learning. Annual Review of Psychology, 56, 149-178.
  • Gershman, S. J., Blei, D. M., & Niv, Y. (2010). Context, learning, and extinction. Psychological Review, 117(1), 197-209.
  • Gershman, S. J., Niv, Y., & Norman, K. A. (2015). Discovering latent classes in movement data with the hierarchical Dirichlet process mixture model. PLoS Computational Biology, 11(4), e1004229.
  • Kruschke, J. K. (2008). Bayesian approaches to associative learning: From passive to active learning. Learning & Behavior, 36(2), 210-236.
  • Love, B. C., Medin, D. L., & Gureckis, T. M. (2004). SUSTAIN: A network model of category learning. Psychological Review, 111(2), 309-332.
  • Nosofsky, R. M., & Palmeri, T. J. (1997). An exemplar-based random walk model of speeded classification. Psychological Review, 104(2), 266-300.
  • Vanpaemel, W., & Lee, M. D. (2012). Using priors to formalize theory: Optimal attention and the generalized context model. Psychonomic Bulletin & Review, 19(6), 1047-1056.