Luiza Lucuța, Anne Kösem, Peter E. Keller, Anna Fiveash
Predictions of content (“what”) and timing (“when”) have been widely investigated in both music and speech research. Considering that music and speech serve distinct communicative functions, predictions of content and timing, as well as their interplay, are likely to exhibit domain-specific characteristics. However, because these domains have largely been studied in isolation, the neural bases of content–timing interactions remain poorly understood. This review sets multiple objectives. The first is to provide an overview of how the neural correlates of content and timing predictions in music and speech have been investigated through the lenses of Predictive Coding and Dynamic Attending Theory. The second is to evaluate the extent to which these frameworks account for the interaction between content and timing predictions, particularly at a neurophysiological level. The third is to highlight points of convergence between the two frameworks that are often difficult to identify due to differences in nomenclature and research design. The fourth is to propose an integrative perspective through which methodological strengths from both research traditions can be complementarily implemented to further our understanding of the neural bases of content and timing predictions in music and speech, by considering different decisions that must be taken when investigating prediction in music and speech (stimulus, analysis, and interpretation decisions). Reconciling key aspects of these theoretical frameworks across multiple levels of description (i.e., computational, algorithmic, implementational) within music and speech processing can provide insights into prediction as a fundamental aspect of cognition.
PsyArXiv
