Library
PubMed Central Open Access
research article
Professional
Open access

Compositional Complexity in Text and Images

Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine

Neurobiology of LanguageLast synced 8/27/2026Status: syncedPMID: 42644209 pmidDOI: 10.1162/NOL.a.271

Abstract Compositionality enables us to derive the meaning of a complex whole from the syntax and semantics of its individual parts. While compositional processing has primarily been studied in language, similar principles apply to visual scenes, in which meaning can be derived from their individual parts and the relationships between them. In this study, we therefore aimed to explore whether there is a shared neural basis for compositional processing in text and images. Based on abstract meaning representation graphs for text, commonly used to capture “who is doing what to whom”, and action graph annotations for visual scenes, we defined an analogous graph depth-based notion of compositional complexity. By conducting a functional magnetic resonance imaging experiment in which participants view text-image pairs from the Common Objects in Context-Actions data set, we aimed to identify brain activity patterns related to compositional processing through a combination of univariate and multivariate pattern analysis. Specifically, we hypothesized that there are shared brain regions—such as the pars opercularis, pars triangularis or anterior temporal lobe—involved in processing compositional complexity across text and images. Alternatively, there might exist distinct regions for processing compositional complexity in text and images, without any substantial cross-modal overlap. Our results provided no evidence in support of either hypothesis: no significant neural responses to comp

Abstract

Abstract Compositionality enables us to derive the meaning of a complex whole from the syntax and semantics of its individual parts. While compositional processing has primarily been studied in language, similar principles apply to visual scenes, in which meaning can be derived from their individual parts and the relationships between them. In this study, we therefore aimed to explore whether there is a shared neural basis for compositional processing in text and images. Based on abstract meaning representation graphs for text, commonly used to capture “who is doing what to whom”, and action graph annotations for visual scenes, we defined an analogous graph depth-based notion of compositional complexity. By conducting a functional magnetic resonance imaging experiment in which participants view text-image pairs from the Common Objects in Context-Actions data set, we aimed to identify brain activity patterns related to compositional processing through a combination of univariate and multivariate pattern analysis. Specifically, we hypothesized that there are shared brain regions—such as the pars opercularis, pars triangularis or anterior temporal lobe—involved in processing compositional complexity across text and images. Alternatively, there might exist distinct regions for processing compositional complexity in text and images, without any substantial cross-modal overlap. Our results provided no evidence in support of either hypothesis: no significant neural responses to compositional complexity were observed, in either a shared or modality-specific manner. Given the rigorous control of confounds in our study design, the operationalization of compositional complexity itself may underlie the null result.

Educational only
This information is for general education and is not medical advice. Always talk to a licensed U.S. clinician about your situation, medications, or treatment decisions.