Library
PubMed Central Open Access
research article
Professional
Open access

Markov reads Puškin, again: A statistical journey into the poetic world of Evgenij Onegin

Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine

PLOS OneLast synced 6/6/2026Status: syncedPMID: 42241450 pmidDOI: 10.1371/journal.pone.0350827

This study applies symbolic time series analysis and Markov modeling to explore the phonological structure of—as captured through a graphemic vowel/consonant (V/C) encoding—and one contemporary Italian translation. Using a binary encoding inspired by Markov’s original scheme, we construct minimalist probabilistic models that capture both local V/C dependencies and large–scale sequential patterns. A compact four-state Markov chain is shown to be descriptively accurate and generative, reproducing key features of the original sequences such as autocorrelation and memory depth. All findings are exploratory in nature and aim to highlight structural regularities while suggesting hypotheses about underlying narrative dynamics. The analysis reveals a marked asymmetry between the Russian and Italian texts: the original exhibits a gradual decline in memory depth, whereas the translation maintains a more uniform profile. To further investigate this divergence, we introduce phonological probes — short symbolic patterns that link surface structure to narrative-relevant cues. Tracked across the unfolding text, these probes reveal subtle connections between graphemic form and thematic development, particularly in the Russian original. By revisiting Markov’s original proposal of applying symbolic analysis to a literary text and pairing it with contemporary tools from computational statistics and data science, this study shows that even minimalist Markov models can support exploratory analysi

Abstract

This study applies symbolic time series analysis and Markov modeling to explore the phonological structure of—as captured through a graphemic vowel/consonant (V/C) encoding—and one contemporary Italian translation. Using a binary encoding inspired by Markov’s original scheme, we construct minimalist probabilistic models that capture both local V/C dependencies and large–scale sequential patterns. A compact four-state Markov chain is shown to be descriptively accurate and generative, reproducing key features of the original sequences such as autocorrelation and memory depth. All findings are exploratory in nature and aim to highlight structural regularities while suggesting hypotheses about underlying narrative dynamics. The analysis reveals a marked asymmetry between the Russian and Italian texts: the original exhibits a gradual decline in memory depth, whereas the translation maintains a more uniform profile. To further investigate this divergence, we introduce phonological probes — short symbolic patterns that link surface structure to narrative-relevant cues. Tracked across the unfolding text, these probes reveal subtle connections between graphemic form and thematic development, particularly in the Russian original. By revisiting Markov’s original proposal of applying symbolic analysis to a literary text and pairing it with contemporary tools from computational statistics and data science, this study shows that even minimalist Markov models can support exploratory analysis of complex poetic material. When complemented by a coarse layer of linguistic annotation, such models provide a general framework for comparative poetics and demonstrate that stylized structural patterns remain accessible through simple representations grounded in linguistic form.

Educational only
This information is for general education and is not medical advice. Always talk to a licensed U.S. clinician about your situation, medications, or treatment decisions.