Library
PubMed Central Open Access
research article
Professional
Open access

PySimi: a unified framework for similarity measure evaluation in spectral clustering with applications to omics data

Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine

Frontiers in GeneticsLast synced 8/27/2026Status: syncedPMID: 42644185 pmidDOI: 10.3389/fgene.2026.1913487

High-throughput omics technologies generate increasingly large and complex datasets, creating a growing demand for clustering methods capable of identifying meaningful biological patterns. Spectral clustering is widely used for analyzing high-dimensional omics data, but its performance strongly depends on the construction of the similarity matrix. Although numerous similarity measures have been proposed, most existing spectral clustering tools support only a limited set of similarity construction strategies, making systematic evaluation and comparison difficult. Here, we present PySimi, an open-source Python framework for flexible similarity matrix construction and spectral clustering. PySimi integrates classical, adaptive, and neighborhood-based similarity measures within a unified and extensible framework and provides a consistent workflow for constructing, comparing, and evaluating similarity matrices. The framework also supports downstream analyses, including dimensionality reduction and visualization, and offers an interactive web application for exploratory analysis. We evaluated PySimi using multiple bulk and single-cell RNA-sequencing datasets. The results demonstrate that the choice of similarity measure can substantially influence clustering outcomes and downstream biological interpretation. While no single method consistently achieved the best performance across all datasets, adaptive and neighborhood-based approaches generally showed stronger performance than clas

Abstract

High-throughput omics technologies generate increasingly large and complex datasets, creating a growing demand for clustering methods capable of identifying meaningful biological patterns. Spectral clustering is widely used for analyzing high-dimensional omics data, but its performance strongly depends on the construction of the similarity matrix. Although numerous similarity measures have been proposed, most existing spectral clustering tools support only a limited set of similarity construction strategies, making systematic evaluation and comparison difficult. Here, we present PySimi, an open-source Python framework for flexible similarity matrix construction and spectral clustering. PySimi integrates classical, adaptive, and neighborhood-based similarity measures within a unified and extensible framework and provides a consistent workflow for constructing, comparing, and evaluating similarity matrices. The framework also supports downstream analyses, including dimensionality reduction and visualization, and offers an interactive web application for exploratory analysis. We evaluated PySimi using multiple bulk and single-cell RNA-sequencing datasets. The results demonstrate that the choice of similarity measure can substantially influence clustering outcomes and downstream biological interpretation. While no single method consistently achieved the best performance across all datasets, adaptive and neighborhood-based approaches generally showed stronger performance than classical methods. By providing a unified platform for similarity matrix construction, comparison, and evaluation, PySimi enables systematic investigation of similarity measures and facilitates their application to diverse omics datasets.

Educational only
This information is for general education and is not medical advice. Always talk to a licensed U.S. clinician about your situation, medications, or treatment decisions.