PySimi: a unified framework for similarity measure evaluation in spectral clustering with applications to omics data
Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine
High-throughput omics technologies generate increasingly large and complex datasets, creating a growing demand for clustering methods capable of identifying meaningful biological patterns. Spectral clustering is widely used for analyzing high-dimensional omics data, but its performance strongly depends on the construction of the similarity matrix. Although numerous similarity measures have been proposed, most existing spectral clustering tools support only a limited set of similarity construction strategies, making systematic evaluation and comparison difficult. Here, we present PySimi, an open-source Python framework for flexible similarity matrix construction and spectral clustering. PySimi integrates classical, adaptive, and neighborhood-based similarity measures within a unified and extensible framework and provides a consistent workflow for constructing, comparing, and evaluating similarity matrices. The framework also supports downstream analyses, including dimensionality reduction and visualization, and offers an interactive web application for exploratory analysis. We evaluated PySimi using multiple bulk and single-cell RNA-sequencing datasets. The results demonstrate that the choice of similarity measure can substantially influence clustering outcomes and downstream biological interpretation. While no single method consistently achieved the best performance across all datasets, adaptive and neighborhood-based approaches generally showed stronger performance than clas
Abstract
High-throughput omics technologies generate increasingly large and complex datasets, creating a growing demand for clustering methods capable of identifying meaningful biological patterns. Spectral clustering is widely used for analyzing high-dimensional omics data, but its performance strongly depends on the construction of the similarity matrix. Although numerous similarity measures have been proposed, most existing spectral clustering tools support only a limited set of similarity construction strategies, making systematic evaluation and comparison difficult. Here, we present PySimi, an open-source Python framework for flexible similarity matrix construction and spectral clustering. PySimi integrates classical, adaptive, and neighborhood-based similarity measures within a unified and extensible framework and provides a consistent workflow for constructing, comparing, and evaluating similarity matrices. The framework also supports downstream analyses, including dimensionality reduction and visualization, and offers an interactive web application for exploratory analysis. We evaluated PySimi using multiple bulk and single-cell RNA-sequencing datasets. The results demonstrate that the choice of similarity measure can substantially influence clustering outcomes and downstream biological interpretation. While no single method consistently achieved the best performance across all datasets, adaptive and neighborhood-based approaches generally showed stronger performance than classical methods. By providing a unified platform for similarity matrix construction, comparison, and evaluation, PySimi enables systematic investigation of similarity measures and facilitates their application to diverse omics datasets.
