Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
158 tokens/sec
GPT-4o
7 tokens/sec
Gemini 2.5 Pro Pro
45 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
38 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

A Novel Sampled Clustering Algorithm for Rice Phenotypic Data (2312.14920v2)

Published 22 Dec 2023 in cs.LG and cs.AI

Abstract: Phenotypic (or Physical) characteristics of plant species are commonly used to perform clustering. In one of our recent works (Shastri et al. (2021)), we used a probabilistically sampled (using pivotal sampling) and spectrally clustered algorithm to group soybean species. These techniques were used to obtain highly accurate clusterings at a reduced cost. In this work, we extend the earlier algorithm to cluster rice species. We improve the base algorithm in three ways. First, we propose a new function to build the similarity matrix in Spectral Clustering. Commonly, a natural exponential function is used for this purpose. Based upon the spectral graph theory and the involved Cheeger's inequality, we propose the use a base "a" exponential function instead. This gives a similarity matrix spectrum favorable for clustering, which we support via an eigenvalue analysis. Also, the function used to build the similarity matrix in Spectral Clustering was earlier scaled with a fixed factor (called global scaling). Based upon the idea of Zelnik-Manor and Perona (2004), we now use a factor that varies with matrix elements (called local scaling) and works better. Second, to compute the inclusion probability of a specie in the pivotal sampling algorithm, we had earlier used the notion of deviation that captured how far specie's characteristic values were from their respective base values (computed over all species). A maximum function was used before to find the base values. We now use a median function, which is more intuitive. We support this choice using a statistical analysis. Third, with experiments on 1865 rice species, we demonstrate that in terms of silhouette values, our new Sampled Spectral Clustering is 61% better than Hierarchical Clustering (currently prevalent). Also, our new algorithm is significantly faster than Hierarchical Clustering due to the involved sampling.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (20)
  1. Unequal probability sampling without replacement through a splitting method. Biometrika 85, 89–101.
  2. Cheeger’s inequality and the sparse cut problem. Lecture notes on recent advances in approximation algorithms (University of Washington).
  3. Cheeger’s inequality continued, spectral clustering. Lecture notes on design and analysis of algorithms I (University of Washington).
  4. Diversity and population structure of red rice germplasm in Bangladesh. PLoS One 13, e0196096.
  5. Genetic variability and cluster analysis for phenological traits of thai indigenous upland rice (oryza sativa l.). Indian Journal of Agricultural Research 54.
  6. Cube sampled K-prototype clustering for featured data, in: 2021 IEEE 18th India Council International Conference (INDICON), pp. 1–6.
  7. Cluster analysis in common bean genotypes (Phaseolus Vulgaris L.). Turkish Journal of Agricultural and Natural Sciences 1, 1030–1035.
  8. Multiway spectral partitioning and higher-order cheeger inequalities. Journal of the ACM (JACM) 61, 1–30.
  9. The international rice information system. A platform for meta-analysis of rice crop data. Plant Physiology 139, 637–642.
  10. Determination of the optimal number of clusters using a spectral clustering optimization. Expert systems with applications 65, 304–314.
  11. On spectral clustering: Analysis and an algorithm, in: Advances in neural information processing systems, MIT Press. pp. 849–856.
  12. Molecular and morphological characterization of indian farmers rice varieties (oryza sativa l.). Australian Journal of Crop Science 7, 923.
  13. Clustering analysis of soybean germplasm (glycine max l. merrill). The Pharma Innovation Journal 7, 781–786.
  14. Assessing genetic variation for heat tolerance in synthetic wheat lines using phenotypic data and molecular markers. Australian Journal of Crop Science 8, 515–522.
  15. Probabilistically sampled and spectrally clustered plant species using phenotypic characteristics. PeerJ 9, e11927.
  16. Vector quantized spectral clustering applied to whole genome sequences of plants. Evolutionary Bioinformatics 15, 1–7.
  17. Genetic diversity is indispensable for plant breeding to improve crops. Crop Science 61, 839–852.
  18. Variability assessment for root and drought tolerance traits and genetic diversity analysis of rice germplasm using SSR markers. Scientific reports 9, 16513.
  19. A tutorial on spectral clustering. Statistics and computing 17, 395–416.
  20. Self-tuning spectral clustering., in: Advances in neural information processing systems, pp. 1601–1608.

Summary

We haven't generated a summary for this paper yet.