An Analysis of Statistical Techniques Applying to Multi-Feature Similarity Comparison between Corpora
Statistical techniques applying to multi-feature similarity comparison belong to the type of goodness-of-fit test which include chi-square test, rank correlation test and Kolmogorov-Smirnov test (K-S test). Experiments show that both chi-square independence test and rank correlation test are subject to the variation of sample size. With the expansion of sample size, the former test achieves the results of significant difference and the latter achieves the results of significant correlation easily. However, both results fail to reveal the actual situation of multi-feature similarity comparison between corpora. Only K-S test, which quantifies a distance between the empirical distribution functions of two samples, can achieve the highest statistical effectiveness.
X. X. Chen et al., "An Analysis of Statistical Techniques Applying to Multi-Feature Similarity Comparison between Corpora", Applied Mechanics and Materials, Vols. 66-68, pp. 2323-2329, 2011