Research Project: Pattern Recognition in High Dimensional Data
Loading...
Contributors
Funders
ID
EC.00056
Authors
Ceyhan, Elvan
Faculty Member
Publications
Edge density of new graph types based on a random digraph family
(Elsevier, 2016) Ceyhan, Elvan; Department of Mathematics; Yes; College of Sciences
We consider two types of graphs based on a family of proximity catch digraphs (PCDs) and study their edge density. in particular, the PCDs we use are a parameterized digraph family called proportional-edge (PE) PCDs and the two associated graph types are the "underlying graphs" and the newly introduced "reflexivity graphs" based on the PE-PCDs. these graphs are extensions of random geometric graphs where distance is replaced with a dissimilarity measure and the threshold is not fixed but depends on the location of the points. PCDs and the associated graphs are constructed based on data points from two classes, say X and y, where one class (say class X) forms the vertices of the PCD and the Delaunay tessellation of the other class (i.e., class y) yields the (Delaunay) cells which serve as the support of class X points. We demonstrate that edge density of these graphs is a U-statistic, hence obtain the asymptotic normality of it for data from any distribution that satisfies mild regulatory conditions. the rate of convergence to asymptotic normality is sharper for the edge density of the reflexivity and underlying graphs compared to the arc density of the PE-PCDs. for uniform data in Euclidean plane where Delaunay cells are triangles, we demonstrate that the distribution of the edge density is geometry invariant (i.e., independent of the shape of the triangular support). We compute the explicit forms of the asymptotic normal distribution for uniform data in one Delaunay triangle in the Euclidean plane utilizing this geometry invariance property. We also provide various versions of edge density in the multiple triangle case. the approach presented here can also be extended for application to data in higher dimensions.
Simulation and characterization of multi-class spatial patterns from stochastic point processes of randomness, clustering and regularity
(Springer, 2014) Ceyhan, Elvan; Department of Mathematics; Yes; College of Sciences
Spatial pattern analysis of data from multiple classes (i.e., multi-class data) has important implications. We investigate the resulting patterns when classes are generated from various spatial point processes. Our null pattern is that the nearest neighbor probabilities being proportional to class frequencies in the multi-class setting. In the two-class case, the deviations are mainly in two opposite directions, namely, segregation and association of the classes. But for three or more classes, the classes might exhibit mixed patterns, in which one pair exhibiting segregation, while another pair exhibiting association or complete spatial randomness independence. To detect deviations from the null case, we employ tests based on nearest neighbor contingency tables (NNCTs), as NNCT methods can provide an omnibus test and post-hoc tests after a significant omnibus test in a multi-class setting. In particular, for analyzing these multi-class patterns (mixed or not), we use an omnibus overall test based on NNCTs. After the overall test, the pairwise interactions are analyzed by the post-hoc cell-specific tests based on NNCTs. We propose various parameterizations of the segregation and association alternatives, list some appealing properties of these patterns, and propose three processes for the two-class association pattern. We also consider various clustering and regularity patterns to determine which one(s) cause segregation from or association with a class from a homogeneous Poisson process and from other processes as well. We perform an extensive Monte Carlo simulation study to investigate the newly proposed association patterns and to understand which stochastic processes might result in segregation or association. The methodology is illustrated on two real life data sets from plant ecology.
Testing spatial symmetry using contingency tables based on nearest neighbor relations
(Hindawi, 2014) Ceyhan, Elvan; Department of Mathematics; Yes; College of Sciences
We consider two types of spatial symmetry, namely, symmetry in the mixed or shared nearest neighbor (NN) structures. We use Pielou’s and Dixon’s symmetry tests which are defined using contingency tables based on the NN relationships between the data points. We generalize these tests to multiple classes and demonstrate that both the asymptotic and exact versions of Pielou’s first type of symmetry test are extremely conservative in rejecting symmetry in the mixed NN structure and hence should be avoided or only the Monte Carlo randomized version should be used. Under RL, we derive the asymptotic distribution for Dixon’s symmetry test and also observe that the usual independence test seems to be appropriate for Pielou’s second type of test. Moreover, we apply variants of Fisher’s exact test on the shared NN contingency table for Pielou’s second test and determine the most appropriate version for our setting. We also consider pairwise and one-versus-rest type tests in post hoc analysis after a significant overall symmetry test. We investigate the asymptotic properties of the tests, prove their consistency under appropriate null hypotheses, and investigate finite sample performance of them by extensive Monte Carlo simulations. The methods are illustrated on a real-life ecological data set.
Segregation indices for disease clustering
(Wiley, 2014) Ceyhan, Elvan; Department of Mathematics; Yes; College of Sciences
Spatial clustering has important implications in various fields. In particular, disease clustering is of major public concern in epidemiology. In this article, we propose the use of two distance-based segregation indices to test the significance of disease clustering among subjects whose locations are from a homogeneous or an inhomogeneous population. We derive the asymptotic distributions of the segregation indices and compare them with other distance-based disease clustering tests in terms of empirical size and power by extensive Monte Carlo simulations. The null pattern we consider is the random labeling (RL) of cases and controls to the given locations. Along this line, we investigate the sensitivity of the size of these tests to the underlying background pattern (e.g., clustered or homogenous) on which the RL is applied, the level of clustering and number of clusters, or to differences in relative abundances of the classes. We demonstrate that differences in relative abundances have the highest influence on the empirical sizes of the tests. We also propose various non-RL patterns as alternatives to the RL pattern and assess the empirical power performances of the tests under these alternatives. We observe that the empirical size of one of the indices is more robust to the differences in relative abundances, and this index performs comparable with the best performers in literature in terms of power. We illustrate the methods on two real-life examples from epidemiology.
Comparison of relative density of two random geometric digraph families in testing spatial clustering
(Springer, 2014) Ceyhan, Elvan; Department of Mathematics; Yes; College of Sciences
We compare the performance of relative densities of two parameterized random geometric digraph families called proximity catch digraphs (PCDs) in testing bivariate spatial patterns. These PCD families are proportional edge (PE) and central similarity (CS) PCDs and are defined with proximity regions based on relative positions of data points from two classes. The relative densities of these PCDs were previously used as statistics for testing segregation and association patterns against complete spatial randomness. The relative density of a digraph, D, with n vertices (i.e., with order n) represents the ratio of the number of arcs in D to the number of arcs in the complete symmetric digraph of the same order. When scaled properly, the relative density of a PCD is a U-statistic; hence, it has asymptotic normality by the standard central limit theory of U-statistics. The PE- and CS-PCDs are defined with an expansion parameter that determines the size or measure of the associated proximity regions. In this article, we extend the distribution of the relative density of CS-PCDs for expansion parameter being larger than one, and compare finite sample performance of the tests by Monte Carlo simulations and asymptotic performance by Pitman asymptotic efficiency. We find the optimal expansion parameters of the PCDs for testing each alternative in finite samples and in the limit as the sample size tending to infinity. As a result of our comparisons, we demonstrate that in terms of empirical power (i.e., for finite samples) relative density of CS-PCD has better performance (which occurs for expansion parameter values larger than one) for the segregation alternative, while relative density of PE-PCD has better performance for the association alternative. The methods are also illustrated in a real-life data set from plant ecology.
