Resultado científico

Schizophrenia

Machine learning to study schizophrenia

Schizophrenia is a series of eight genetically distinct disorders.

Most complex human diseases and disorders are the result of the interaction between multiple genetic and environmental factors. It is known that hundreds or thousands of genetic variants (SNPs) can interact in complex ways resulting in a multifaceted genetic architecture of disease, which can influence the different manifestations of these complex diseases. When we talk about the genetic architecture of inherited diseases, we refer to the number, frequency and effect size of genetic risk alleles and the way they are organised in genotypic networks.

Schizophrenia is a complex mental illness that affects 1% of the population. In Spain there are more than 500,000 known cases. However, studies in twins and relatives of schizophrenics indicate that the risk of schizophrenia is highly heritable (81%), although only 25% of the variability can be explained by specific genetic variants (SNPs) identified in genome-wide association studies (GWAS). Genetic association studies in mental disorders are plagued by weak and inconsistent findings by ignoring heterogeneity, pleiotropy and epistasis in the analyses. This has resulted in high unexplained (missing) heritability and lack of reproducibility between studies making it very difficult to translate any improvements to the clinic.

We have developed a new data-driven deep unsupervised machine learning method called PGMRA http://phop.ugr.es/fenogeno. This algorithm combines clustering techniques based on consensus, fuzzy, possibilistic, relational, optimisation, and conceptual clustering techniques in a single method in order to discover interesting groups (SNP sets, phenotypic sets, etc.) defined in different knowledge domains such as: phenotype, genotype (SNP), images, TCI inventories and interesting relationships between these groups (clusters). The most interesting features of the method are:

  • the clustering strategy does not use prior knowledge about other studies or genomic characteristics and does not consider the status of subjects (diseased or healthy) in the dataset to identify SNP sets or sets of phenotypes (unsupervised learning);
  • subjects, SNPs and/or phenotype characteristics may belong to more than one group relationship;
  • SNPs and/or phenotype characteristics may belong to more than one relationship group;
  • SNPs within a SNP set can be located anywhere in the genome;
  • the dimensionality of phenotypic characteristics is not reduced, as would be the case with principal component analysis or similar approaches, because, in phenomics, the important characteristics are not known a priori;
  • there is no predefined number of SNP sets and/or phenotype sets and/or relationships between them; many-to-many relationships between SNP sets and phenotype sets are identified in an unbiased manner regardless of the subject’s disease status (e.g. cases, controls);
  • the risk of a disease is estimated in an unbiased manner by incorporating a posteriori the subject’s status within each relationship, weighing the frequency of each type of status (e.g., diseased, relatives, controls) and mapping onto a predicted risk surface.

In addition, PGMRA can estimate the degree of statistical significance of the interactions of SNP sets and disease-associated SNP sets. In summary, PGMRA provides a quick snapshot of different domains in an interpretable way(https://www.ncbi.nlm.nih.gov/pubmed/23761451).

Outstanding results:

Info and contact: {zwir,delval} at ugr.es

Period

Feb 2015 –

Researchers

Igor ZwirCoral del Val,Rocío Romero-Zaliz, (Andalusian Research Institute on Data Science and Computational Intelligence), Javier Arnedo Fernandez, Alberto Mesa, Claude Robert Cloninger (Department of Psychiatry, Washington University School of Medicine, St. Louis, MO, USA), Terho Lehtimäki (Fimlab Laboratories, Department of Clinical Chemistry, Faculty of Medicine and Life Sciences, Finnish Cardiovascular Research Center-Tampere, University of Tampere, Tampere, Finland)

Topics
Deep Learning