Data Mining of Online Genealogy Datasets for Revealing Lifespan Patterns in Human Population
Fichiers
Mots-clés
Citer ce document
Michael Fire et Yuval Elovici, “Data Mining of Online Genealogy Datasets for Revealing Lifespan Patterns in Human Population,” Portail documentaire EnJeu[x], consulté le 2 novembre 2024, https://collections.enfance-jeunesse.fr/items/show/1462.
Métadonnées complètes
Statut du document | Public |
Titre | Data Mining of Online Genealogy Datasets for Revealing Lifespan Patterns in Human Population |
Créateur | Michael Fire |
Créateur | Yuval Elovici |
Date | 2013 |
Type | Journal Article |
Author | Michael Fire |
Author | Yuval Elovici |
Type de contenu | Journal Article |
Abstract Note | Online genealogy datasets contain extensive information about millions of people and their past and present family connections. This vast amount of data can assist in identifying various patterns in human population. In this study, we present methods and algorithms which can assist in identifying variations in lifespan distributions of human population in the past centuries, in detecting social and genetic features which correlate with human lifespan, and in constructing predictive models of human lifespan based on various features which can easily be extracted from genealogy datasets. We have evaluated the presented methods and algorithms on a large online genealogy dataset with over a million profiles and over 9 million connections, all of which were collected from the WikiTree website. Our findings indicate that significant but small positive correlations exist between the parents' lifespan and their children's lifespan. Additionally, we found slightly higher and significant correlations between the lifespans of spouses. We also discovered a very small positive and significant correlation between longevity and reproductive success in males, and a small and significant negative correlation between longevity and reproductive success in females. Moreover, our machine learning algorithms presented better than random classification results in predicting which people who outlive the age of 50 will also outlive the age of 80. We believe that this study will be the first of many studies which utilize the wealth of data on human populations, existing in online genealogy datasets, to better understand factors which influence human lifespan. Understanding these factors can assist scientists in providing solutions for successful aging. |
Access Date | 2015-07-07 12:21:46 |
Date | 2013 |
Extra | arXiv: 1311.4276 |
Library Catalog | arXiv.org |
Pages | 1-24 |
Publication Title | arXiv |
Titre | Data Mining of Online Genealogy Datasets for Revealing Lifespan Patterns in Human Population |
URL | http://arxiv.org/abs/1311.4276 |
Attachment Title | arXiv.org Snapshot |
Attachment Title | arXiv:1311.4276 PDF |
Attachment URL | [No URL] |
Attachment URL | [No URL] |