-
openalex
Fred W. Allendorf, Paul A. Hohenlohe, Gordon Luikart
2010-09-17
置信度 0.72
GenomicsBiologyConservation geneticsConservation biologyBiodiversity
-
The modulation of DNA-protein interactions by methylation of protein-binding sites in DNA and the occurrence in genomic imprinting, X chromosome inactivation, and fragile X syndrome of different methylation patterns in DNA of different chromosomal origin have …
openalex
Marianne Frommer, Louise E. McDonald, David Millar, Christina M. Collis 等
1992-03-01
置信度 0.72
DNA methylation5-MethylcytosineBiologygenomic DNADNA
-
K. Edwards, C. Johnstone, C. Thompson; A simple and rapid method for the preparation of plant genomic DNA for PCR analysis, Nucleic Acids Research, Volume
openalex
Keith J. Edwards, CN Johnstone, C. Thompson
1991-01-01
置信度 0.72
BiologySection (typography)Library sciencePlant scienceSimple (philosophy)
-
ABSTRACT Despite important strides in marker technologies, the use of marker‐assisted selection has stagnated for the improvement of quantitative traits. Biparental mating designs for the detection of loci affecting these traits (quantitative trait loci [QTL])…
openalex
Elliot L. Heffner, Mark E. Sorrells, Jean‐Luc Jannink
2009-01-01
置信度 0.72
BiologyQuantitative trait locusGenomic selectionSelection (genetic algorithm)Heritability
-
A very simple, fast, universally applicable and reproducible method to extract high quality megabase genomic DNA from different organisms is described. We applied the same method to extract high quality complex genomic DNA from different tissues (wheat, barley…
openalex
S. Aljanabi
1997-11-15
置信度 0.72
Biologygenomic DNADNA extractionDNACloning (programming)
-
openalex
Mathew J. Garnett, Elena J. Edelman, Sonja J. Heidorn, Chris Greenman 等
2012-03-01
置信度 0.72
CancerPharmacogenomicsBiologyCancer cellBiomarker
-
MOTIVATION: We introduce GMAP, a standalone program for mapping and aligning cDNA sequences to a genome. The program maps and aligns a single sequence with minimal startup time and memory requirements, and provides fast batch processing of large sequence sets.…
openalex
Ting‐Di Wu, C.K. Watanabe
2005-02-22
置信度 0.72
Computer scienceComputational biologyProbabilistic logicSequence (biology)Genetics
-
openalex
Alan L. Harvey, RuAngelie Edrada‐Ebel, Ronald J. Quinn
2015-01-23
置信度 0.72
Drug discoveryNatural productComputational biologyNatural (archaeology)Natural Product Research
-
Liver cancer has the second highest worldwide cancer mortality rate and has limited therapeutic options. We analyzed 363 hepatocellular carcinoma (HCC) cases by whole-exome sequencing and DNA copy number analyses, and we analyzed 196 HCC cases by DNA methylati…
openalex
Adrian Ally, Miruna Balasundaram, Rebecca Carlsen, Eric Chuah 等
2017-06-01
置信度 0.72
BiologyHepatocellular carcinomaCancer researchExome sequencingDNA methylation
-
Oesophageal cancers are prominent worldwide; however, there are few targeted therapies and survival rates for these cancers remain dismal. Here we performed a comprehensive molecular analysis of 164 carcinomas of the oesophagus derived from Western and Eastern…
openalex
2017-01-01
置信度 0.72
AdenocarcinomaCancer researchEsophagusPathologyCancer
-
Genomics is a Big Data science and is going to get much bigger, very soon, but it is not known whether the needs of genomics will exceed other Big Data domains. Projecting to the year 2025, we compared genomics with three other major generators of Big Data: as…
openalex
Zachary Stephens, Skylar Y. Lee, Faraz Faghri, Roy H. Campbell 等
2015-07-07
置信度 0.72
Big dataGenomicsData scienceBiologyComputational genomics
-
We performed integrated genomic, transcriptomic, and proteomic profiling of 150 pancreatic ductal adenocarcinoma (PDAC) specimens, including samples with characteristic low neoplastic cellularity. Deep whole-exome sequencing revealed recurrent somatic mutation…
openalex
Benjamin J. Raphael, Ralph H. Hruban, Andrew J. Aguirre, Richard A. Moffitt 等
2017-08-01
置信度 0.72
GNAS complex locusKRASCDKN2ACancer researchBiology
-
This FAIRsharing record describes: The National Cancer Institute’s (NCI’s) Genomic Data Commons (GDC) is a data sharing The National Cancer Institute's Genomic Data Commons (GDC) was created to promote precision medicine in oncology. It supports the import and…
openalex
Robert L. Grossman, Allison P. Heath, Vincent Ferretti, Harold Varmus 等
2016-09-21
置信度 0.72
MedicineCancerMEDLINEInternal medicineBiology
-
We developed bulked segregant analysis as a method for rapidly identifying markers linked to any specific gene or genomic region. Two bulked DNA samples are generated from a segregating population from a single cross. Each pool, or bulk, contains individuals t…
openalex
Richard W. Michelmore, Ilan Paran, Rick Kesseli
1991-11-01
置信度 0.72
Bulked segregant analysisGeneticsBiologyLocus (genetics)Population
-
Marine stickleback fish have colonized and adapted to thousands of streams and lakes formed since the last ice age, providing an exceptional opportunity to characterize genomic mechanisms underlying repeated ecological adaptation in nature. Here we develop a h…
openalex
Felicity C. Jones, Manfred Grabherr, Yingguang Frank Chan, Pamela Russell 等
2012-04-01
置信度 0.72
BiologyEvolutionary biologyAdaptive evolutionBasis (linear algebra)Gene duplication
-
The emergence of telomere-to-telomere (T2T) genome assemblies has opened new avenues for comparative genomics, yet effective tokenization strategies for genomic sequences remain underexplored. In this pilot study, we apply Byte Pair Encoding (BPE) to nine T2T …
arxiv
Marina Popova, Iaroslav Chelombitko, Aleksey Komissarov
2025-05-13T19:27:58Z
置信度 0.78
q-bio.GNcs.AI
-
In comparative genomics, the rearrangement distance between two genomes (equal the minimal number of genome rearrangements required to transform them into a single genome) is often used for measuring their evolutionary remoteness. Generalization of this measur…
arxiv
Sergey Aganezov,, Max A. Alekseyev
2012-08-01T08:20:16Z
置信度 0.78
q-bio.GNcs.DMq-bio.PE
-
Long-range dependencies are critical for understanding genomic structure and function, yet most conventional methods struggle with them. Widely adopted transformer-based models, while excelling at short-context tasks, are limited by the attention module's quad…
arxiv
Matvei Popov, Aymen Kallala, Anirudha Ramesh, Narimane Hennouni 等
2025-04-07T18:34:06Z
置信度 0.78
q-bio.GNcs.CVcs.LG
-
Background: Recent studies in a growing number of organisms have yielded accumulating evidence that a significant portion of the non-coding region in the genome is transcribed. We address this issue in the yeast Saccharomyces cerevisiae. Results: Taking into a…
arxiv
Moshe Havilio, Erez Y. Levanon, Galia Lerman, Martin Kupiec 等
2005-06-16T13:28:40Z
置信度 0.78
q-bio.GN
-
Rare diseases are collectively common, affecting approximately one in twenty individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in DNA sequencing, development of new computational and experimental a…
arxiv
Moez Dawood, Ben Heavner, Marsha M. Wheeler, Rachel A. Ungar 等
2024-12-18T21:11:13Z
置信度 0.78
q-bio.OT
-
Predicting the response of a specific cancer to a therapy is a major goal in modern oncology that should ultimately lead to a personalised treatment. High-throughput screenings of potentially active compounds against a panel of genomically heterogeneous cancer…
arxiv
Michael P. Menden, Francesco Iorio, Mathew Garnett, Ultan McDermott 等
2012-12-03T19:38:09Z
置信度 0.78
q-bio.GNcs.CEcs.LGq-bio.CB
-
The Global Alliance for Genomics and Health (GA4GH) Beacon protocol lets researchers ask whether a genomic variant has been observed in a participating cohort and receive aggregate variant-level counts. As Beacon networks grow, two privacy risks remain: host i…
arxiv
Christos Galanopoulos, Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
2026-06-18T14:49:13Z
置信度 0.78
q-bio.GNcs.CR
-
The study of genome rearrangement has many flavours, but they all are somehow tied to edit distances on variations of a multi-graph called the breakpoint graph. We study a weighted 2-break distance on Eulerian 2-edge-colored multi-graphs, which generalizes wei…
arxiv
Pijus Simonaitis, Annie Chateau, Krister M. Swenson
2018-02-21T11:12:45Z
置信度 0.78
cs.DScs.DMmath.COq-bio.GN
-
The domestication and subsequent selection by humans to create breeds of cattle undoubtedly altered the patterning of variation within their genomes. Strong selection to fix advantageous large-effect mutations underlying domesticability, breed characteristics …
arxiv
Holly R. Ramey, Jared E. Decker, Stephanie D. McKay, Megan M. Rolf 等
2012-12-11T04:50:02Z
置信度 0.78
q-bio.GNq-bio.PE
-
The COVID-19 crisis has demonstrated the potential of cutting-edge genomics research. However, privacy of these sensitive pieces of information is an area of significant concern for genomics researchers. The current security models makes it difficult to create…
arxiv
David Reddick, Justin Presley, F. Alex Feltus, Susmit Shannigrahi
2022-04-10T03:42:06Z
置信度 0.78
cs.CRcs.NI
-
Human genomic data carry unique information about an individual and offer unprecedented opportunities for healthcare. The clinical interpretations derived from large genomic datasets can greatly improve healthcare and pave the way for personalized medicine. Sh…
arxiv
Mohammed Alghazwi, Fatih Turkmen, Joeri van der Velde, Dimka Karastoyanova
2021-11-19T10:59:32Z
置信度 0.78
cs.CR
-
Background: Diversity estimates in cultivated plants provide a rationale for conservation strategies and support the selection of starting material for breeding programs. Diversity measures applied to crops usually have been limited to the assessment of genome…
arxiv
H. Laurentin, A. Ratzinger, P. Karlovsky
2014-11-01T11:06:36Z
置信度 0.78
q-bio.PEq-bio.GN
-
Two distinct families of pan-primate endogenous retroviruses, namely HERVL and HERVH, infected primates germline, colonized host genomes, and evolved into the global retroviral genomic regulatory dominion (GRD) operating during human embryogenesis (HE). HE ret…
arxiv
Gennadi Glinsky
2023-10-25T08:04:58Z
置信度 0.78
q-bio.GNq-bio.MNq-bio.NCq-bio.PEq-bio.TO
-
Saccharomyces cerevisiae is one of the premier model systems for studying the genomics and evolution of transposable elements. The availability of the S. cerevisiae genome led to many insights into its five known transposable element families (Ty1-Ty5) in the …
arxiv
Martin Carr, Douda Bensasson, Casey M. Bergman
2012-09-01T19:56:01Z
置信度 0.78
q-bio.PEq-bio.GN
-
Genomic language models (gLMs) have transformed computational biology, achieving state-of-the-art performance across genomic tasks. Yet a fundamental question threatens the foundation of this success: do these models learn the mechanistic principles governing …
arxiv
Bryan Cheng, Jasper Zhang
2026-04-08T00:56:26Z
置信度 0.78
q-bio.GN
-
Wolbachia are maternally-inherited symbiotic bacteria commonly found in arthropods, which are able to manipulate the reproduction of their host in order to maximise their transmission. Here we use whole genome resequencing data from 290 lines of Drosophila mel…
arxiv
Mark F. Richardson, Lucy A. Weinert, John J. Welch, Raquel S. Linheiro 等
2012-05-25T21:29:14Z
置信度 0.78
q-bio.PEq-bio.GN
-
Background: In the marine environment, where there are few absolute physical barriers, contemporary contact between previously isolated species can occur across great distances, and in some cases, may be inter-oceanic. [..] in the minke whale species complex […
arxiv
Ketil Malde, Bjørghild B. Seliussen, María Quintela, Geir Dahle 等
2018-09-06T13:49:32Z
置信度 0.78
q-bio.GN
-
Genome-wide association studies (GWAS) provide a means of examining the common genetic variation underlying a range of traits and disorders. In addition, it is hoped that GWAS may provide a means of differentiating affected from unaffected individuals. This ha…
arxiv
Carlos Pinto, Michael Gill, Schizophrenia Working Group of the Psychiatric Genomics Consortium, Elizabeth A. Heron
2019-11-20T16:08:27Z
置信度 0.78
q-bio.GN
-
The study of high-throughput genomic profiles from a pharmacogenomics viewpoint has provided unprecedented insights into the oncogenic features modulating drug response. A recent screening of ~1,000 cancer cell lines to a collection of anti-cancer drugs illumi…
arxiv
Yu-Chiao Chiu, Hung-I Harry Chen, Tinghe Zhang, Songyao Zhang 等
2018-05-20T04:27:09Z
置信度 0.78
stat.MLcs.LGq-bio.GN
-
Large Genomic Foundation Models have recently achieved remarkable results and in-vivo translation capabilities. However these models quickly grow to over a few Billion of parameters and are expensive to run when compute is limited. To overcome this challenge, …
arxiv
Rasched Haidari, Sam Martin, Maxime Allard
2026-03-27T14:07:13Z
置信度 0.78
cs.LGcs.AI
-
We present a nonparametric Bayesian method for disease subtype discovery in multi-dimensional cancer data. Our method can simultaneously analyse a wide range of data types, allowing for both agreement and disagreement between their underlying clustering struct…
arxiv
Richard S. Savage, Zoubin Ghahramani, Jim E. Griffin, Paul Kirk 等
2013-04-12T09:06:45Z
置信度 0.78
q-bio.GNstat.ML
-
Generating the hash values of short subsequences, called seeds, enables quickly identifying similarities between genomic sequences by matching seeds with a single lookup of their hash values. However, these hash values can be used only for finding exact-matchi…
arxiv
Can Firtina, Jisung Park, Mohammed Alser, Jeremie S. Kim 等
2021-12-16T08:18:00Z
置信度 0.78
q-bio.GN
-
How to compare whole genome sequences at large scale has not been achieved via conventional methods based on pair-wisely base-to-base comparison; nevertheless, no attention was paid to handle in-one-sitting a number of genomes crossing genetic category (chromo…
arxiv
Yuncan Ai, Hannan Ai, Fanmei Meng, Lei Zhao
2013-03-10T02:47:09Z
置信度 0.78
q-bio.GNcs.CEmath.NA
-
Searching for similar genomic sequences is an essential and fundamental step in biomedical research and an overwhelming majority of genomic analyses. State-of-the-art computational methods performing such comparisons fail to cope with the exponential growth of…
arxiv
Mohammed Alser, Julien Eudine, Onur Mutlu
2022-11-15T14:09:39Z
置信度 0.78
cs.DSq-bio.GNq-bio.QM
-
The immense increase in the generation of genomic scale data poses an unmet analytical challenge, due to a lack of established methodology with the required flexibility and power. We propose a first principled approach to statistical analysis of sequence-level…
arxiv
Geir K. Sandve, Sveinung Gundersen, Halfdan Rydbeck, Ingrid K. Glad 等
2011-01-25T18:57:24Z
置信度 0.78
q-bio.GN
-
In biomedical research, validation of a new scientific discovery is tied to the reproducibility of its experimental results. However, in genomics, the definition and implementation of reproducibility still remain imprecise. Here, we argue that genomic reproduc…
arxiv
Pelin Icer Baykal, Paweł P. Łabaj, Florian Markowetz, Lynn M. Schriml 等
2023-08-18T13:43:36Z
置信度 0.78
q-bio.GN
-
Genome data are crucial in modern medicine, offering significant potential for diagnosis and treatment. Thanks to technological advancements, many millions of healthy and diseased genomes have already been sequenced; however, obtaining the most suitable data f…
arxiv
Teddy Lazebnik, Liron Simon-Keren
2023-05-01T07:16:40Z
置信度 0.78
q-bio.GNcs.LGcs.NE
-
Advancements in genomic research such as high-throughput sequencing techniques have driven modern genomic studies into "big data" disciplines. This data explosion is constantly challenging conventional methods used in genomics. In parallel with the urgent dema…
arxiv
Tianwei Yue, Yuanxin Wang, Longxiang Zhang, Chunming Gu 等
2018-02-02T12:50:25Z
置信度 0.78
q-bio.GNcs.LG
-
In recent years, Reinforcement Learning (RL) has emerged as a powerful tool for solving a wide range of problems, including decision-making and genomics. The exponential growth of raw genomic data over the past two decades has exceeded the capacity of manual a…
arxiv
Mohsen Karami, Khadijeh, Jahanian, Roohallah Alizadehsani 等
2023-02-26T08:43:08Z
置信度 0.78
q-bio.GNcs.AIcs.LGq-bio.QM
-
To support comparative genomics, population genetics, and medical genetics, we propose that a reference genome should come with a scheme for mapping each base in any DNA string to a position in that reference genome. We refer to a collection of one or more ref…
arxiv
Benedict Paten, Adam Novak, David Haussler
2014-04-20T04:48:24Z
置信度 0.78
q-bio.GN
-
The ultimate secret of all lives on earth is hidden in their genomes -- a totality of DNA sequences. We currently know the whole genome sequence of many organisms, while our understanding of the genome architecture on a systematic level remains rudimentary. Ap…
arxiv
Yanying Wu
2020-09-06T21:36:58Z
置信度 0.78
q-bio.GNmath.CT
-
This paper will argue that one of the biggest challenges for livestock genomics is to make whole-genome sequencing and functional genomics applicable to breeding practice. It discusses potential explanations for why it is so difficult to consistently improve t…
arxiv
M. Johnsson
2023-02-02T14:55:51Z
置信度 0.78
q-bio.GN
-
Recent advances in high-throughput genomics technologies have resulted in the sequencing of large numbers of (near) complete genomes. These genome sequences are being mined for important functional elements, such as genes. They are also being compared and cont…
arxiv
Lior Pachter
2006-12-25T21:20:34Z
置信度 0.78
q-bio.GNq-bio.QM
-
In the rapidly evolving landscape of genomics, deep learning has emerged as a useful tool for tackling complex computational challenges. This review focuses on the transformative role of Large Language Models (LLMs), which are mostly based on the transformer a…
arxiv
Micaela E. Consens, Cameron Dufault, Michael Wainberg, Duncan Forster 等
2023-11-13T02:13:58Z
置信度 0.78
q-bio.GNcs.LG
-
Low-cost, high-throughput DNA and RNA sequencing (HTS) data is the backbone of the life sciences. Genome sequencing is now becoming a part of Predictive, Preventive, Personalized, and Participatory (termed 'P4') medicine. All genomic data are currently process…
arxiv
William Andrew Simon, Leonid Yavits, Konstantina Koliogeorgi, Yann Falevoz 等
2025-05-31T15:13:06Z
置信度 0.78
q-bio.GNcs.AR
-
This paper reviews strategies for solving problems encountered when analyzing large genomic data sets and describes the implementation of those strategies in R by packages from the Bioconductor project. We treat the scalable processing, summarization and visua…
arxiv
Michael Lawrence, Martin Morgan
2014-09-09T10:47:37Z
置信度 0.78
q-bio.GNcs.DC
-
Being able to store and transmit human genome sequences is an important part in genomic research and industrial applications. The complete human genome has 3.1 billion base pairs (haploid), and storing the entire genome naively takes about 3 GB, which is infea…
arxiv
Anirduddha Laud, Gaurav Menghani, Madhava Keralapura
2020-10-05T19:00:23Z
置信度 0.78
q-bio.GN
-
In recent years, the sequencing, assembling and annotation of prokaryotic genomes has become increasingly easy and cheap. Thus it becomes increasingly feasible and interesting to perform comparative genomics analyses of new genomes to those of related organism…
arxiv
Giorgio Gonnella
2023-02-06T16:47:50Z
置信度 0.78
q-bio.GN
-
Analyzing a functional genomics experiment, such as ATAC-, ChIP- or RNA-sequencing, requires reference data including a genome assembly and gene annotation. These resources can generally be retrieved from different organizations and in different versions. Most…
arxiv
Siebren Frölich, Maarten van der Sande, Tilman Schäfers, Simon J. van Heeringen
2022-09-02T07:00:49Z
置信度 0.78
q-bio.GN
-
We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. Our core contribution is to simplify and unify the workflow for genomic model development: data collection, model tuning, infere…
arxiv
Weimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin 等
2025-09-13T03:31:55Z
置信度 0.78
q-bio.GNcs.LG
-
An approach for approximately calculating the number of genes in a genome is presented, which takes into account the average protein length expected for the species. A number of virus, bacterial and eukaryotic genomes are scrutinized. Genome figures are presen…
arxiv
N. S. Santos-Magalhaes, H. M. de Oliveira
2015-02-12T17:03:36Z
置信度 0.78
q-bio.GNstat.AP
-
The effective visualization of genomic data is crucial for exploring and interpreting complex relationships within and across genes and genomes. Despite advances in developing dedicated bioinformatics software, common visualization tools often fail to efficien…
arxiv
Thomas Hackl, Markus Ankenbrand, Bart van Adrichem, David Wilkins 等
2024-11-05T11:39:53Z
置信度 0.78
q-bio.GN
-
Zimin words are words that have the same prefix and suffix. They are unavoidable patterns, with all sufficiently large strings encompassing them. Here, we examine for the first time the presence of k-mers not containing any Zimin patterns, defined hereafter as…
arxiv
Nikol Chantzi, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
2024-10-16T20:00:33Z
置信度 0.78
q-bio.GN
-
Data visualization is a fundamental tool in genomics research, enabling the exploration, interpretation, and communication of complex genomic features. While machine learning models show promise for transforming data into insightful visualizations, current mod…
arxiv
Skylar Sargent Walters, Arthea Valderrama, Thomas C. Smits, David Kouřil 等
2025-09-19T21:29:13Z
置信度 0.78
q-bio.GNcs.AIcs.HCcs.LG
-
In recent years, cancer genome sequencing and other high-throughput studies of cancer genomes have generated many notable discoveries. In this review, Novel genomic alteration mechanisms, such as chromothripsis (chromosomal crisis) and kataegis (mutation storm…
arxiv
Edwin Wang
2014-09-10T21:45:52Z
置信度 0.78
q-bio.GNq-bio.MN
-
Shannon information (SI) and its special case, divergence, are defined for a DNA sequence in terms of probabilities of chemical words in the sequence and are computed for a set of complete genomes highly diverse in length and composition. We find the following…
arxiv
Hong-Da Chen, Chang-Heng Chang, Li-Ching Hsieh, Hoong-Chien Lee
2004-12-18T03:15:03Z
置信度 0.78
q-bio.GN
-
The number of available genomes of prokaryotic organisms is rapidly growing enabling comparative genomics studies. The comparison of genomes of organisms with a common phenotype, habitat or phylogeny often shows that these genomes share some common contents. C…
arxiv
Giorgio Gonnella
2023-03-15T16:55:03Z
置信度 0.78
q-bio.GN
-
Research in quantitative evolutionary genomics and systems biology led to the discovery of several universal regularities connecting genomic and molecular phenomic variables. These universals include the log-normal distribution of the evolutionary rates of ort…
arxiv
Eugene V. Koonin
2011-08-17T22:19:54Z
置信度 0.78
q-bio.PEq-bio.GNq-bio.MN
-
As synthetic genomics scales toward the construction of increasingly larger genomes, computational strategies are needed to address technical feasibility. We introduce an algorithmic framework for the Minimum-Cost Synthetic Genome Planning problem, aiming to i…
arxiv
Michail Patsakis, Ioannis Mouratidis, Ilias Georgakopoulos-Soares
2025-09-07T22:46:57Z
置信度 0.78
q-bio.GN
-
Relation of genome sizes to organisms complexity is still described rather equivocally. Neither the number of genes (G-value), nor the total amount of DNA (C-value) correlates consistently with phenotype complexity. Using information theory considerations we d…
arxiv
Dmitri V. Parkhomchuk
2006-12-19T14:31:13Z
置信度 0.78
q-bio.GNq-bio.PE
-
The phenomenon of gene conservation is an interesting evolutionary problem related to speciation and adaptation. Conserved genes are acted upon in evolution in a way that preserves their function despite other structural and functional changes going on around …
arxiv
Bradly J. Alicea, Marcela A. Carvallo-Pinto, Jorge L. M. Rodrigues
2008-07-21T20:38:48Z
置信度 0.78
q-bio.GNq-bio.PE
-
Today's sequencing technology allows sequencing an individual genome within a few weeks for a fraction of the costs of the original Human Genome project. Genomics labs are faced with dozens of TB of data per week that have to be automatically processed and mad…
arxiv
Uwe Roehm, Jose Blakeley
2009-09-09T18:09:17Z
置信度 0.78
cs.DBq-bio.GN
-
Deep learning techniques have driven significant progress in various analytical tasks within 3D genomics in computational biology. However, a holistic understanding of 3D genomics knowledge remains underexplored. Here, we propose MIX-HIC, the first multimodal …
arxiv
Minghao Yang, Pengteng Li, Yan Liang, Qianyi Cai 等
2025-04-12T03:31:03Z
置信度 0.78
cs.LGcs.AIq-bio.GN
-
Genomic data visualization is essential for interpretation and hypothesis generation as well as a valuable aid in communicating discoveries. Visual tools bridge the gap between algorithmic approaches and the cognitive skills of investigators. Addressing this n…
arxiv
Sabrina Nusrat, Theresa Harbig, Nils Gehlenborg
2019-05-08T00:53:22Z
置信度 0.78
q-bio.GNcs.HCq-bio.QM
-
This work analyzed genome-wide nucleotide distribution patterns in ten insect genomes. Two internal measures were applied: (i) GC variation and (ii) third codon nucleotide preference. Although the genome size and overall GC level did not show any correlation w…
arxiv
Manoj Pratim Samanta
2007-02-18T08:17:22Z
置信度 0.78
q-bio.GN
-
Given the increasing volume and quality of genomics data, extracting new insights requires interpretable machine-learning models. This work presents Genomic Interpreter: a novel architecture for genomic assay prediction. This model outperforms the state-of-the…
arxiv
Zehui Li, Akashaditya Das, William A V Beardall, Yiren Zhao 等
2023-06-08T12:10:13Z
置信度 0.78
cs.LGq-bio.GN
-
During the course of evolution, an organism's genome can undergo changes that affect the large-scale structure of the genome. These changes include gene gain, loss, duplication, chromosome fusion, fission, and rearrangement. When gene gain and loss occurs in a…
arxiv
Birte Kehr, Knut Reinert, Aaron E. Darling
2012-07-30T15:24:49Z
置信度 0.78
q-bio.GN
-
We now need more than ever to make genome analysis more intelligent. We need to read, analyze, and interpret our genomes not only quickly, but also accurately and efficiently enough to scale the analysis to population level. There currently exist major computa…
arxiv
Mohammed Alser, Joel Lindegger, Can Firtina, Nour Almadhoun 等
2022-05-16T19:43:13Z
置信度 0.78
q-bio.GNcs.ARq-bio.QM
-
The problem of the directionality of genome evolution is studied from the information-theoretic view. We propose that the function-coding information quantity of a genome always grows in the course of evolution through sequence duplication, expansion of code, …
arxiv
Liaofu Luo
2008-05-08T01:13:20Z
置信度 0.78
q-bio.GNq-bio.CB
-
Clinical predictions using clinical data by computational methods are common in bioinformatics. However, clinical predictions using information from genomics datasets as well is not a frequently observed phenomenon in research. Precision medicine research requ…
arxiv
Moeez M. Subhani, Ashiq Anjum
2020-06-14T12:23:49Z
置信度 0.78
q-bio.GNcs.AIcs.LGstat.ML
-
Because genomes are products of natural processes rather than intelligent design, all genomes contain functional and nonfunctional parts. The fraction of the genome that has no biological function is called rubbish DNA. Rubbish DNA consists of junk DNA, i.e., …
arxiv
Dan Graur
2016-01-22T15:41:37Z
置信度 0.78
q-bio.GN
-
Segmentation and genome annotation (SAGA) algorithms are widely used to understand genome activity and gene regulation. These algorithms take as input epigenomic datasets, such as chromatin immunoprecipitation-sequencing (ChIP-seq) measurements of histone modi…
arxiv
Maxwell W Libbrecht, Rachel CW Chan, Michael M Hoffman
2021-01-03T19:08:36Z
置信度 0.78
q-bio.GNcs.LG
-
Pan-genome analysis is a standard procedure to decipher genome heterogeneity and diversification of bacterial species. Specie evolution is traced by defining and comparing the core (conserved), accessory (dispensable) and unique (strain-specific) gene pool wit…
arxiv
Zarrin Basharat, Azra Yasmin
2016-10-13T16:32:47Z
置信度 0.78
q-bio.GN
-
We propose that the distribution of DNA words in genomic sequences can be primarily characterized by a double Pareto-lognormal distribution, which explains lognormal and power-law features found across all known genomes. Such a distribution may be the result o…
arxiv
Miklós Csűrös, Laurent Noé, Gregory Kucherov
2006-09-14T17:18:30Z
置信度 0.78
q-bio.GN
-
The wide array of currently available genomes display a wonderful diversity in size, composition and structure with many more to come thanks to several global biodiversity genomics initiatives starting in recent years. However, sequencing of genomes, even with…
arxiv
Katharine M. Jenike, Lucía Campos-Domínguez, Marilou Boddé, José Cerca 等
2024-04-01T22:58:30Z
置信度 0.78
q-bio.GN
-
The advancements in artificial intelligence in recent years, such as Large Language Models (LLMs), have fueled expectations for breakthroughs in genomic foundation models (GFMs). The code of nature, hidden in diverse genomes since the very beginning of life's …
arxiv
Heng Yang, Jack Cole, Ke Li
2024-10-02T17:40:44Z
置信度 0.78
q-bio.GNcs.CL
-
Analysis of genomics data is central to nearly all areas of modern biology. Despite significant progress in artificial intelligence (AI) and computational methods, these technologies require significant human oversight to generate novel and reliable biological…
arxiv
Sehi L'Yi, John Conroy, Priya Misner, David Kouřil 等
2025-10-28T18:48:02Z
置信度 0.78
q-bio.GN
-
The article presents the theoretical foundations of the algorithm for calculating the number of different genomes in the medium under study and of two algorithms for determining the presence of a particular (known) genome in this medium. The approach is based …
arxiv
Valery Kirzhner, Zeev Volkovich
2015-06-19T21:12:34Z
置信度 0.78
q-bio.GNq-bio.QM
-
We outline the global control architecture of genomes. A theory of genomic control information is presented. The concept of a developmental control network called a cene (for control gene) is introduced. We distinguish parts-genes from control genes or cenes. …
arxiv
Eric Werner
2011-10-24T15:49:30Z
置信度 0.78
q-bio.OTcs.CEq-bio.GN
-
Microbial genome web portals have a broad range of capabilities that address a number of information-finding and analysis needs for scientists. This article compares the capabilities of the major microbial genome web portals to aid researchers in determining w…
arxiv
Peter D. Karp, Natalia Ivanova, Markus Krummenacker, Nikos Kyrpides 等
2018-10-30T17:01:34Z
置信度 0.78
q-bio.GN
-
This thesis is focused in the study of the evolution of the metabolic repertoire of Streptomyces, which are renowned as proficient producers of bioactive Natural Products (NPs). The main goal of my work was to contribute into the understanding of the evolution…
arxiv
Pablo Cruz-Morales
2022-05-30T09:25:35Z
置信度 0.78
q-bio.GN
-
Summary: With the rapid development of long-read sequencing technologies, the era of individual complete genomes is approaching. We have developed wgatools, a cross-platform, ultrafast toolkit that supports a range of whole genome alignment (WGA) formats, offe…
arxiv
Wenjie Wei, Songtao Gui, Jian Yang, Erik Garrison 等
2024-09-13T06:39:17Z
置信度 0.78
q-bio.GN
-
Genomic data and biomedical imaging data are undergoing exponential growth. However, our understanding of the phenotype-genotype connection linking the two types of data is lagging behind. While there are many types of software that enable the manipulation and…
arxiv
Andrew T. Oberlin, Dominika A. Jurkovic, Mitchell F. Balish, Iddo Friedberg
2012-12-03T16:52:41Z
置信度 0.78
q-bio.GN
-
Annotations of gene structures and regulatory elements can inform genome-wide association studies (GWAS). However, choosing the relevant annotations for interpreting an association study of a given trait remains challenging. We describe a statistical model tha…
arxiv
Joseph K. Pickrell
2013-11-19T19:12:08Z
置信度 0.78
q-bio.GN
-
Motivation: Protein-to-genome alignment is critical to annotating genes in non-model organisms. While there are a few tools for this purpose, all of them were developed over ten years ago and did not incorporate the latest advances in alignment algorithms. The…
arxiv
Heng Li
2022-10-14T18:43:47Z
置信度 0.78
q-bio.GN
-
Background: With the fast development of next generation sequencing technologies, increasing numbers of genomes are being de novo sequenced and assembled. However, most are in fragmental and incomplete draft status, and thus it is often difficult to know the a…
arxiv
Binghang Liu, Yujian Shi, Jianying Yuan, Xuesong Hu 等
2013-08-09T01:51:19Z
置信度 0.78
q-bio.GN
-
Thanks to the increasing availability of genomics and other biomedical data, many machine learning approaches have been proposed for a wide range of therapeutic discovery and development tasks. In this survey, we review the literature on machine learning appli…
arxiv
Kexin Huang, Cao Xiao, Lucas M. Glass, Cathy W. Critchlow 等
2021-05-03T21:20:20Z
置信度 0.78
cs.LGq-bio.GNq-bio.QM
-
The code of nature, embedded in DNA and RNA genomes since the origin of life, holds immense potential to impact both humans and ecosystems through genome modeling. Genomic Foundation Models (GFMs) have emerged as a transformative approach to decoding the genom…
arxiv
Heng Yang, Jack Cole, Yuan Li, Renzhi Chen 等
2025-05-20T14:16:25Z
置信度 0.78
q-bio.GNcs.CL
-
Although DNA foundation models have advanced the understanding of genomes, they still face significant challenges in the limited scale and diversity of genomic data. This limitation starkly contrasts with the success of natural language foundation models, whic…
arxiv
Huixin Zhan, Ying Nian Wu, Zijun Zhang
2024-02-12T21:40:45Z
置信度 0.78
q-bio.GNcs.AIcs.LG
-
How can we identify causal genetic mechanisms that govern bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype return high accuracy scores. However, attempts to extract any meaning from t…
arxiv
Tamsin James, Ben Williamson, Peter Tino, Nicole Wheeler
2025-02-11T18:25:14Z
置信度 0.78
q-bio.GNcs.LG
-
The classic algorithms of Needleman--Wunsch and Smith--Waterman find a maximum a posteriori probability alignment for a pair hidden Markov model (PHMM). In order to process large genomes that have undergone complex genome rearrangements, almost all existing wh…
arxiv
Colin Dewey, Peter Huggins, Kevin Woods, Bernd Sturmfels 等
2005-12-02T21:24:37Z
置信度 0.78
q-bio.GNmath.COq-bio.QM
-
Mitochondrial genomes in the Pinaceae family are notable for their large size and structural complexity. In this study, we sequenced and analyzed the mitochondrial genome of Cathaya argyrophylla, an endangered and endemic Pinaceae species, uncovering a genome …
arxiv
Kerui Huang, Wenbo Xu, Haoliang Hu, Xiaolong Jiang 等
2024-10-09T15:51:52Z
置信度 0.78
q-bio.GN
-
Next-generation sequencing (NGS) is a pivotal technique in genome sequencing due to its high throughput, rapid results, cost-effectiveness, and enhanced accuracy. Its significance extends across various domains, playing a crucial role in identifying genetic va…
arxiv
Fathima Nuzla Ismail, Shanika Amarasoma
2025-04-24T18:05:08Z
置信度 0.78
q-bio.GN
-
Deep neural networks excel in mapping genomic DNA sequences to associated readouts (e.g., protein-DNA binding). Beyond prediction, the goal of these networks is to reveal to scientists the underlying motifs (and their syntax) which drive genome regulation. Tra…
arxiv
Alex M. Tseng, Gokcen Eraslan, Tommaso Biancalani, Gabriele Scalia
2024-10-08T17:15:26Z
置信度 0.78
q-bio.GNcs.LG
-
Accurate evaluation of genome assemblies within highly repetitive regions, such as centromeres, remains a major open challenge in genomics. Conventional benchmarking relies on sequence alignment, which becomes problematic in regions of high homogeneity and div…
arxiv
Luca Franco, Matteo Migliarini, Matteo Tommaso Ungaro, Egnald Çela 等
2026-06-09T10:42:49Z
置信度 0.78
q-bio.GN