{"id":27989,"date":"2025-10-14T15:52:47","date_gmt":"2025-10-14T15:52:47","guid":{"rendered":"https:\/\/naijaglobalnews.org\/?p=27989"},"modified":"2025-10-14T15:52:47","modified_gmt":"2025-10-14T15:52:47","slug":"new-dna-search-engine-brings-order-to-biologys-big-data","status":"publish","type":"post","link":"https:\/\/naijaglobalnews.org\/?p=27989","title":{"rendered":"New DNA Search Engine Brings Order to Biology\u2019s Big Data"},"content":{"rendered":"<p>\n<\/p>\n<p class=\"article_pub_date-zPFpJ\">October 14, 2025<\/p>\n<p class=\"article_read_time-ZYXEi\">3 min read<\/p>\n<p>New DNA Search Engine Brings Order to Biology\u2019s Big Data<\/p>\n<p>MetaGraph compresses vast data archives into a search engine for scientists, opening up new frontiers of biological discovery<\/p>\n<p class=\"article_authors-ZdsD4\">By Elie Dolgin &amp; Nature magazine <\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">The Internet has Google. Now biology has MetaGraph. Detailed today in Nature, the search engine can quickly sift through the staggering volumes of biological data housed in public repositories.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cIt\u2019s a huge achievement,\u201d says Rayan Chikhi, a biocomputing researcher at the Pasteur Institute in Paris. \u201cThey set a new standard\u201d for analysing raw biological data \u2014 including DNA, RNA and protein sequences \u2014 from databases that can contain millions of billions of DNA letters, amounting to \u2018petabases\u2019 of information, more entries than all the webpages in Google\u2019s vast index.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">Although MetaGraph is tagged as \u2018Google for DNA\u2019, Chikhi likens the tool to a search engine for YouTube, because the tasks are more computationally demanding. In the same way that YouTube searches can retrieve every video that features, say, red balloons even when those key words don\u2019t appear in the title, tags or description, MetaGraph can uncover genetic patterns hidden deep within expansive sequencing data sets without needing those patterns to be explicitly annotated in advance.<\/p>\n<h2>On supporting science journalism<\/h2>\n<p>If you&#8217;re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cIt enables things that cannot be done in any other way,\u201d Chikhi says.<\/p>\n<h2 id=\"indexing-lifes-library\" class=\"\" data-block=\"sciam\/heading\">Indexing life\u2019s library<\/h2>\n<p class=\"\" data-block=\"sciam\/paragraph\">The motivation behind MetaGraph was to address an accessibility problem in sequencing data sets. The size of these repositories has risen at a blistering pace in the past few decades, but this growth has presented challenges for the scientists using the data they contain. Raw sequencing reads are fragmented, noisy and too numerous to search directly. \u201cThe volume of the data, paradoxically, is the main inhibitor of us actually using the data,\u201d says Artem Babaian, a computational biologist at the University of Toronto in Canada.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">According to one of the study authors, Andr\u00e9 Kahles, a bioinformatician at the Swiss Federal Institute of Technology (ETH) Zurich in Switzerland, MetaGraph could help researchers to ask biological questions of repositories such as the Sequence Read Archive (SRA), a public database containing in excess of 100 million billion DNA letters.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">They tackled the problem through the use of mathematical \u2018graphs\u2019 that links overlapping DNA fragments together, much like sentences that share the same words lining up in a book index.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">The researchers integrated data from seven publicly funded data repositories, creating 18.8 million unique DNA and RNA sequence sets and 210 billion amino-acid sequence sets across all clades of life \u2014 including viruses, bacteria, fungi, plants and animals, including humans. They also developed a search engine for these sequences, in which users use text prompts to search these integrated archives of raw data.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cIt is a totally new way to interact with this body of data,\u201d says Kahles. \u201cIt\u2019s compressed, but accessible on the fly.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">To demonstrate the utility of MetaGraph, the study authors used it to scan 241,384 human gut microbiome samples for genetic indicators of antibiotic resistance around the world, building on work that used an earlier version of the tool to track drug-resistance genes in bacterial strains that live in subway systems across major urban centres. The authors say they performed the analysis in about an hour on a high-powered computer.<\/p>\n<h2 id=\"open-road-to-discovery\" class=\"\" data-block=\"sciam\/heading\">Open road to discovery<\/h2>\n<p class=\"\" data-block=\"sciam\/paragraph\">MetaGraph is not the only massive-scale sequence search tool now on offer.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">Chikhi and Babaian, for example, have built a platform called Logan, which stitches together billions of short sequencing reads to make longer, organized stretches of DNA. This design architecture allows the system to spot whole genes and their variants across even larger collections of sequencing reads than is possible with MetaGraph, albeit with certain trade-offs. \u201cWe have less functionality but more performance,\u201d Chikhi says.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">The added reach of Logan helped the researchers to uncover more than 200 million naturally occurring versions of a plastic-eating enzyme found in a variety of bacteria, fungi and insects \u2014 including some versions that work even better than enzymes designed in the lab. Chikhi and Babaian reported their findings in a preprint posted last month.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">They and others have also used an earlier, narrower search tool tailored to viral-DNA repositories to reveal reams of previously undocumented viruses and viral contaminants in engineered T-cell therapies for treating cancer.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">According to Babaian, such discoveries would not have been possible without two things: open-source search tools, available at sites such as metagraph.ethz.ch and logan-search.org, and the public sequencing repositories they tap into. With funding cuts threating other sorts of biological databases, Babaian stresses that these search innovations underscore the \u201ccritical importance of open data sharing\u201d.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cThese are resources to drive scientific progress across the world,\u201d says Babaian. \u201cThey are opening up a completely new field of petabase-scale genomics\u201d \u2014 and the most impactful applications are yet to come.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">This article is reproduced with permission and was first published on October 8, 2025.<\/p>\n<h2 class=\"subscriptionPleaHeading-DMY4w\">It\u2019s Time to Stand Up for Science<\/h2>\n<p class=\"subscriptionPleaText--StZo\">If you enjoyed this article, I\u2019d like to ask for your support. <span class=\"subscriptionPleaItalicFont-i0VVV\">Scientific American<\/span> has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.<\/p>\n<p class=\"subscriptionPleaText--StZo\">I\u2019ve been a <span class=\"subscriptionPleaItalicFont-i0VVV\">Scientific American<\/span> subscriber since I was 12 years old, and it helped shape the way I look at the world. <span class=\"subscriptionPleaItalicFont-i0VVV\">SciAm <\/span>always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.<\/p>\n<p class=\"subscriptionPleaText--StZo\">If you subscribe to <span class=\"subscriptionPleaItalicFont-i0VVV\">Scientific American<\/span>, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.<\/p>\n<p class=\"subscriptionPleaText--StZo\">In return, you get essential news, captivating podcasts, brilliant infographics, can&#8217;t-miss newsletters, must-watch videos, challenging games, and the science world&#8217;s best writing and reporting. You can even gift someone a subscription.<\/p>\n<p class=\"subscriptionPleaText--StZo\">There has never been a more important time for us to stand up and show why science matters. I hope you\u2019ll support us in that mission.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>October 14, 2025 3 min read New DNA Search Engine Brings Order to Biology\u2019s Big Data MetaGraph compresses vast data archives into a search engine for scientists, opening up new frontiers of biological discovery By Elie Dolgin &amp; Nature magazine The Internet has Google. Now biology has MetaGraph. Detailed today in Nature, the search engine<\/p>\n","protected":false},"author":1,"featured_media":27990,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[50],"tags":[1285,16665,1620,1111,734,7553,1435,2018],"class_list":{"0":"post-27989","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-environment","8":"tag-big","9":"tag-biologys","10":"tag-brings","11":"tag-data","12":"tag-dna","13":"tag-engine","14":"tag-order","15":"tag-search"},"_links":{"self":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts\/27989","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=27989"}],"version-history":[{"count":0,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts\/27989\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/media\/27990"}],"wp:attachment":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=27989"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=27989"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=27989"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}