{"id":7257,"date":"2026-07-27T12:00:03","date_gmt":"2026-07-27T12:00:03","guid":{"rendered":"https:\/\/hoo.central12.com\/portal\/2026\/07\/27\/closing-the-data-loop-in-ai-driven-drug-discovery\/"},"modified":"2026-07-27T12:00:03","modified_gmt":"2026-07-27T12:00:03","slug":"closing-the-data-loop-in-ai-driven-drug-discovery","status":"publish","type":"post","link":"https:\/\/hoo.central12.com\/portal\/2026\/07\/27\/closing-the-data-loop-in-ai-driven-drug-discovery\/","title":{"rendered":"Closing the data loop in AI-driven drug discovery"},"content":{"rendered":"<p>Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage.<\/p>\n<p>Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years\u2014a phenomenon known as <a href=\"https:\/\/www.oecd.org\/en\/publications\/artificial-intelligence-in-science_a8d820bd-en\/full-report\/eroom-s-law-and-the-decline-in-the-productivity-of-biopharmaceutical-r-d_f42df75c.html\" target=\"_blank\" rel=\"noreferrer noopener\">Eroom\u2019s Law<\/a>. Today, bringing a new drug to market takes an average of 10-15 years and costs anywhere from <a href=\"https:\/\/cdn.cytivalifesciences.com\/api\/public\/content\/srGfSmOLS1yA_mNQchWbKA-pdf\" target=\"_blank\" rel=\"noreferrer noopener\">$1 billion to $2.5 billion<\/a>, with failure rates upward of 90%.<\/p>\n<p>AI has become the pharmaceutical industry\u2019s biggest bet on bringing success rates up and timelines down. The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development.<\/p>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"1254\" height=\"836\" src=\"https:\/\/wp.technologyreview.com\/wp-content\/uploads\/2026\/06\/iStock-2247435323_f6bf40.jpg\" alt=\"\" class=\"wp-image-1139670\" style=\"width:741px;height:auto\" \/><\/figure>\n<p>\u201cThe main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,\u201d says Paul Belcher, director of protein research strategy at global life sciences company Cytiva. \u201cAI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic.\u201d<\/p>\n<p>Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems.<\/p>\n<h3 class=\"wp-block-heading\"><a><\/a>AI brings efficiency to the lab<\/h3>\n<p>One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it. A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug.<\/p>\n<p>Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&amp;D).<\/p>\n<p>This means companies are no longer limited by how much they can physically screen to identify starting points. \u201cAI does away with that,\u201d says Belcher. \u201cAnd it can help eliminate low-quality candidates before you have to physically test them, saving time and resources.\u201d<\/p>\n<p>What AI can\u2019t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab.<\/p>\n<p>Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.<\/p>\n<p>\u201cThe current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data\u2014yes-or-no responses,\u201d Belcher explains. \u201cAI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits.\u201d<\/p>\n<h3 class=\"wp-block-heading\"><a><\/a>Models need complete, quality data<\/h3>\n<p>As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data.<\/p>\n<p>Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren&#8217;t built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias.<\/p>\n<p>Publication bias reinforces the problem. \u201cMost publicly available datasets and scientific publications focus exclusively on positive results,\u201d says Belcher. \u201cNo one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.\u201d<\/p>\n<p>The data Belcher believes would markedly improve models\u2014the failed experiments, the compounds that don\u2019t bind\u2014remains frustratingly difficult to come by. \u201cWe often joke that there should be a journal of negative data,\u201d he says. \u201cIt\u2019s often buried in lab notebooks, and it\u2019s never used to inform or guide future research.\u201d<\/p>\n<p>This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can\u2019t be adequately trained to avoid bias. \u201cIn all machine learning applications, the model\u2019s performance relies heavily on the quality and scope of the training data,\u201d notes Belcher.<\/p>\n<p>Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example. These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher <a href=\"https:\/\/www.nature.com\/articles\/d41586-020-01363-z\" target=\"_blank\" rel=\"noreferrer noopener\">cites research<\/a> by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial.<\/p>\n<p>\u201cManipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,\u201d says Belcher. \u201cThere needs to be more tools to verify that data is not manipulated.\u201d<\/p>\n<p>Some vendors are starting to tackle this challenge. Belcher points to solutions like Cytiva\u2019s Image Integrity Checker, for instance, which uses secure hash algorithms\u2014the same technology used in blockchain\u2014to detect whether scientific images have been tampered with. \u201cWe\u2019re starting to see a lot of interest from publishing houses that want to adopt this as standard because it\u2019s a quick way to ensure that what gets published in the literature is genuine,\u201d he adds.<\/p>\n<h3 class=\"wp-block-heading\"><a><\/a>Autonomous labs could accelerate breakthroughs<\/h3>\n<p>Belcher describes the future state of drug discovery as fully autonomous labs that run with minimal human intervention. Foundational to this vision is consistency in data and infrastructure.<\/p>\n<p>These AI-driven dark labs, or labs-in-the-loop, operate around the clock. They cycle through prediction, testing, and optimization, and then feed results back into AI models to guide the next round of experiments. This can improve the success rates of drug candidates entering clinical trials, says Belcher. Better starting points, combined with more rounds of optimization, should result in better candidates with fewer liabilities reaching the clinic.<\/p>\n<p>But automating a lab depends heavily on integration. That means interoperable systems, highly structured and comprehensive datasets, and information flowing easily in and out. Most labs aren\u2019t there yet. \u201cToday, a lot of the instruments in labs are standalone,\u201d Belcher notes. \u201cYou can have the best technology in the world, but if it\u2019s a closed ecosystem\u2014if the user can\u2019t get the data out\u2014it doesn\u2019t do any good.\u201d<\/p>\n<p>An integrated infrastructure can enable labs to generate FAIR (findable, accessible, interoperable, and reusable) data at scale. This would not only inform individual lab reports, but could also train subsequent generations of AI models, effectively closing the loop between the computational, AI-driven dry lab and the physical wet lab.<\/p>\n<p>\u201cOur goal is to help scientists and researchers accelerate their breakthroughs and make that future state of autonomous labs a real possibility,\u201d says Belcher. \u201cWe want to help them generate reliable data, simplify workflows in discovery, and hopefully enable what they\u2019re working on to become tomorrow\u2019s life-changing therapies, faster and with greater confidence.\u201d<\/p>\n<h3 class=\"wp-block-heading\"><a><\/a>On costs and what comes next<strong><\/strong><\/h3>\n<p>AI-driven drug discovery is still in its early days. Notably, no drug discovered primarily through AI-driven design has yet received full FDA approval\u2014although Belcher expects that to change in the next two to three years.<\/p>\n<p>How big of an impact could AI eventually have on drug discovery? \u201cThe holy grail would be full in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work,\u201d says Belcher. But there are many barriers to this beyond the maturity of the models, including regulatory hurdles and cost challenges.<\/p>\n<p>A <a href=\"https:\/\/epoch.ai\/blog\/how-much-does-it-cost-to-train-frontier-ai-models\" target=\"_blank\" rel=\"noreferrer noopener\">Stanford study<\/a> found that the cost of training frontier AI models has more than doubled every year since 2016, adding more financial pressure to a sector already defined by <a href=\"https:\/\/www.sciencedirect.com\/science\/article\/pii\/S135964462400285X\" target=\"_blank\" rel=\"noreferrer noopener\">exceptionally high R&amp;D spend<\/a>.<\/p>\n<p>Belcher acknowledges the tension, but remains optimistic about what\u2019s ahead. \u201cI think we\u2019ll get to a point where there\u2019s a balance between AI and wet work, from a cost perspective and a risk perspective,\u201d he says. \u201cAs long as the cost of compute doesn\u2019t ever outweigh the cost of clinical development, I think AI is going to be an advantage.\u201d<\/p>\n<\/p>\n<p>Learn more about how <a href=\"https:\/\/www.cytivalifesciences.com\/solutions\/protein-research\" target=\"_blank\" rel=\"noreferrer noopener\">Cytiva is using faster discovery<\/a> to reshape protein purification workflows.<\/p>\n<p><em>This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review\u2019s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years\u2014a phenomenon known as Eroom\u2019s Law. Today, bringing a new drug to market takes an average [&hellip;]<\/p>\n","protected":false},"author":0,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[68],"tags":[67],"class_list":["post-7257","post","type-post","status-publish","format-standard","hentry","category-mit-feed","tag-mit-tech"],"_links":{"self":[{"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/posts\/7257","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/comments?post=7257"}],"version-history":[{"count":0,"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/posts\/7257\/revisions"}],"wp:attachment":[{"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/media?parent=7257"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/categories?post=7257"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/hoo.central12.com\/portal\/wp-json\/wp\/v2\/tags?post=7257"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}