{"id":75848,"date":"2026-06-16T17:12:12","date_gmt":"2026-06-16T17:12:12","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/75848\/"},"modified":"2026-06-16T17:12:12","modified_gmt":"2026-06-16T17:12:12","slug":"the-ai-chemist-to-be-trustworthy-llms-need-to-show-their-work","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/75848\/","title":{"rendered":"The AI Chemist: To be trustworthy, LLMs need to show their work"},"content":{"rendered":"<p class=\"article-content\">\u00a0<\/p>\n<p>        Introducing the AI Chemist<\/p>\n<p>Artificial intelligence and large language models offer promising methods to interpret vast amounts of data but also more than a few cautions. This C&amp;EN column will cover what the technologies can do now, what they could do in the future, and what they shouldn\u2019t tackle\u2014all written by expert contributors.<\/p>\n<p class=\"article-content\">Drug discovery is really, really hard, and most drug candidates fail: humans are variable, the animals we use for testing aren\u2019t humans, pharmacokinetics and pharmacodynamics are hard to predict, and unsuspected off-target effects cause toxicity. The emergence of artificial intelligence and machine learning tools such as AlphaFold has raised excitement around the potential to accelerate early-stage drug discovery. Even AI skeptics, who professionally criticize the utility and ethics of ChatGPT and other large language models (LLMs), will often say, \u201cBut of course, AlphaFold is helping cure cancer, so it\u2019s not all bad.\u201d<\/p>\n<p class=\"article-content\">I\u2019m not sure I agree that software like AlphaFold is an exception.<\/p>\n<p class=\"article-content\">Computer-aided drug design (CADD) uses models of proteins and chemical compounds to prioritize compounds for investigation as possible drugs. This approach accelerates the early stage of drug development, as it helps focus attention on the likeliest candidates while considering frankly enormous numbers of candidates\u2014as many as trillions of compounds. To be fair, medicinal chemists have always considered protein structure when designing drugs, but the tools and the availability of protein structures were more limited in the past. Over the past 50 years, slowly (and then ever more quickly), the computational tools addressing docking, molecular dynamics, free-energy perturbation calculations, and single-point quantum mechanics calculations have improved. The computing power available to run these calculations has improved, and so has the availability of experimentally confirmed protein conformations. But both the use of these tools and the acquisition of the data they rely on require a lot of expertise.<\/p>\n<p class=\"article-content\">When AlphaFold arrived, suddenly all structures of all proteins were available to everyone at the click of a button. Then LLM-guided docking arrived, which simplified protein-ligand CADD docking, greatly democratizing the screening of protein-ligand interactions. While I was writing this column, a high school student approached me and my colleagues, asking us to help them test some proposed drug candidates they had identified though \u201cvibe CADDing,\u201d performing CADD by conversing with an AI assistant. You can do all the steps of CADD without the collaboration of a bunch of PhDs from different disciplines and without having studied any organic or physical, let alone quantum or medicinal, chemistry at all.<\/p>\n<p class=\"article-content\">The problem is that in practice, conducting these studies still does require extensive collaboration. A structure from the Protein Data Bank (PDB) is a single frozen conformation of a highly dynamic protein. You can\u2019t use that structure directly in a CADD study without considering dynamics, the protein\u2019s inherent natural environment, the influence of the ligands that the protein was soaked with as a way to reduce motion enough to make a crystal or get a good cryo-electron microscopy ensemble, and more.<\/p>\n<p class=\"article-content\">Structural biologists know these considerations: many proteins have dozens (or even hundreds) of different entries in the PDB. Those entries exist because the structures are not all the same. Computational, all-atomic simulations can be used to model some of this dynamism back in, but the assumptions used in these calculations are abstractions of physical reality. And when we assume . . . that leads to weird errors for you and me. But if you do this process regularly, you know what to look for and how to validate your models.<\/p>\n<p class=\"article-content\">If a protein is not present in the PDB, you can use homology modeling to try to recreate a reasonable 3D structural ensemble. Because that model is built on shaky ground, validating it against your own integrated experimental and computational data can take highly skilled practitioners months or years.<\/p>\n<p class=\"article-content\">In contrast, AlphaFold gives me a structure in seconds! Using it is so much more efficient. But my team often sees AlphaFold make the very human mistake of imposing structural order where there is only chaos. Proteins with long unstructured domains that flop around a lot can\u2019t be captured using our current techniques. So these parts of (or whole) proteins aren\u2019t in the PDB datasets used to train AlphaFold. A model of roads trained on only grid cities won\u2019t properly predict the street plans of Rome or Boston. Having the available data biased toward a well-organized structure is a problem. But it isn\u2019t the big one.<\/p>\n<p class=\"article-content\">The main problem is that I don\u2019t know how the program got the structure. I can\u2019t track the process back to check the assumptions, the steps, or the models that went into the structure generation. I can\u2019t check if it is right. Getting a good protein model (note, not a structure; almost all useful CADD models are a collection of individual conformations inclusive of a dynamic modeling component) is the most important step in CADD.<\/p>\n<p class=\"article-content\">Good science requires a \u201cchain of custody of why\u201d connecting input data and observation to an output conclusion. Every step in the logic chain must be auditable by our peers and those who want to use our work. We must be able to challenge and potentially falsify every decision point so that when something inevitably goes wrong, we can try to figure out where we went wrong. Protein-structure models don\u2019t allow for this.<\/p>\n<p class=\"article-content\">Consequently, I don\u2019t use AlphaFold\u2014or any similar LLM\u2014to generate structures. I made that decision because the only way to validate what the programs give me is to use the \u201cold fashioned\u201d (that is, circa 2020 but with our better computing tools in 2026) protein-preparation approach to check. The result is interesting if the program\u2019s product is more or less accurate than mine. It is interesting if it is completely wrong, and the program is leading others <a href=\"https:\/\/inpreparation.substack.com\/p\/go-ahead-and-use-ai-to-design-drugs\" shape=\"rect\" rel=\"nofollow noopener\" target=\"_blank\">down the garden path of futility<\/a>.<\/p>\n<p class=\"article-content\">I\u2019m certainly not saying these models are always wrong. That suggestion is far from the case. The problem is they aren\u2019t always right. And if you are relying on them for CADD (rather than as pretty pictures), they need to be. The old way will be wrong sometimes. But you will know why you are wrong when the data start coming back, and you will be able to adjust because you have your chain of custody of why. With an LLM, there is only the one step. If you can\u2019t examine the \u201cwhy?\u201d at every step of the chain, you can\u2019t determine if your hypothesis is realistic and worth spending large amounts of money to check out, and you can\u2019t tell if you would be better served by hosting a giant bonfire party using your investors\u2019 cash as fuel. And that situation makes me nervous. Nervous enough not to use an LLM.<\/p>\n<p>              <img data-lazy-src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/AI-Column-The-cost-of-AI-in-drug-discovery---489702.webp\"  alt=\"Portrait of John Trant\" class=\"w-100\" decoding=\"async\"\/><br \/>\n              Portrait of John Trant<\/p>\n<p>            Credit:<br \/>\n              Courtesy of John Trant<\/p>\n<p class=\"article-content\">John Trant is an associate professor and a Faculty of Science research chair at the University of Windsor, where he leads a team of interdisciplinary scientists tackling societally relevant molecular problems.<\/p>\n<p class=\"article-content\">Views expressed are those of the author and not necessarily those of C&amp;EN or the American Chemical Society.<\/p>\n<p class=\"article-content\">Do you have a view about using AI in chemistry that you\u2019d like to share with our readers? Please email editor Chris Gorski at <a href=\"https:\/\/cen.acs.org\/pharmaceuticals\/drug-discovery\/AI-Chemist-trustworthy-LLMs-need\/104\/web\/2026\/mailto:c_gorski@acs.org\" shape=\"rect\" rel=\"nofollow noopener\" target=\"_blank\">c_gorski@acs.org<\/a> to propose a column that you\u2019d like to write.<\/p>\n","protected":false},"excerpt":{"rendered":"\u00a0 Introducing the AI Chemist Artificial intelligence and large language models offer promising methods to interpret vast amounts&hellip;\n","protected":false},"author":2,"featured_media":75849,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,11553,601,25,2486,3845,13004,203,13590,11794,1941,1407,1944,52,41616],"class_list":["post-75848","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-alphafold","tag-artificial","tag-artificial-intelligence","tag-big","tag-big-data","tag-cost","tag-data","tag-discovery","tag-drug","tag-drug-discovery","tag-investigation","tag-protein","tag-research","tag-structure"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/75848","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=75848"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/75848\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/75849"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=75848"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=75848"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=75848"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}