{"id":78500,"date":"2026-06-18T15:43:06","date_gmt":"2026-06-18T15:43:06","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/78500\/"},"modified":"2026-06-18T15:43:06","modified_gmt":"2026-06-18T15:43:06","slug":"openai-life-science-benchmark-reveals-ai-passes-only-1-in-3-scientific-research-tasks","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/78500\/","title":{"rendered":"OpenAI Life Science Benchmark Reveals AI Passes Only 1 in 3 Scientific Research Tasks"},"content":{"rendered":"<p>OpenAI published <a href=\"https:\/\/openai.com\/index\/introducing-life-sci-bench\/\" target=\"_blank\" rel=\"noopener nofollow\">LifeSciBench<\/a> on June 17, 2026, a 750-task evaluation built with 173 PhD-level scientists to test whether AI can handle the full complexity of real life-science research \u2014 and the results put a hard number on how far the technology still has to go. The best-performing model, GPT-Rosalind, cleared 36.1% of tasks. Every other model tested did worse. Nearly two in three research-level tasks still defeat even the most capable AI system available.<\/p>\n<p>Why Standard Biology Benchmarks Have Always Been the Wrong Test<\/p>\n<p>Most AI evaluations for biology use the same format as a standardized exam: multiple-choice questions, clean reference answers, one correct option. That format has almost nothing in common with what a working scientist does. Researchers weigh incomplete data, reconcile conflicting published results, design experiments under uncertainty, and communicate findings with the right caveats for a regulatory or peer-review audience. LifeSciBench was built around that gap.<\/p>\n<p>Released alongside a <a href=\"https:\/\/cdn.openai.com\/pdf\/b4299379-0a97-4ffa-8b9b-c3fbb299caa9\/lifescibench_preprint.pdf\" target=\"_blank\" rel=\"noopener nofollow\">technical preprint<\/a>, the benchmark spans seven scientific workflows \u2014 evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, translation, and scientific communication \u2014 and seven biological domains running from genomics and medicinal chemistry through to clinical and translational science. Each task is structured as a free-response prompt the way a scientist might brief a knowledgeable colleague, paired with any relevant supporting materials and graded against a detailed expert-written rubric. No multiple-choice options. No reference strings to match against. Just a problem that requires the model to reason through and respond as a scientist would.<\/p>\n<p>How LifeSciBench Actually Scores AI<\/p>\n<p>The scoring architecture is what makes LifeSciBench technically distinctive from any prior life-science evaluation. Each task comes with a rubric averaging 25 specific grading criteria, covering not just whether a model reaches the correct conclusion, but whether it does so in a scientifically valid and operationally useful way \u2014 with appropriate justification, correct caveats, and the level of detail a domain expert would expect.<\/p>\n<p>Across all 750 tasks, those rubrics total 19,020 individual criteria. A model earns points for each criterion it satisfies. The task pass threshold is set at 70% of available rubric points \u2014 a design choice that means a model can accumulate partial credit without passing the task, and the benchmark records both the normalized score and the task pass rate separately. That separation matters: a response can contain high-quality reasoning while still failing to meet the full standard.<\/p>\n<p>More than half of tasks (53%) require models to interpret or synthesize information from at least one attached artifact \u2014 figures, PDFs, tables, genomic sequence files, chemical structure files, or web references. The full benchmark includes 1,062 such artifacts across 750 tasks. This is not incidental. Scientific work is multimodal by nature, and the ability to reason over a graph of experimental data or a chemical structure file is qualitatively different from answering a text-based question about the same subject.<\/p>\n<p>The 79% of tasks that require multiple reasoning or decision-making steps, averaging four steps per task, reflect the same principle: real scientific judgment is sequential and conditional, not single-step retrieval.<\/p>\n<p>Where Every Frontier Model Stands on the AI Life Science Benchmark<\/p>\n<p>Five models were evaluated in a single-turn setting, each seeing a task prompt and any attached artifacts once, with unrestricted internet browsing permitted.<\/p>\n<p><a href=\"https:\/\/openai.com\/index\/introducing-gpt-rosalind\/\" target=\"_blank\" rel=\"noopener nofollow\">GPT-Rosalind<\/a>, OpenAI&#8217;s domain-specialized life-sciences model, led with a 0.576 normalized score and a 36.1% task pass rate. GPT-5.5, the company&#8217;s general-purpose frontier model, followed at 0.519 and 25.7%. Gemini 3.1 Pro placed third at 0.515 and 23.6%, with GPT-5.4 at 0.479 and 20.7% and Grok 4.3 at 0.399 and 13.0%.<\/p>\n<p>The ranking tells only part of the story. GPT-Rosalind led on a per-task basis for 386 of 750 tasks \u2014 but Gemini 3.1 Pro uniquely led on 214 of them. Aggregate scores can hide task-specific strengths: models that perform worst overall may still outperform the leader on specific workflow categories. Notably, Anthropic&#8217;s Claude models were not included in the evaluation.<\/p>\n<p>Across all five models, frontier AI is currently most capable at structured judgment tasks \u2014 scientific communication, translation, and evidence synthesis where the output format and expectations are relatively stable. It struggles most in the categories that require designing something novel or making a multi-step optimization decision under constraint: Design, Optimization, and Prediction saw GPT-Rosalind passing only 30.7% of tasks; Analysis came in at 30.3%. The benchmark&#8217;s authors note that no model passes 171 tasks (22.8%), and 261 tasks (34.8%) have a best-model pass rate below 20% \u2014 meaning a substantial portion of the benchmark poses a challenge that no current AI can reliably meet.<\/p>\n<p>The Artifact Problem: Where AI Drops Sharply<\/p>\n<p>The clearest finding in the benchmark data is a consistent, steep performance penalty when tasks require artifact interpretation. GPT-Rosalind&#8217;s task pass rate falls from 45.1% on text-only tasks to 28.1% on tasks involving at least one attached artifact \u2014 a 17-percentage-point drop. The same degradation pattern holds for GPT-5.5: 29.9% on text-only tasks, 21.9% on artifact tasks.<\/p>\n<p>This is not a marginal difference. It is the benchmark&#8217;s most concrete signal about where frontier AI needs to improve before it can serve as a reliable research collaborator. In actual drug discovery workflows, experimental data almost always lives in figures, tables, assay output files, and chemical structure databases \u2014 not in text alone. An AI system that performs competently on text-based scientific questions but degrades significantly when handed a genomic sequence file or a spatial transcriptomics dataset has a functional gap in the part of the research pipeline that matters most.<\/p>\n<p>Does OpenAI Grading Its Own Models Create a Conflict?<\/p>\n<p>LifeSciBench is a proprietary benchmark designed and administered by the same company whose model \u2014 GPT-Rosalind \u2014 leads its leaderboard. Readers interpreting the 36.1% claim as independent scientific validation should understand this structural relationship.<\/p>\n<p>The concern is well-documented in the broader AI evaluation field. A peer-reviewed analysis published in <a href=\"https:\/\/www.nature.com\/articles\/s41591-026-04431-5\" target=\"_blank\" rel=\"noopener nofollow\">Nature Medicine<\/a> on June 12, 2026, examining OpenAI&#8217;s HealthBench evaluation, concluded that industry-created benchmarks may systematically favor the systems developed by their creators, and called for independently constructed evaluation instruments. Early community reactions to LifeSciBench&#8217;s publication included similar objections to &#8220;opaque expert selection and rival-focused framing.&#8221;<\/p>\n<p>OpenAI designed LifeSciBench with some mitigations in place. Tasks were written by 173 scientists external to the company, in collaboration with Tacit Labs, a startup specializing in feedback loops for drug development. The 453-person reviewer cohort \u2014 97% of whom hold doctorates \u2014 achieved more than 96% agreement on relevance, reasoning, grounding, and usefulness. Each task averaged six automated review cycles and at least two rounds of expert review before acceptance. Expert consensus was required at 90% agreement per domain.<\/p>\n<p>Whether those structural guardrails fully offset the inherent limitation of a self-administered benchmark remains an open question the evaluation community will need to answer as independent researchers examine the preprint and the task set. OpenAI has stated it intends to connect benchmark performance to deployment studies in live research settings, which would provide a more externally verifiable signal.<\/p>\n<p>What a 36 Percent AI Drug Discovery Pass Rate Means in Practice<\/p>\n<p>OpenAI frames LifeSciBench as the beginning of a measurement program, not a destination. The goal is to connect benchmark performance to deployment studies in actual research workflows \u2014 measuring whether AI systems accelerate discovery will require tracking model use over longer time horizons and across multiple rounds of reasoning and experimental follow-up.<\/p>\n<p>That framing is appropriate caution. A PitchBook analysis from January 2026 found that more than $17 billion has been invested in AI drug discovery since 2019, but no AI-developed drug has yet entered large-scale clinical trials. The gap between AI performance on any benchmark and AI impact on a drug approval timeline is long, complex, and poorly understood.<\/p>\n<p>A 36.1% pass rate on a hard benchmark does not mean AI is ready to function as an autonomous research partner. Joy Jiao, <a href=\"https:\/\/openai.com\/index\/introducing-gpt-rosalind\/\" target=\"_blank\" rel=\"noopener nofollow\">OpenAI&#8217;s life sciences research lead<\/a>, stated at GPT-Rosalind&#8217;s April 2026 launch that the company does not believe AI can yet create new disease treatments on its own.<\/p>\n<p>What the benchmark result does establish is a calibrated floor: the specific workflow categories, artifact types, and reasoning steps where frontier AI currently fails at rates that would make unsupervised deployment in a research pipeline risky. That calibration is useful information for research directors, enterprise AI buyers, and pharma organizations evaluating how to deploy AI tools \u2014 and where to keep human scientists in the loop.<\/p>\n<p>GPT-Rosalind is currently available in research preview to eligible organizations globally through OpenAI&#8217;s <a href=\"https:\/\/openai.com\/index\/introducing-gpt-rosalind\/\" target=\"_blank\" rel=\"noopener nofollow\">trusted-access program<\/a>, which requires legitimate scientific research with clear public benefit, enterprise-grade governance, and biosecurity oversight. OpenAI is also accepting applications from scientists who want to contribute to future benchmark iterations.<\/p>\n<p>Frequently Asked Questions<\/p>\n<p>What is LifeSciBench and how does it differ from existing AI benchmarks?<\/p>\n<p>LifeSciBench is a 750-task evaluation for AI in life-science research, published by OpenAI on June 17, 2026, and developed with 173 PhD-level scientists from biotechnology and pharmaceutical research. Unlike existing AI biology benchmarks that use multiple-choice questions with clean reference answers, LifeSciBench presents free-response tasks graded by expert-written rubrics averaging 25 criteria each \u2014 19,020 criteria in total \u2014 and requires models to interpret scientific artifacts including genomic sequence files, chemical structure files, and experimental figures. Tasks require an average of four reasoning steps; 53% require artifact synthesis.<\/p>\n<p>Why do frontier AI models struggle so much with life-science research tasks?<\/p>\n<p>The benchmark identified artifact processing as the primary bottleneck. GPT-Rosalind&#8217;s task pass rate drops from 45.1% on text-only tasks to 28.1% on tasks involving attached data files \u2014 a 17-percentage-point degradation. Scientific data almost always lives in figures, assay outputs, and structure files rather than in plain text, which means the part of real research workflows where AI is weakest is also the most common. Design, optimization, and multi-step analysis tasks are also particularly hard, with GPT-Rosalind clearing only 30.7% in design and 30.3% in analysis.<\/p>\n<p>Does OpenAI building the benchmark that tests its own model raise any concerns?<\/p>\n<p>Yes, this is a recognized limitation. A peer-reviewed study in Nature Medicine published June 12, 2026, found that industry-created benchmarks may systematically favor the systems developed by their creators. OpenAI included structural safeguards \u2014 external scientists authored the tasks, 453 independent reviewers provided quality control, and tasks required 90% reviewer agreement per domain \u2014 but independent replication of the results against externally constructed evaluations remains the standard the field needs before treating LifeSciBench scores as fully objective validation. OpenAI has committed to connecting benchmark scores to deployment studies in real research settings as a next step.<\/p>\n<p>Can AI help accelerate drug discovery today, and what should research organizations realistically expect?<\/p>\n<p>AI tools including GPT-Rosalind are most reliably useful for evidence synthesis, literature review, protocol assistance, and structured communication tasks, where LifeSciBench shows the strongest performance. A 36.1% pass rate on expert-designed research tasks means AI can meaningfully assist skilled scientists but cannot replace the judgment required to safely interpret its outputs. No AI-developed drug has yet reached large-scale clinical trials despite more than $17 billion invested in AI drug discovery since 2019. Research organizations that deploy AI tools at the evidence-handling and experimental-planning stages \u2014 with scientists reviewing and validating outputs \u2014 are using the technology in a manner consistent with current benchmarked capability.<\/p>\n","protected":false},"excerpt":{"rendered":"OpenAI published LifeSciBench on June 17, 2026, a 750-task evaluation built with 173 PhD-level scientists to test whether&hellip;\n","protected":false},"author":2,"featured_media":78501,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[42880,42881,10232,157],"class_list":["post-78500","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-ai-life-science","tag-benchmark","tag-life-science","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/78500","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=78500"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/78500\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/78501"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=78500"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=78500"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=78500"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}