{"id":133168,"date":"2026-08-07T18:51:18","date_gmt":"2026-08-07T18:51:18","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/133168\/"},"modified":"2026-08-07T18:51:18","modified_gmt":"2026-08-07T18:51:18","slug":"how-synthetic-data-can-help-ai-think-like-nobel-laureates","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/133168\/","title":{"rendered":"How Synthetic Data Can Help AI Think Like Nobel Laureates"},"content":{"rendered":"<p class=\"wp-block-paragraph\">As AI large language models, or LLMs, grow in power and sophistication, where will their architects turn when, someday \u2014 as experts predict \u2014 algorithms outgrow the limits of general human knowledge and begin craving information possessed by only the world\u2019s most elite thinkers and creators?\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Worrying as it might seem, this quest for the next frontier of data is already underway, spearheaded by research engineers in the technology industry, according to Harsh Raj, a graduate student in Northeastern University\u2019s Khoury College of Computer Sciences.<\/p>\n<p class=\"wp-block-paragraph\">Raj is one such intrepid explorer. As part of his two, high-profile co-ops, Raj has plumbed the depths of the most cutting-edge LLMs in search of data that could one day help them think like a winner of a Nobel Prize or Fields Medal \u2014 the highest prize for someone in mathematics \u2014 he said.<\/p>\n<p class=\"wp-block-paragraph\">LLMs are a type of artificial intelligence trained on massive amounts of data so that they can understand and create human-like writing. They can process text, summarize lengthy documents or write code in accordance with specific prompts fed to them by users.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt\u2019s very hard to collect data which trains the model to be better than humans, because there are very few humans who can create that data,\u201d Raj said. \u201cYou want to make it better than a Fields medalist or a Nobel Prize winner. How do you collect that?\u201d<\/p>\n<p class=\"wp-block-paragraph\">That question has become a focus for Raj. Now a research engineer wrapping up experiential learning at Scale AI, Raj has worked to understand the limitations and failings of LLMs in order to offer the most viable paths toward super-human knowledge capabilities. His specialized background in innovative AI projects has made him an attractive candidate for technology companies during his graduate studies.<\/p>\n<p>\tNortheastern Global News, in your inbox.<\/p>\n<p class=\"has-small-font-family has-small-font-size wp-block-paragraph\" style=\"margin-top:0\">Sign up for NGN\u2019s daily newsletter for news, discovery and analysis from around the world.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"990\" height=\"569\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/EmailGraphic.png\" alt=\"\" class=\"wp-image-217664 size-full\" style=\"object-position:50% 50%\"  \/><\/p>\n<p class=\"wp-block-paragraph\">Though Raj was courted by several technology companies hoping to leverage his skills, he chose to pursue his first co-op with Bespoke Labs, a Bay Area company that creates learning environments in which AI agents can be trained.<\/p>\n<p class=\"wp-block-paragraph\">While there, he worked alongside other researchers to understand the fail points of models hosted on cloud computing infrastructure and offer specific solutions, he said. He also created \u201csynthetic\u201d data which, unlike human-created writing or code, is generated artificially before being reused as training material for the LLMs, Raj said.<\/p>\n<p class=\"wp-block-paragraph\">Creating synthetic data that matches the quality of human-made data is a tall task, Raj said. Computing power can be acquired relatively easily, but finding the data to support the project is uncharted territory, he added.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" height=\"533\" width=\"800\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/IMG_4170.jpg\" alt=\"Group of ten colleagues posing together outdoors on a paved path, some wearing medals around their necks.\" class=\"wp-image-309507\"  \/>Harsh Raj has worked to understand the limitations and failings of LLMs in order to offer the most viable paths toward super-human knowledge capabilities. Courtesy Photo<\/p>\n<p class=\"wp-block-paragraph\">\u201cHarsh is a self-starter, very enterprising,\u201d said <a href=\"https:\/\/baulab.info\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">David Bau<\/a>, Northeastern professor of machine learning and a prominent figure in the field of artificial intelligence. \u201cHe went on to his new opportunities on his own, and the success he found in his co-op is all his doing.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Raj was inspired to enroll at Northeastern in large part because of Bau, he said. In addition to his work as a software engineer during the early days of Microsoft and Google, Bau in 2024 secured a $9 million grant from the U.S. National Science Foundation to launch <a href=\"https:\/\/www.khoury.northeastern.edu\/research_projects\/national-deep-inference-fabric-ndif\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">National Deep Inference Fabric<\/a>, a research computing infrastructure project at Northeastern to examine the mysteries of large-scale AI systems. Raj trusted that Bau\u2019s expertise would offer him the best opportunity to expand his knowledge and further his career.<\/p>\n<p class=\"wp-block-paragraph\">And that\u2019s what happened. The Bau Lab, located on Northeastern\u2019s Boston campus, provided Raj guidance in analyzing and interpreting LLMs\u2019 \u201cblack box,\u201d or the internal math and reasoning that lead from a prompt to a response. Raj quickly found ways to advance the lab\u2019s research and improve the team\u2019s understanding of this inner reasoning.<\/p>\n<p class=\"wp-block-paragraph\">\u201cNormally the internal monologue of an LLM is an unstructured \u2018chain of thought,\u2019\u201d Bau said. \u201cWhile working in our lab, Harsh built ways of training models to be able to follow specific instructions to make their internal thoughts easier to process and understand.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Raj is now working in a co-op position at Scale AI, a company that provides high-quality data and evaluations to AI labs, governments and Fortune 500 companies. His role there has him creating and publishing <a href=\"https:\/\/arxiv.org\/abs\/2607.28802\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">scientific research papers<\/a> on the failure taxonomies of LLMs, or the myriad reasons an LLM fails to deliver a user\u2019s expected result. To do so, he said, he investigates the flaws in newer models and suggests targeted improvements by experimenting with the models, reading scientific literature, attending conferences and exchanging ideas with industry experts.<\/p>\n<p class=\"wp-block-paragraph\">Raj said attempting to map out the next frontier of AI data is as fascinating as it is challenging. As AI laboratories exhaust human knowledge, teams are looking to target the collective knowledge of entire teams of people so that LLMs can not only perform the work of a competent coder, but a full start-up staff, he said.<\/p>\n<p class=\"wp-block-paragraph\">\u201cYou want to create data which is automating science, automating discovery,\u201d said Raj, who dismissed concerns over AI superseding the limits of human knowledge. \u201cThere are very few people on earth who actually have experience in that \u2026 so it\u2019s very hard to create data.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Raj said he plans to continue this exploration after his graduation using the skills picked up from his pioneering work during his Northeastern co-ops.<\/p>\n<p class=\"wp-block-paragraph\">\u201cI like to think that working at the edge of scientific knowledge about LLMs helped give Harsh a bit of confidence,\u201d Bau said, \u201cthat he could hold his own and make contributions in fundamental research.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"As AI large language models, or LLMs, grow in power and sophistication, where will their architects turn when,&hellip;\n","protected":false},"author":2,"featured_media":133169,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,181,64252,415,14954,134],"class_list":["post-133168","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-boston","tag-co-op","tag-llm","tag-nobel-prize","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/133168","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=133168"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/133168\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/133169"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=133168"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=133168"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=133168"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}