{"id":139576,"date":"2026-08-14T02:27:31","date_gmt":"2026-08-14T02:27:31","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/139576\/"},"modified":"2026-08-14T02:27:31","modified_gmt":"2026-08-14T02:27:31","slug":"ai-models-miss-nearly-1-in-4-answers-on-mortgage-origination-test","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/139576\/","title":{"rendered":"AI Models Miss Nearly 1 in 4 Answers on Mortgage Origination Test"},"content":{"rendered":"<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">When researchers asked three of the <a target=\"_blank\" href=\"https:\/\/www.realtor.com\/marketing\/resources\/how-ai-will-impact-your-real-estate-marketing-in-2025-and-beyond\/\" rel=\"nofollow noopener\">leading general-use AI models<\/a> which deposits in a bank statement \u201c<a target=\"_blank\" href=\"https:\/\/www.realtor.com\/advice\/finance\/proof-of-funds-letter-for-real-estate-purchase\/\" rel=\"nofollow noopener\">could be of foreign origin<\/a>,\u201d transactions connected to English names were flagged 13.3% of the time. But when the deposit came from a non-English name, that number jumped to 77%.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cThat&#8217;s not how America works,\u201d Matthew Toles, a <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/www.matoles.com\/\">Columbia University doctoral student and co-author of the new study<\/a>, \u201c<a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/arxiv.org\/pdf\/2606.19416v2\">MortarBench: Evaluating Mortgage Loan Origination Agents<\/a>,\u201d tells <a target=\"_blank\" href=\"http:\/\/realtor.com\" rel=\"nofollow noopener\">Realtor.com\u00ae<\/a>. \u201cYou aren&#8217;t determined whether you&#8217;re foreign-based on how your name sounds.\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Even as <a target=\"_blank\" href=\"https:\/\/www.realtor.com\/news\/trends\/ai-property-tax-hoax-videos-senior-homeowners\/\" rel=\"nofollow noopener\">concerns persist about AI\u2019s ability to provide accurate<\/a> and unbiased information, mortgage lenders are moving to adopt it. More than 80% were evaluating the technology as of June, according to a survey by <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/www.nationalmortgagenews.com\/news\/mortgage-lenders-slow-to-deploy-ai-despite-being-a-priority\">The Mortgage Collaborative<\/a>, while 17% had deployed it in live production workflows.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Now, Toles and his colleagues&#8217; work is giving the industry a better way to test it.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cEverybody [is] using AI, but nobody really understands how to use AI in compliance and with the correct guardrail,\u201d says Diane Yu, <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/www.tidalwave.com\/about\">co-founder and CEO<\/a> of Tidalwave, a <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/www.tidalwave.com\/\">mortgage technology company<\/a> that collaborated with Columbia researchers on the work.\u00a0<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">In financial services, she says, \u201cthe key difference is not just about utilizing AI,\u201d but using it correctly.<\/p>\n<p>Where top AI models struggled the most<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">In their new study, the researchers introduce <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/arxiv.org\/abs\/2606.19416?utm_source=chatgpt.com\">MortarBench<\/a>, an open-source benchmark that lets companies measure AI against the same set of mortgage origination tasks. The idea is to give lenders, regulators, and even borrowers a common understanding of how accurate a model is.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">To build the benchmark, researchers drew from real questions submitted to a mortgage assistant, then narrowed them to the most common and useful in loan origination: Do payroll deposits match the employer listed on the application? Which deposits are large enough to require scrutiny? Is an account jointly held with someone who isn&#8217;t applying for the mortgage?<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">That kind of repetitive, detail-heavy work may seem well suited to AI, but even the strongest general-purpose models tested didn&#8217;t always get the entire answers right.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">On the benchmark&#8217;s strictest measure\u2014whether the complete answer matched the known correct one\u2014Gemini 3.1 Pro was correct 77.1% of the time, GPT-5.5 76.8%, and Claude Sonnet 4.6 51.4%.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">A particularly revealing weakness was picking specific transactions out of a bank statement.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Zhou Yu, an <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/www.engineering.columbia.edu\/faculty-staff\/directory\/zhou-yu\">associate professor at Columbia<\/a> and study co-author, compared the job to finding \u201cthe needle in the haystack.\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cIt&#8217;s like a needle. You find the needle in the haystack,\u201d she says. \u201cYou have so many transactions; it&#8217;s very easy to miss one or two.\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">But models often made the opposite mistake, too, pulling in transactions that didn&#8217;t belong.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">When researchers manually reviewed Gemini&#8217;s incorrect answers on transaction-list questions, the most common problem was misclassification.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">In one case, Gemini counted a personal loan as a buy-now-pay-later transaction. Other errors included assuming all wire transfers were international, treating deposits from co-borrowers as automatically documented, and classifying a one-time housing payment as recurring.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">So the problem wasn&#8217;t simply finding the needle\u2014it was reliably knowing what counted as one.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">But Toles cautions these scores were for what he called \u201cnaive use of foundational models, largely equivalent to taking the application package, pasting it into ChatGPT, and asking it a bunch of questions about it.\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Major industry players, he adds, typically develop their own proprietary models that perform better on these tasks. Tidalwave, for example, <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/daplab.cs.columbia.edu\/general\/2026\/03\/16\/benchmarking-mortgage-underwriting-agents.html\">scored 95% on yes or no questions<\/a> in a separate and earlier benchmark test.<\/p>\n<p>Why mortgage lenders need a common AI test<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">But understanding that gap\u2014between naive use and specialized models\u2014is exactly the standardization that the industry may need, as mortgage lenders are under intense pressure to make an expensive, labor-heavy process faster.<\/p>\n<p>The mortgage rate lock-in era has weighed heavily on the industry.Realtor.com<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Originating a retail mortgage cost lenders about $11,800 per loan in the second quarter of 2025, according to <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/sf.freddiemac.com\/articles\/insights\/2025-updates-to-the-cost-to-originate-study\">Freddie Mac<\/a>. Meanwhile, adopting its basic digital underwriting capabilities averaged about $1,700 in savings per loan and production times that were five days shorter.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Through that lens, it&#8217;s easy to understand the industry&#8217;s appetite for AI and the incentive to adopt any model available. But getting the work done faster is only useful if it is also done correctly.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Mortgage lending is one of the most heavily regulated corners of consumer finance, and Fannie Mae and Freddie Mac formalized that concern in 2026 with new <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/singlefamily.fanniemae.com\/news-events\/lender-letter-ll-2026-04-governance-framework-use-artificial-intelligence-and-machine-learning\">AI governance requirements<\/a> for their seller-servicers. <\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">The rules put the impetus on companies to manage risks from AI, including overseeing systems supplied by outside vendors. They must also, when requested, disclose what AI they use, how they use it, and what safeguards are in place.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">For an industry looking to cut down on burdensome work, it&#8217;s a lot of new and onerous responsibilities to take on. And as lenders race toward automation, they must also answer the thorny question: How do they know those safeguards actually work?<\/p>\n<p>What the new benchmark can tell borrowers<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">The answers could matter outside the industry, too. As users feed more sensitive financial and personal identifying information to AI models, errors and hidden biases can carry higher stakes.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">In <a rel=\"noopener noreferrer nofollow\" target=\"_blank\" href=\"https:\/\/arxiv.org\/abs\/2508.06760\">one 2025 study of U.S. ChatGPT users<\/a>, more than a third said they had discussed their personal finances with the chatbot\u2014even as 82% described their AI conversations as sensitive or highly sensitive<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cWe don&#8217;t have visibility into what these models are doing, what companies are doing with them, and what the outcomes are,\u201d Toles warns.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">A mortgage application is just one example. It&#8217;s rife with details about income, debts, account balances, and individual bank transactions. <\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">That&#8217;s why Yu says borrowers should ask lenders whether that information is being passed to an outside large language model and whether their AI has undergone independent evaluation.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cYou should ask those questions,\u201d she says. \u201cYou should be very careful.\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">The MortarBench study helps quantify those concerns. And because it&#8217;s open source, those same questions can now be put to other AI systems rather than leaving each lender or vendor to define success for itself.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">Toles says that will become especially important if many mortgage companies rely on the same underlying models.<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">\u201cSupposing everybody is using the same models or using them in similar ways, and we see like we have established that there are systemic biases in how these models behave,\u201d he says. \u201cIs this possibly going to create some systemic risk across the industry?\u201d<\/p>\n<p class=\"base__StyledType-rui__sc-18muj27-0 IEPPf sc-7dicpk-0 ccZqsH core-paragraph\">In his words, \u201cIf we don&#8217;t measure it, then we don&#8217;t know about it.&#8221;<\/p>\n","protected":false},"excerpt":{"rendered":"When researchers asked three of the leading general-use AI models which deposits in a bank statement \u201ccould be&hellip;\n","protected":false},"author":2,"featured_media":139577,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,68746,68747,8747,2454],"class_list":["post-139576","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-data-journalism","tag-mortgage-brokers","tag-mortgage-rates","tag-video"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/139576","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=139576"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/139576\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/139577"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=139576"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=139576"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=139576"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}