{"id":64565,"date":"2026-08-01T23:54:24","date_gmt":"2026-08-01T23:54:24","guid":{"rendered":"https:\/\/www.europesays.com\/spain\/64565\/"},"modified":"2026-08-01T23:54:24","modified_gmt":"2026-08-01T23:54:24","slug":"when-speed-meets-quality-how-we-make-ai-products-measurable-at-santander","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/spain\/64565\/","title":{"rendered":"When speed meets quality: How we make AI products measurable at Santander"},"content":{"rendered":"<p>This is the idea at the centre of the work: Responsible AI is not a separate checklist bolted onto quality, it is part of what quality means. A product that shows bias, blocks a legitimate request, or can be talked into the wrong answer is not a high-quality product, however fast or fluent it looks. So, safety and fairness are scored in the same place as accuracy and performance, in a single view of quality. If a product falls short on Responsible AI, it cannot be considered good.<\/p>\n<p>The need for this comes from the nature of AI itself. Traditional software is deterministic: the same input gives the same result unless the code changes. Products built on large language models (LLMs) and agents behave differently. Their outputs can depend on wording, context, the sources they retrieve, the orchestration around them and updates to the underlying models. A product can work well at launch and then drift because a model was updated or a data source changed.<\/p>\n<p>This is why evaluation is treated as part of development rather than a gate at the end of it. Regression testing still matters, but confirming a system has not got worse is not the same as showing it is good. Building AI products is increasingly a loop \u2014 generate an output, assess it, improve it, evaluate again \u2014 and the framework is built for that loop.<\/p>\n<p>It starts from the product itself. Rather than applying a generic test set, it reads what the system is meant to do and generates targeted evaluation cases from its own requirements and policies. These evaluations include its Responsible AI obligations, so fairness, safety and guardrail compliance are tested with the same rigour as functional correctness. It then probes behaviour beyond those cases, to find weaknesses no one thought to anticipate, while a curated set of expected behaviours gives a stable reference point so regressions show up when prompts, models or data change. The same logic runs before and after deployment: an end-to-end quality check ahead of release, and continuous monitoring in production for drift, degradation and emerging vulnerabilities \u2014 one consistent definition of quality, Responsible AI included, across a product&#8217;s life.<\/p>\n<p>Underneath the technology, the ambition is cultural. Engineers should be able to move fast while understanding the impact of each change. Product teams should prioritise evidence. And Responsible AI and governance teams should rely on the same metrics as the people shipping the product \u2014 so safety and fairness are measured in the same breath as performance, not bolted on afterwards.<\/p>\n<p>For the customer, none of this is visible, and that is the point. People do not ask for \u2018Responsible AI\u2019; they ask for a product that answers correctly, holds a normal conversation, and treats them fairly. Good quality disappears into a good experience \u2014 and so does good Responsible AI. When it works, no one notices it; they simply trust the product. As AI becomes a larger part of how the bank operates, that trust cannot be something we assume is there. It has to be observable, measurable and continuously improved. That is what this capability is for: not simply better testing, but a more responsible, more reliable way of building and scaling AI products.<\/p>\n","protected":false},"excerpt":{"rendered":"This is the idea at the centre of the work: Responsible AI is not a separate checklist bolted&hellip;\n","protected":false},"author":2,"featured_media":64566,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[121],"tags":[144,485,442],"class_list":["post-64565","post","type-post","status-publish","format-standard","has-post-thumbnail","category-banco-santander","tag-banco-santander","tag-innovation","tag-stories"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/posts\/64565","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/comments?post=64565"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/posts\/64565\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/media\/64566"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/media?parent=64565"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/categories?post=64565"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/spain\/wp-json\/wp\/v2\/tags?post=64565"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}