Kasey Roh, U.S. CEO of Upstage AI

Kasey Roh, U.S. CEO of Upstage AI

Upstage AI

On Tuesday, artificial intelligence company Upstage launched Solar Pro 4, its new commercial large language model, betting that the next phase of AI competition in fintech will be decided by more than which model scores highest on a benchmark.

The launch comes as financial institutions move AI from copilots and experiments into production systems capable of taking action. That transition is changing what enterprises need from the models underneath them.

Upstage is positioning Solar Pro 4 around behavioral reliability, which measures whether a model can consistently follow instructions across multiple steps, call the right tools, return information in the correct format and operate within company policies without expensive failures or retries.

“We’ve targeted a critical gap in agent performance that costs enterprises millions in waste,” Kasey Roh, U.S. CEO of Upstage, told me in an interview ahead of the launch.

For example, Uber burned through its entire 2026 AI budget in just four months. For fintech leaders, that kind of spending highlights a larger question around how enterprises should measure the value of AI once it moves into production.

“Enterprises approach technology as a tool,” Roh said. “When you evaluate a tool, ROI is the game. How advanced [the model] is, is a secondary question.”

Intelligence alone is not a measure of enterprise ROI. A model also has to work reliably inside the systems, policies and economics of the business deploying it.

That reality is beginning to reshape the broader AI market. OpenAI, Anthropic and other frontier model companies are building enterprise services and capabilities around their models, while specialized AI companies are competing on deployment, governance, reliability and industry-specific workflows.

In other words, the competition is expanding beyond who can build the most capable model to who can make that intelligence most useful inside an enterprise.

The Smartest AI Model May Not Be The Best Model For Fintech Enterprises

Imagine an AI agent helping process a financial transaction. It may need to retrieve information from several systems, follow a company’s internal procedures, call specific tools in a particular order, produce an output in an exact format and escalate the transaction to a human if certain risk conditions appear.

Getting the answer “mostly right” isn’t necessarily useful. The workflow has to work. That is the distinction Roh argues enterprises are beginning to understand.

A model can perform well on an intelligence benchmark while struggling with a seemingly simpler problem: consistently following the same set of instructions across a long, multi-step workflow.

Failures can be mundane. The agent calls the wrong tool. It returns an output that doesn’t match the schema another system expects. It forgets an instruction several steps into a task. Or it produces an unusable answer and has to run the process again. Each failure creates another inference, another API call and potentially another opportunity for something to go wrong.

For enterprises, AI performance therefore starts to look less like a leaderboard and more like an operating question: How reliably can the model finish the job?

Upstage designed Solar Pro 4 around that premise, emphasizing multi-turn instruction adherence, tool-call structure and behavioral reliability in agentic workflows.

The company is also pushing deeper into the U.S. market, where Roh sees an opportunity as enterprises move AI into production. Upstage already works with major enterprises including Samsung, SK and Hyundai. In April, it announced a $130 million first close of its Series C, bringing total funding to approximately $270 million and its valuation above $1 billion.

But the significance of that opportunity extends beyond one model company.

Fintech Doesn’t Have The Luxury Of Waiting

The pressure to figure this out isn’t coming only from inside financial institutions. Customers are adopting AI at extraordinary speed, too. McKinsey’s research on how AI is rewriting banking illustrates the gap.

Generative AI reached 45% of the U.S. working-age population within two years after ChatGPT launched in November 2022, compared with roughly 15 years for digital banking to reach the same level of adoption. By 2025, generative AI usage among working-age adults had reached 55%.

Consumers are also entrusting AI with consequential financial tasks. McKinsey points to customers using AI to find higher-yield savings accounts, refinance or consolidate debt, compare financial products and access financial advice. Its researchers expect agentic systems to push that behavior further as consumers become more comfortable allowing AI to act on their behalf.

That could fundamentally change financial competition.

Net interest income accounts for roughly 60% of retail bank revenue globally, according to McKinsey. Historically, many consumers have left deposits in convenient accounts rather than continuously optimizing for yield. An AI agent doesn’t have the same inertia. It can theoretically monitor rates, move idle cash toward higher-yielding accounts and return money when bills are due.

The same logic can extend to credit cards, lending, payments, rewards and wealth management.

For fintech leaders, this creates pressure from two directions at once: customers may adopt AI faster than institutions can safely deploy it.

Financial institutions therefore aren’t simply racing to adopt AI. They have to determine which systems can operate reliably inside environments where mistakes carry financial, regulatory and reputational consequences.

And that turns model selection into a risk decision.

For Fintech, AI Performance Is Becoming A Governance Question

Financial services has a much lower tolerance for unpredictability than many of the environments where generative AI first became popular.

A marketing team can regenerate a bad piece of copy. A financial institution may have to explain why an automated system made a decision, what information it used, whether it followed policy and who was responsible when something went wrong.

Kim Olson, Chief Risk Officer at Green Dot, described the challenge to me from the perspective of an institution responsible for managing risk that grows as AI becomes more autonomous.

“AI has the ability to amplify both good as well as bad,” Olson told me in an interview. “You can propagate mistakes and you can propagate risk unknowingly.”

For Olson, that makes monitoring, testing and human oversight essential as financial institutions deploy AI. “You want technology and AI to be an accelerant, but decision making has to be a human.”

For financial institutions, adopting AI does not eliminate the existing risk framework. It introduces another technology that has to fit within it. Leaders still need to understand where data goes, what a system is allowed to do, how decisions can be reviewed and where human intervention belongs.

That changes the enterprise AI buying conversation.

Olson doesn’t view risk as the opposite of innovation. She sees it as what allows institutions to move faster safely. “To go fast, you have to have good brakes,” she said. “It is important to go fast, but you can’t go fast at the cost of sustainability.”

That means determining when an institution can accelerate, when it needs to slow down and what protections have to exist before a new technology reaches customers. Accuracy matters. But so does data security, explainability, auditability, consistency, cost and the ability to maintain controls as the technology changes.

“If risk and compliance is an afterthought, it is really hard to put that genie back in bottle,” she said.

That is one reason specialized AI companies see an opening even as frontier models continue becoming more powerful. The competitive advantage comes from everything surrounding the intelligence: the deployment architecture, industry knowledge, governance systems and workflows that allow an enterprise to use it safely.

The Economics Of AI Change Once Agents Start Acting

There is also a financial reason enterprises should care about reliability.

A chatbot generally waits for a user to ask a question and generates a response. An agent may execute a sequence of actions before completing a single task.

One workflow might require multiple model calls, tool calls, searches and interactions with internal systems. If the model fails to follow an instruction halfway through the process, portions of that sequence may have to be repeated.

At small scale, a retry may look insignificant. Across millions of enterprise workflows, it becomes an operating cost.

This is where Upstage’s argument around behavioral reliability becomes particularly relevant. The company is competing on the total cost of successfully completing a task not just on what its model costs per token.

That is a subtle but important change in how enterprises may eventually calculate AI economics. The important question isn’t whether an AI system can demonstrate an impressive capability once. It’s whether an institution can trust that capability thousands or millions of times inside a production environment.

The cheapest model isn’t necessarily the cheapest system if employees constantly have to intervene, workflows fail or tasks repeatedly need to be rerun.

Likewise, the most intelligent model isn’t necessarily the most valuable if an enterprise needs to build significant infrastructure around it to make its behavior predictable.

The relevant unit of AI economics may eventually become less about cost per token and more about cost per successful outcome.

For fintech companies operating enormous transaction and servicing volumes, that distinction can compound quickly.

Fintech Leaders Still Need Humans Who Know How To Think

There is another layer to this transition that has little to do with the technical capabilities of the model itself.

As companies rely more heavily on AI, their employees still have to know when the system is wrong.

“I hope we don’t lose the ability to think critically,” Jyoti Menon, VP of Product at Bread Financial, told me recently on the Humans of Fintech podcast.

Menon has spent her career building financial products, including helping launch Apple Pay at Citi. Her concern isn’t that fintech leaders shouldn’t use AI. It’s that employees can begin accepting AI-generated answers without interrogating whether those answers make sense for their company, customers or specific context.

That becomes more consequential as AI gains agency. Menon argues that even prompting requires context and judgment.

“What story are you telling? How do you set the context? That’s like writing a paper,” she said. Her point connects directly to the technical and governance questions now facing fintech.

An institution can choose a more reliable model. It can build better controls. It can establish escalation procedures and governance frameworks. But humans still have to decide what the system should be allowed to do in the first place.

Menon raised a simple example: What happens if an AI agent makes a purchase and a consumer later tells their bank, “That wasn’t me, the AI did it”?

“What are the rules around that?” she asked.

Those questions aren’t hypothetical edge cases for long. They are the policy infrastructure that has to develop alongside the technical infrastructure.

The Next Fintech AI Advantage May Be Deployment

None of this means model intelligence has stopped mattering.

Frontier models will continue improving, benchmarks will continue moving and enterprises will continue evaluating the capabilities of the models available to them.

But intelligence is becoming one part of a much larger enterprise equation.

The companies ultimately capturing value from AI in financial services may be those that can connect models to proprietary data, embed them into existing workflows, govern their behavior, measure their economics and demonstrate that they operate reliably under real-world conditions.

That is also why the enterprise strategies emerging around AI matter.

Model providers are building more services around their intelligence. Infrastructure companies are developing orchestration and governance layers. Fintechs and financial institutions are determining which parts of their businesses should be automated and which decisions still require human judgment.

And providers like Upstage are betting that enterprises will choose models based on the jobs they need them to perform, rather than selecting one model for everything.

That makes Solar Pro 4 interesting – not because another LLM has entered an already crowded market, but because of what its positioning says about where that market is going.

The first era of generative AI rewarded models for showing us what was possible.

The next era will require enterprises to prove what is practical, safe, regulated, and compliant.

For fintech leaders, the competitive advantage won’t come simply from having access to the most intelligent AI. Nearly everyone will have access to increasingly capable models. It will come from knowing which intelligence to deploy, where to deploy it, how to govern it and when to trust it. That is why the future of fintech AI is about much more than the model.