
Google CEO Sundar Pichai addresses GenAI in his address to developers. (Photo by Camille Cohen / AFP) (Photo by CAMILLE COHEN/AFP via Getty Images)
AFP via Getty Images
The fast facts are that Google bid $10 million for data, beating AI training company Mercor’s $7.5 million offer, for roughly 100 million Spirit Airlines emails, 500 million Teams messages, 30 million lines of code and employee records reaching back to 1986. AI startup Micro1 lobbed in a late bid of $12.5 million after the deadline.
Micro1 is a young AI training startup, founded by CEO Ali Ansari, that matches talent and real-world data to companies building machine learning models. It submitted the late bid for Spirit’s data, topping Google’s $10 million offer, because it believes decades of real operational records are exactly what’s needed to train more capable AI models, and Ansari argued Google’s price was far too low for data that valuable.
The judge will rule on September 9, 2026.
The AI Data Wars Open With a Dead Airline’s Inbox
The AI data wars just found their strangest battlefield. It is in a bankruptcy auction for a dead airline’s email archive. Spirit Airlines stopped flying on May 2, 2026 after its second bankruptcy, leaving over 17,000 workers out of work, roughly $8.1 billion in debt and something no one thought to put a price on until this month which was decades of ordinary office communication. On August 14th, Google won the auction for the archive.
The Google bid may or may not hold as days later Micro1 sent Spirit’s lawyers a $12.5 million offer for the same data, but after the deadline for the bidding had passed. Passenger profiles and frequent flyer accounts are excluded from the sale.
The people who worked there are not.
Why The AI Data Wars Moved Beyond The Public Web
Why would anyone pay millions for a defunct company’s email? Because AI models learn by reading and the reading material is running out. For years, AI companies built their models on the public internet with sources like websites, books, forums, and code. That supply has a ceiling and a blind spot. According to BTUAI, only about 15% of the world’s knowledge has ever been digitized, and far less of it is searchable, which makes everything outside the public web the next data frontier.
Most of the world’s data isn’t digitized today. Estimates from BTUAI and other sources.
Sandy Carter and BTUAI
The public internet shows what companies say about themselves. It almost never shows how a company runs.
An archive like Spirit’s captures work as it happened.
For example, work captured might be a revenue team debating a fare change over email or a maintenance issues escalating through a Microsoft Teams thread, or even a marketing launch coming together across shared documents. This is rich data as it’s in the wild and real. If you are teaching AI to do the work of a real enterprise, that record of decisions, coordination and real mistakes is worth more than what’s left of the public internet. According to Bloomberg Law, Google says the data will improve its products and AI models.
The Fine Print of the AI Data Wars: De-identification
There is a privacy mechanism attached, and it deserves a plan explanation. Before Google receives anything, a third party will de-identify the archive. That means the process will strip out names, addresses and other details that point to a specific person. Google agrees not to reverse the process.
But here is the catch.
The sale agreement requires the links between records to stay intact, because a dataset where you can follow one employee’s thread across email, chat and files is exactly what makes it valuable for training. Those preserved links are also what worries the flight attendants’ union, which formally objected on the grounds that individuals a 17,000 person workforce could potentially be re-identified from the patterns. The bankruptcy judge pushed the approval hearing to September 9th, 2026.
Until then, Google has won the bid, not the data.
The AI Data Wars Move To Bankruptcy Court
Whatever the judge decides, the precedent is already visible.
A small market has formed this year around the data of failed companies, especially with startups selling their internal records and messages to AI firms so creditors can recover some of their investment.
Bankruptcy treats data as an asset of the business. So data is viewed just like a gate slot or a plane.
Spirit Airlines’s employee data and work conversations up for sale due to its Bankruptcy. (Photo by Joe Raedle/Getty Images)
Getty Images
This is partially because the law hasn’t caught up with the technology. Bankruptcy law was written in 1978, long before anyone imagined the value of data or the privacy concerns.
While judges rarely re-open a finished auction, the bankruptcy code does leave room for a late bid if it puts more money in the creditors pockets. So right now the hearing has two big issues which are:
the union’s privacy objection the Micro1’s extra $2.5 million bid. How Leaders Should Prepare for the AI Data Wars
Your company’s communication archive is now an asset along with any other data that your company owns. If that data can be sold to the highest bidder in a wind down situation for the benefit of your creditors, why not ask the uncomfortable questions now.
For instance,
What does your retention policy keep and for how long? What does your archive contain in it that needs extra privacy protection.For certain industries like healthcare, this matters the most. Do your employees and vendors know that their communications could outlive the company? What do your employment agreements and vendor contracts say about who owns the record of work and the data?
The first battles were fought over scraping the public web, and they played out in copyright suits. The next battles are over private archives of how businesses really work.
Spirit found out what its inbox was worth only after it died, and never grasped the privacy risks its employees faced. Is your data protected in these AI data wars?