
AI Training Data Scraping: What the Suno Hack and Google Lawsuit Reveal

July 15, 2026 was a busy day in AI training data scraping news. Two separate stories landed within hours of each other. And while they came from different industries, music and publishing, they tell the same story.
A hacker breached Suno, one of the most popular AI music generators on the internet, and shared internal source code with 404 Media. The code showed that Suno had scraped at least 383,000 hours of music from YouTube Music, Deezer, Genius, and several stock music libraries. To do it, Suno purchased commercial proxy services to route requests through rotating IP addresses, bypassing YouTube's anti-bot detection and rolling cipher technology. That's not a robots.txt problem. That's a potential DMCA Section 1201 violation.
On the same day, Hachette Book Group, Cengage Learning, and Elsevier filed a class action lawsuit against Google, alleging that Google had used books from the Google Books program to train Gemini. Publishers had licensed those works for specific, restricted purposes. Google's own internal communications, cited in the complaint, described the practice as "highly problematic" and warned of potential fines between $10 billion and $100 billion.
Two cases. Two industries. One pattern: AI companies treating your rights as negotiable.
What Did the Suno Hack Actually Reveal?
The Suno breach happened in November 2025. Suno publicly disclosed it in July 2026. A hacker used a supply chain attack to access an employee's credentials and then pulled source code showing how Suno built its training datasets.
The numbers are precise. Suno's internal files documented scraped content by source: 113,879 hours from YouTube Music, 152,162 hours from a tagged YouTube Music dataset, 62,117 hours from Pond5, 19,514 hours from IMSLP, 17,615 hours from Genius, 12,287 hours from Deezer, 3,726 hours from Jamendo, 410 hours from Freesound, and 103 hours from Musescore Lyrics. In total, more than 383,000 hours of music. The same source code also revealed Suno sought to download roughly 1 million hours of podcasts, targeting 420,000 shows with at least five 30-minute episodes.
This is the rare kind of documentation that usually never surfaces. AI companies almost never disclose what they trained on or where they got it. The 404 Media investigation is one of the most detailed looks inside an AI training pipeline that has ever been published.
The most legally significant detail is how Suno obtained the YouTube content. According to the source code, Suno used Bright Data, a commercial proxy and data services company, to route scraping requests through rotating IP addresses and evade YouTube's bot detection systems. Suno also specifically searched for acapella versions of songs to isolate vocals.
Under DMCA Section 1201, circumventing a technical protection measure that controls access to a copyrighted work is a separate legal violation from copyright infringement itself. The Recording Industry Association of America made this argument in its existing lawsuits against Suno. The hacked source code now provides direct evidence to support it.
Suno previously admitted in court filings that it trained on "essentially all music files of reasonable quality accessible on the open internet." The hack gives that admission a detailed face: eight platforms, 383,000 hours, proxies, and circumvention.
What Did Google's Internal Documents Actually Say?
The Hachette, Cengage, and Elsevier class action lawsuit against Google, filed on July 10, 2026 in the US District Court for the Southern District of New York, contains a detail that stands out.
Google's own employees warned that using publisher-provided books for AI training was "highly problematic for Google," with potential fines in the range of "$10 billion to $100 billion." That warning is documented in an internal communication cited in the complaint. The employees knew. The company proceeded anyway.
The allegation is specific. Publishers had licensed their works to Google through the Google Books program, for purposes that did not include training generative AI. Google allegedly took those works anyway, using them to develop Gemini. The complaint also alleges that Google intentionally removed or altered copyright metadata from the ingested works, specifically to conceal that Gemini had been trained on materials the company was not authorized to use.
Gemini can now produce a 100-page murder mystery novel in approximately 20 minutes for $0.39. The publishers' argument is that this capability was built directly on their copyrighted works, without payment, and in violation of the agreements Google signed.
This is not the first major publisher suit filed in 2026. The same group of plaintiffs, Hachette, Cengage, and Elsevier, sued Meta in May 2026 on similar grounds. And Google faces separate ongoing litigation from publishers that predates this filing.
How Is This Different From a robots.txt Problem?
We have written before about why robots.txt cannot protect publisher content. That post covers crawlers that ignore a voluntary text signal. The Suno and Google cases are different in kind, not just degree.
Suno did not ignore a robots.txt file. Suno built technical infrastructure to actively route around a platform's security systems. Purchasing Bright Data proxy services to circumvent YouTube's rolling cipher is not an oversight. It is a deliberate engineering decision made by a team that understood what they were doing.
Google did not ignore a signal either. Google had a formal agreement with publishers. The question in the lawsuit is whether the use of licensed content for AI training exceeded what the agreement permitted. Google's internal documents suggest the company believed it did, and proceeded regardless.
As of May 2026, more than 20 active lawsuits assert DMCA Section 1201 circumvention claims against AI companies, a shift from pure copyright infringement theory to the separate question of whether technical protections were bypassed. Courts are still working through the standard. YouTube's rolling cipher and rate-limiting systems are a much stronger Section 1201 argument than robots.txt, which a federal court has already ruled does not constitute a technical protection measure under the law.
The distinction matters for publishers. If Section 1201 claims succeed, the legal test shifts from "who ignored your preferences" to "who bypassed your controls." Publishers with genuine technical access controls are in a better legal position than publishers with robots.txt files.
Why Lawsuits Are Not a Protection Strategy
Lawsuits are reactive. The Suno scraping happened in 2023 and 2024. The source code confirming it became public in July 2026. The litigation is just beginning. By the time courts settle these questions, the content has already been consumed, the models have already been trained, and the AI products built on that content are already competing with the publishers who created it.
Google's Gemini is generating novels for $0.39 each. Suno is generating music. The legal process that might eventually compensate the original creators will take years and deliver uncertain results.
The AI licensing revenue that publishers are actually earning comes from companies that agreed to pay before accessing content, not from companies that took content and are now being sued. People Inc. grew AI licensing revenue 26% in Q1 2026. News Corp holds $400 million in AI commitments. Those numbers come from structured access with clear terms, not from litigation.
Publishers who want to be in that group need to think about access control differently. The question is not "how do we sue faster after scraping happens?" The question is "how do we make scraping technically impossible in the first place?"
What Infrastructure-Level Access Control Actually Means
Real access control is not a better legal agreement. It's not a stronger robots.txt. It's not a more aggressive Terms of Service.
Real access control means that your content does not exist at a public endpoint. There is no URL that a proxy can hit, no stream that can be ripped, no file that can be downloaded without authenticated, authorized, metered access. Unauthorized use is not a legal risk to be managed after the fact. It is a technical impossibility before the fact.
Alien Intelligence's data streaming infrastructure deploys directly on your servers. Content never leaves your infrastructure unless a valid, authenticated request has been made, authorized, and logged. Every retrieval is tracked, attributed, and billed. The same architecture that makes unauthorized access impossible is the one that turns authorized access into a revenue stream.
This is the gap between what Suno's victims had and what they needed. YouTube had technical protections. Suno bought commercial proxies to work around them. The proxies worked because the technical measures were on YouTube's infrastructure, not on the content itself. Content-level access control closes that gap. When the content only exists inside an authenticated data layer, there is no public stream to rip and no rotating IP trick that gets around it.
AI content access control at the data layer is also what makes an AI licensing agreement commercially viable. When every query is logged and billed, you can audit compliance in real time. You don't have to wait for a hacker to breach a company and share source code to find out what was taken.
The GEO and content monetization strategy that works long term is not built on hoping AI companies behave. It's built on infrastructure that doesn't give them a choice.
Conclusion
Two stories broke this week that prove something publishers have suspected for years. AI companies are not just scraping the open web and ignoring voluntary signals. They are circumventing technical protections with commercial tools, and using legitimate access agreements as cover for uses those agreements never authorized.
Suno scraped 383,000 hours of music and used paid proxy services to bypass platform security. Google allegedly used books provided for one purpose to train a model for a completely different one, with its own employees warning in writing that the exposure could reach $100 billion.
Lawsuits are now moving forward in both cases. They will take years and produce uncertain results. The content has already been consumed.
Publishers who want protection that works before the next breach or the next lawsuit need to build systems that don't rely on AI companies acting in good faith. Authenticated access. Metered retrieval. Technical enforcement at the infrastructure layer, not at the legal layer.
That is not a legal strategy. It is the only approach that works before the scraping happens.
Frequently Asked Questions
What did the Suno hack reveal about AI training data scraping?
A November 2025 hack of Suno, disclosed publicly in July 2026, exposed internal source code documenting how the company built its training datasets. The code showed Suno scraped more than 383,000 hours of music from YouTube Music, Deezer, Genius, Pond5, IMSLP, Jamendo, Freesound, and Musescore Lyrics, plus roughly 1 million hours of podcasts. It also showed Suno used commercial proxies from Bright Data to bypass YouTube's bot detection and rolling cipher, routing requests through rotating IP addresses to evade detection. This is the most detailed documented look inside an AI music training pipeline yet published.
Did Google violate its agreement with publishers in the Hachette lawsuit?
That is the central allegation in the class action filed July 10, 2026. Hachette Book Group, Cengage Learning, and Elsevier allege that Google used books licensed through the Google Books program to train its Gemini AI models, beyond what the licensing agreements permitted. The complaint cites internal Google communications describing the practice as "highly problematic" with potential fines of $10 billion to $100 billion. Google allegedly also removed or altered copyright metadata on ingested works to conceal their origin. Google has not yet responded to the lawsuit in court.
What is DMCA Section 1201 and why does it matter for AI scraping?
DMCA Section 1201 prohibits circumventing a technical measure that controls access to a copyrighted work. It's separate from copyright infringement and carries its own statutory damages. As of May 2026, more than 20 active lawsuits assert Section 1201 claims against AI companies for bypassing YouTube's encryption, rate limiters, and bot-detection systems. Courts have already ruled that robots.txt does not qualify as a technical protection measure under Section 1201. But platform-level encryption and rolling ciphers are a stronger argument. If Section 1201 claims succeed in these cases, publishers with genuine technical access controls will be in a much stronger legal position than those relying on robots.txt.
Why can't publishers just sue AI companies that scrape their content?
Lawsuits are reactive. The Suno scraping documented by the July 2026 hack happened in 2023 and 2024. By the time the litigation resolves, the training data has already been consumed and the AI products built on it are already competing with the original creators. Legal strategy is not a substitute for access control. Publishers who have already structured access via authenticated, metered infrastructure are earning licensing revenue from compliant AI platforms, while publishers pursuing litigation are still years from any outcome.
What does real protection from AI training data scraping look like?
Real protection means your content doesn't exist at a public endpoint that a scraper or proxy can reach. Infrastructure-level access control requires every retrieval to be authenticated, authorized, and logged before the content is delivered. With this approach, unauthorized use is technically impossible before the fact, not legally pursued after the fact. This is what separates perimeter defense, which Suno bypassed using commercial proxies, from content-level access control, where there is no public stream to rip and no rotating IP trick that works around it.



