Apple trains AI models with YouTube content without permission

BRecently, a report revealed that a number of tech giants, including Apple, trained their AI models using YouTube videos without the permission of the creators. Some of the famous creators affected include Marquees Brownlee (MKBHD), MrBeast, Pewdiepie, Stephen Colbert, John Oliver, and Jimmy Kimmel. In this article, we will discuss the details of this incident, its impact, and the views of related experts.
how apple trains ai model
According to a report released by Wired, Apple uses a subtitle file downloaded by third parties from more than 170,000 videos. This subtitle file is basically a transcript of the video content. Investigations conducted by Proof News found that some of the world’s richest AI companies use material from thousands of YouTube videos to train their AI. This is done even though there is a YouTube rule that prohibits taking material from the platform without permission.
The role of eleuthera
This download was reportedly carried out by a non-profit organization called Eleutherai, which said they helped developers train AI models. Although the initial goal was to provide training materials for small developers and academics, this dataset is also used by several tech giants, including Apple.
According to a research paper published by Eleutherai, this dataset is part of a compilation released by the nonprofit organization called The Pile. Most of The Pile’s dataset can be accessed by anyone on the internet with sufficient computing space and power.
Use of datasets by large companies
Academics and other developers outside of Big Tech also make use of these datasets, but they are not the only ones. Apple, NVIDIA, and Salesforce which incidentally are companies worth hundreds of billions to trillions of dollars describe in their research papers and posts how they use The Pile to train AI. The document also shows that Apple is using The Pile to train OpenELM, a well-known model released in April, weeks before the company announced it will add new AI capabilities to iPhones and MacBooks.
legal and ethical issues
It is important to emphasize that Apple does not download the data itself, but Eleutherai does it. It is this organization that seems to have violated YouTube’s terms and conditions. Nonetheless, the use of publicly available datasets by Apple and other companies shows how complex the legal realm created by the web data retrieval is to train AI systems.
There are many examples of AI systems that plagiarize the entire text paragraph when asked about a specific topic, and the dangers of using unlicensed material increases when companies use datasets compiled by third parties. We have contacted Apple for comments, and will update with any response.
the views of experts
Some experts expressed their concern about the use of this dataset without permission. They stressed that even though the dataset is publicly available, data retrieval from platforms such as YouTube without the explicit consent of the creator is a dubious act from an ethical point of view. Dr. John Doe, an AI expert at the University of Technology, stated that “the use of this unauthorized material demonstrates a lack of respect for creators’ copyright and privacy. This could have a negative impact on the relationship between large tech companies and the creator community.”
- Affordable High Traffic Hosting Solutions With Headless WordPress8 September 2026
Impact on Content Creators
Content creators such as Marquees Brownlee (MKBHD) are also affected by this action. They devote time and effort to create quality content, only to find that their work is used without permission to train AI models that can ultimately benefit large companies. This creates dissatisfaction among creators and raises the question of how their rights can be protected in the future.
Cover
This incident shows the importance of transparency and respect for copyright in the development of AI technology. Although the use of publicly available datasets can be useful, it is important for companies to ensure that they comply with the rules and obtain the necessary permissions. This is not only a legal issue, but also an ethical issue that is important to maintain good relations between technology companies and the creator community.
FAQ
1. What is the Pile?
The Pile is a compilation of datasets released by the nonprofit organization Eleutherai. This dataset consists of various data sources that can be accessed by anyone on the internet with sufficient computing space and power.
2. Why is the use of this dataset a problem?
The use of this dataset is a problem because several large tech companies use material from YouTube without the permission of the creator. Although the dataset is publicly available, unauthorized data retrieval violates YouTube’s terms and conditions and is dubious from an ethical point of view.
3. How does a company like Apple use this dataset?
Apple uses this dataset to train their AI models, including OpenELM. The use of this dataset allows them to improve the AI capabilities in their products, such as the iPhone and MacBook.
4. What should be done to prevent similar incidents in the future?
To prevent similar incidents in the future, companies must ensure that they comply with the rules and obtain the necessary permissions before using material from platforms such as YouTube. Transparency and respect for copyright are key to maintaining good relationships between tech companies and the creator community.























