It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free.
I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves
Well, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post)
No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
It should be no surprise that it's a movement of lying thieves and scammers, Effective Altruism, that's behind close-source AI in the US.
Now I'm not sure justice won't come: SBF is behind bars for 25 years. He could turn out to not the be last scum from the EA movement to end behind bars.
As to open-weights models: at least it's not sold back at a mark-up and anyone can run them.
IMO if they didn’t have proper licensing to train on the data the model should not be copyrightable.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
Try doing something about actual evil shit like arms manufacturers and the politicians ordering the deaths of millions from the comfort of their sofas.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress; there's just too much shit to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property”, if you believe that is a thing, but not of the “labor” that went into it.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we such up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
56 comments:
It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
Regulation that said something like “we own 50% of your profit or 20% of your revenue, whichever is the larger” would.
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves
really not the same entities here
Well, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post)
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
Yeah that's more fair
Not at leaf level, but if you trace the trunk, pretty sure you end up on the same one.
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
I mean if you cite a copyrighted book verbatim. You are held liable. So should a company producing copyrighted work.
For instance a image/video generating model.
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
The what thread, now? Got a link instead?
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
I am unclear how this would help anyone?
Any argument that writers and artists lose from these existing, would remain unchanged.
Copying data isn't a crime.
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
What is DD supposed to mean?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
Due diligence — pretty standard acronym in this community.
No it's not, and I've been here more than a decade.
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
At least the Chinese AI companies are doing good service open-sourcing their models back to the public.
There are no open-source LLMs.
> Crimes this large are crimes against humanity.
It should be no surprise that it's a movement of lying thieves and scammers, Effective Altruism, that's behind close-source AI in the US.
Now I'm not sure justice won't come: SBF is behind bars for 25 years. He could turn out to not the be last scum from the EA movement to end behind bars.
As to open-weights models: at least it's not sold back at a mark-up and anyone can run them.
IMO if they didn’t have proper licensing to train on the data the model should not be copyrightable.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
> sell it back to us at a mark-up.
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about actual evil shit like arms manufacturers and the politicians ordering the deaths of millions from the comfort of their sofas.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress; there's just too much shit to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
Defending the same entities getting large DoD contracts to use AI for killing?
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
“Information wants to be free“.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property”, if you believe that is a thing, but not of the “labor” that went into it.
In case of programming.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
But it's copying...how is it theft? Your labour WASNT stolen was it?
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
Never ended, just changed in form.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we such up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
How is this different from Microsoft scraping to build Bing?
Honest question. There is a line in the sand somewhere apparently.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
Huge difference between building AI and a search index.
Why? In both cases the SaaS downloaded the whole web and derives 100% of revenue from content they didn't make.
I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
Lots of the original content is no longer available. Bots kill sites, AI kills monetization - both results in the original material disappearing.
Copying means we can both share in the knowledge, surely everyone on HN wants that right? Share the open source code for the good of everyone?
Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
Probably indicative of America's wider downfall that they've all become so self interested
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
they did say theft of labor, not theft of the things being trained on
Copying isn't stealing you babies
And here we are, just watching and doing nothing..