• 179 Posts
  • 214 Comments
Joined 3 years ago
cake
Cake day: June 12th, 2023

help-circle



  • if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

    This doesn’t have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.

    I can’t imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.

    We don’t know that because WMF hasn’t bothered to try defending contributors.

    Very false. Bot detection and blocking efforts have always been pretty documented

    I don’t see how bot detection shows that WMF is trying to defend contributors’ IP.

    Instead, they created a glide path for the pirates taking advantage of them.

    Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

    Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies – if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it’s better to not be in the shadows, if you are trying to evade detection.

    PS: Were they using the API or were they scraping? If they were using the API, why couldn’t WMF just revoke their keys?




  • To have legal standing for a letter, what you’re asking to cease and desist needs to be illegal.

    FWIW, that isn’t true - I didn’t post the letter, but I got a C&D from SoFi for this post. Clearly I had done nothing illegal.

    WMF isn’t Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months’ rain fund in case something massive happens to current revenue.

    They can’t defend the contributors, but they can hire a union busting law firm to keep their staff in check. Got it.

    we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists

    We don’t know that because WMF hasn’t bothered to try defending contributors. Instead, they created a glide path for the pirates taking advantage of them.

    so they might as well make the exploiters pay for it to help maintain the commons’ infrastructure.

    I think it is interesting you say that WMF is being “exploited” yet you believe that they don’t have any standing for the damages they have experienced.


  • Send one to whom? A WHOIS of the Brazilian IP that turns out to be a residential proxy? Anonymous scraping for LLMs is a problem all over the Internet without a solution (save Anubis). There’s a reason all the big companies have went for Cloudflare instead of any lawsuits, which you’d hear in the news.

    I know individuals that have been unmasked in torrent swarms and have had their ISP cancel their service due to that. The idea that Wikimedia’s hands are powerless to send a Cease and Desist to ISPs to warn and ban their customers for scraping is hilarious.

    Again, they have $296M in reserves. WMF can send a letter.

    They’re not even sure if they have the standing; they’ve looked at that question

    But that isn’t what you linked to says; the word standing doesn’t appear in the text, nor do they seem to explore that. Thanks for the reference, but it doesn’t actually support your argument.

    I am simply representing the perspective of contributors as a signatory myself, and as a developer who makes API calls to Wikipedia. “Our content is always free to use, but our infrastructure is not” sums it up nicely and Wikimedia Enterprise keeps the free infrastructure good.

    As I responded to another commenter:

    I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

    Enabling non-violating use cases don’t erase the violating ones.


  • This omits how scrapers and crawlers would have been getting the corpus for free without that, profiteering and clogging up the tubes that bring you Wikipedia and tons of amazing applications that rely on it.

    I didn’t know that I had to preemptively defend the Foundation.

    The reason I don’t think that that is particularly relevant is because these companies are violating the license that Wikimedia projects are distributed under. WMF has $296M in reserves. They couldn’t send a Cease and Desist for the server load?


  • BTW it’s already established that training AI models on copyrighted materials is legal without approval of the copyright holder, and IMO it’s unikely WP’s material would be an exception.

    That isn’t accurate - this is a highly unsettled question, and there are multiple cases in litigation today.

    WMF would need some better argument if they’d want to sue successfully, especially aginst companies that are sitting on 100x more money than them.

    A better argument than what - that they are openly violating the licenses under which the encyclopedia is distributed? What more do they need?










  • My actual sources for statements of fact were already listed. It’s a shorthand for “This is my personal position as an experienced editor that this is an overreaction” that you’re warping in bad-faith.

    I have no idea why you think I am “warping” my reaction in bad faith - my reaction is based on the license text and what is written in the post. I also think it is ironic that you ask me to extend you grace in accepting your shorthand, and you clearly don’t bother to accept mine (that your statement was a reference to your own authority as an experienced editor).

    Material published to Wikimedia is CC BY-SA 4.0; thus, those current Enterprise customers have every right to use the material basically however they see fit regardless of Enterprise.

    You don’t actually tell us why this is the case - I argue that the companies are violating the license by not licensing their derivative works reciprocally - you don’t even bother to respond to that and just posit that they have “every right to use the material however they see fit”. Do you really believe that? Are the CC-BY-SA and GFDL licenses just completely worthless?

    Wikimedia isn’t selling access to the material, because it’s literally nobody’s to sell; they’re selling access to stream the data on their servers which they host.

    I’m not sure how much that matters. Would Warner Brothers not have an issue with me “streaming” access to their movies via my home server for payment?

    I also don’t know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.

    Enabling non-violating use cases don’t erase the violating ones.


  • All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can’t restrict access to any use complying with that license.

    The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content’s license.

    Why do you think the license allows for big tech to produce derivative works that are not licensed CC-BY-SA?

    The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot.

    If WMF has no right to decide who can access it and who cannot, how can they sell privileged access to who can access it? Frankly, that assertion fails on its face.





  • Clearly, everyone realizes that the big tech AI pirates will scrape and stream the data - what I am objecting to is the response. When Google began to pirate Disney’s IP, Disney didn’t immediately offer them a data sharing deal (that also somehow doesn’t provide access to the IP) - they sued.

    The WMF has $296M in assets and they are grubbing for the pocket change that big tech throws at them for the fruits of the unpaid labor in Wikipedia. Why aren’t they suing for us? They are the stewards of the corpus.

    Has WMF even threatened the big tech owners with a good time, or were they simply salivating for the addition to their bottom line? Did they tell them to download the database? Did they attempt to ban their servers? Or did they simply provide big tech with privileged access to data they do not own?

    While clearly commentary from the community will be more interesting than from it outside of it, I don’t begrudge analysis or reporting from traditional media outlets - or do we want this all to be a private matter that 404 Media and the like don’t cover, allowing the theft (and union busting) to continue apace?

    In any case, I’ll show you mine.

    [Removed a reference to a different comment that I should have investigated more deeply.]






  • I think the reimplementation stuff is a separate question because the argument for it working looks a lot stronger, and because it doesn’t have anything to do with the source material having LLM output in it. Also if this method holds as legally valid, it’s going to be easier to just do that than justify copying code directly (which would probably have to only be copies of the explicitly generated parts of the code, requiring figuring out how to replace the rest), which means it won’t matter whether some portion of it was generated.

    Is it a separate question, though?

    Both works are copyrighted, one is just copyrighted as “all rights reserved” (our leaked commercial code) and the rest is licensed as LGPL. We’re putting both pieces of code inside the LLM and then asking the LLM to make a new version.

    What makes the action of leaking different from the act of putting it on the web? Rights are reserved in either case.

    If they aren’t entirely generated, you can’t make a full fork, and why would a partial fork be useful?

    Well, people are contributing to copyleft codebases expecting that when people build on their work, that work (the derivative works) are also licensed in the same way. You don’t need to fork for the value to be lost. People expected virality to be part of their contribution, and clearly the new derivative works are partially non-copyleft.

    Beyond that, as more of the codebase is LLM produced, the less of it is protected by the copyleft license, until we have a ship of Theseus situation where the codebase is available, but no longer copyleft. That is clearly not what was intended by e.g. the GPL. Just look at the Stallman quote in post.


  • ianal but does it even work like that? Is there any specific reason to think it does? I don’t believe you really get credit for purity and fairness vibes in the legal system. Same goes for the idea that code where it is ambiguous whether it is AI output could be considered public domain, seems kind of implausible, is there actually any reason to think the law works that way? If it did, then any copyrighted work not accompanied by proof of human authorship would be at risk, uncharacteristic for a system focused on giving big copyright holders what they want without trouble.

    I’m mostly just playing along with your thought experiment. As I said, we know that projects are already accepting LLM code into projects that are nominally copyleft.

    There is no way, leaks happen, big tech companies have massive influence, a situation where their code falls into the public domain as soon as the public gets their hands on it just isn’t realistic.

    If that is the case, is chardet 7.0.0 a derivative work of chardet, or is it a public domain LLM work? The whole LLM project is fraught with questions like these, but it seems that the vendors at least are counting on not copying leaked software and instead copying open source code that is publicly hosted.

    Why is it okay to strip copyright from open source works but not from leaked closed source works?

    We know that Disney is suing to protect its works - if it is true that LLM outputs are transformative, they should lose, as should any vendor whose leaked code was “transformed” by an LLM.