> a digital copy with the ability of doing millions of copies is stored somewhere
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
>Plus having the info part of a LLM makes it immediately available to literally billions.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
> But, that little bit of data is a bit more data than existed before,
No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.
> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.
That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.
I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.
A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).
That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.
>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.
My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.
>That’s a false dichotomy.
I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.
> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost
I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?
It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.
Or be useful and scan book for Anna archive and other shadow libraries.
Citizen's lobbying against megacorp is a mirage.
>But it's _closer_ to being widely available, not farther.
By what metric ? The copy is now guarded by a company instead of being on the second hand market.
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction?
I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
You really drank all the Koolaid they had to offer.
Destroying the last copy of a book is OK because you shop for second hand books, and because some private company hold the last digital copy and have no incentive to make sure it survives. Damn..
> permanently locking human knowledge inside private corporate servers
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.
So physical books are rather non-permanent themselves.
I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.
Interesting read! I find the idea to use SQL queries to get only the relevant data quite convincing. I do wonder how this would work in practice though. Any changes the user makes would have to be inserted with SQL to allow for the new data to be included in SQL queries, but users also expect to be able to make changes and then not save them (or save them into a different file).
Should one make a massive transaction that is only committed when saving? It is possible to commit such a transaction to a different file when using Save As?
Or maybe for editing one would need to copy the file to a separate temporary location, constantly commit to that file, and when saving move the temporary file over the original file (this way we aren't losing the resilience against corruption SQLite offers).
Or is there a better way to do this? I don't like storing pending changes into the original file since it kinda goes against how users expect files to work (and could cause them to accidentally leak data).
You could insert any modifications and just mark whatever row the current saved one is
This would also work as a really crude undo tree
I don't really know if it actually goes against users expectations, Office kinda "saves" stuff for you and stores them as temporary versions anyway, to be presented in case you forgot to save
These example commits seem like pretty bad commit messages to me. They are just a summery of the changes (something a motivated reader can rediscover by reading the diff), while leaving out the why, which will be lost to history if not documented.
Yes, the why is very very important, but imo it's also useful with a one-liner summary of the actual change.
Consider a trivial change but affecting tens of places in the code, eg an API change. It's very useful to be able to quickly glance past "use new abc-API; required since dependency X bumped to Y" and mentally move on, rather than actually having eyes scanning over those actual changes.
I disagree with it being a minor issue. If I write a shell script around a program that accepts GNU-style arguments, I expect the following to be correct:
./cmd -a"$USER_CONTROLLED_DATA"
A program using this package would break that assumption, introducing a bug where this user-controlled data cannot start with an '='.
This syntax is supported by argparse and clap, the most popular argument parsers for Python and Rust respectively, and it seems to have caused almost no problems for them. It's a problem for the uutils implementation of cut, since `cut -d=` is common, but that's the only instance I could find after a long time scouring search engines and bug trackers and asking for examples.
If anyone does know of other examples or other places this has been discussed I'd love to hear it though, maybe I just haven't found them.
(Also, the more reliable way to write this in general is `-a "$USER_CONTROLLED_DATA"`, since that'll behave correctly if $USER_CONTROLLED_DATA is empty. As will `-a="$USER_CONTROLLED_DATA"` if you know the command supports it.)
In the Gentoo world, sometimes you need to give an exact package name which looks like `=net-misc/foo-0.1.2-r1`. The exact match has to start with the '='.
I think short options taking a value in the same argv (i.e. `-o=1` stuff) isn't a GNUism mostly because it's backwards-incompatible with POSIX. `=` is a valid getopt option character, `chmod` uses it.
That said, I think? 'nloomans means for USER_CONTROLLED_DATA to be a set of short flags, not flag values, as in:
root@08e9950d5bfd:/# export USER_CONTROLLED_DATA=lh
root@08e9950d5bfd:/# ls -a"$USER_CONTROLLED_DATA"
total 56K
drwxr-xr-x 1 root root 4.0K Mar 23 16:51 .
drwxr-xr-x 1 root root 4.0K Mar 23 16:51 ..
[...]
Not that I've seen this in the wild before. But everyone's use of bash is a unique personal hell anyway.
Anyway, one other alternative for the `cut` situation is to allow either ':' or '=' to optionally separate the key and the value. Then you can say `cut -d:=` or `cut -d=:` if you wanted to use either one. This is what https://github.com/c-blake/cligen does (for Nim, not Go).
The problem is existing shell scripts and muscle memory and command histories. `cut -d=` has always worked and works on all the other implementations so it should keep working if you switch to uutils.
Literally any short option key that takes a string or a char that could legitimately start with '=' has the problem, though, not just `cut`. The '=' will be "eaten" by one tool and left in another. But you know that. You write a bit as if we disagree, but I don't see any real point of contention. :-) Also, uutils has a very strict "drop-in" agenda. So, as another e.g., if you want `cp -t=foo` or `cp -S=foo` to work the same, you're going to have trouble if that '=' is eaten.
So, in this case that would seem to imply a problem for any utilities with options taking strings-or-chars, not merely `cut -d`. If uutils really wants to be strictly drop-in compatible, they may well need to roll their own option parser or twist the arm of whoever's they use to provide a mode for them.
In the more general case, "cross compatibility" may just always be limited by the reality that people just disagree on this stuff "more than they seem to think they do" (at least in my experience) and definitely more than they wish they did. I surveyed my /usr/bin once and like half of thousands of commands did not work with --help (yes, running that took some confidence in backups! but anyone could replicate on a throwaway VM or something). Consistency is nice, but consistency with what? -l=foo is consistent with --long=foo, but not (some, but not other) historical things.
I'm not sure there will ever will be a world in which you don't need to know which PLang/CL toolkit was used to make a CLI utility if you really want to know its syntax. The article's lib is going its own way from the Go stdlib. POSIX is pretty darn calcified. A 15 year old Python stdlib thing is unlikely to ever change in this regard. Python also allows "--beg" for "--beginning-long-option" if nothing else starts with "--beg" even back in its optparse days and that also tends to be controversial. cligen tools actually provide a --help-syntax. Maybe something like that could take off?
I can think of a lot of cases where it theoretically could be a problem, but `cut -d=` is the only one I've found so far where an end user ran into trouble because of this ambiguity, and I think it's the only one for which uutils bothers implementing a workaround. That's why I give it special attention.
> You write a bit as if we disagree, but I don't see any real point of contention. :-)
The `cut -d:=` spelling solves a different problem than the one I meant (and the one you're now talking about). But we're mostly on the same page!
It is safer to just put in the space (much like you put in the quotes to be safe). Python's argparse will also accept but not require an `=` separator (maybe optparse, too - I haven't checked that one).
It's only "safer" because argp has this particular bug. It's safer for argp (or python's argument parser, for that matter) to not have surprising buggy features like this.
While it may go against your personal expectations, I believe at least the Python variant is intentional not an accidental/unintended "bug".
"Surprise" is observer-relative. From a blank slate, if --long=val and --long val both work, then why not -s=val and -sval and -s val for a short -s?
So, I think the right word here is "disagreement" (perhaps about "which consistency - history vs. internal"), not throwing shade by simply declaring/asserting it a "bug".
I think because Claude said so and the author just copied it without checking. You can see it in the first screenshot of Claude's output:
> After examining the provided code, I've determined that this appears to be a CLI application for Claude code-related functionality, built as a CommonJS TypeScript application that has been compiled with webpack.
I tested the example using the TalkBack screenreader on Android. With Firefox I was able to select and click on the link, but it did not announce the email address. With Chromium it completely ignored the existence of the SVG email. I was unable to select it and it was like the email wasn't there at all.
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.