> As my regular readers will be aware, I have been engaged, on behalf of my clients, in legal combat with Internet censors around the world, agencies who think that they have both the standing, the competence, and the right to tell American companies what software they can write and run, for the better part of 18 months.
Is this supposed to make me trust their opinions on tech regulation?
David Sacks is one of the most predictable VCs in Silicon Valley. He could have written "I still have the same thoughts; here's a link to my first and only opinion" and the post would have communicated the same amount of info.
Terrence Tao recently published an interesting take that zooms out from the details of the Navier Stokes drama.
Its an interesting observation he makes, because it is not dissimilar from the relatively common phenomenon of one academic lab getting scooped by another lab (usually by coincidence).
> one academic lab getting scooped by another lab (usually by coincidence).
I was with you up to “usually by coincidence”.
There’s a long and sordid history in areas of chemistry and areas of biology of holding up a competing paper in review so you can scoop them. I’m sure it exists in physics as well. Certainly biophysics, but probably most subfields.
Often it’s a famous labs that can steamroll review or even just dump the work into PNAS as a “member contribution.”
At least one author of a famous inorganic chemistry textbook was rumored to do this routinely.
And I know of at least one National Academy member who swore off arxiv prepublication after getting scooped.
None of this makes it all right. But plagiarism and academic theft is old and definitely not always accidental.
I know this happens, but my impression during my phd was that many labs use popular methods to test popular questions, leading to a lot of simultaneous work. You see this in history as well.
This is interesting. He seems to imply that the AI will just provide the solution. He's concerned that the process of getting to that solution is the important bit. But I can see two versions of "process".
A.) The actual steps of the proof, which I assume the AI would provide.
B.) People, while working toward a solution, finding novel properties/methods along the way that open new avenues of research + new open problems.
Does B.) actually happen? Would knowing the solution to a problem stymie the process of finding new open questions? I would assume finding a solution may unlock other problems too. So maybe on balance it's not really bad?
I guess in the end, I'm just making the obvious case "the future is uncertain in the face of AI".
Well the argument is that B is no longer sustainable because of the scenario you describe in A. If you do a bunch of work on Navier Stokes but then OpenAI gets all the press, then what was the point?
Why is a "normie" better off if he hyperventilates like this? In that scenario, they would be screwed AND anxious. If it really is as transformational as you say, then no amount of preparation or awareness matters. You are infinitesimally more ready then they are. Luckily for all of us, there is more to knowledge work then technical implementation.
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
Again, see the Apple lawsuit. OpenAI has never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that this time is different (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and says a pattern is continuing, I’m giving them the preliminary benefit of doubt. It’s a bit extreme to conclude based on that. But it’s far more baseless to swing to the other side and claim we need to clear the table for a serial offender.
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.
Old people seem to be surviving just fine in areas where property tax is based on current market value? It sucks to move, but luckily they have an extremely valuable property to sell
Young people seem to be surviving just fine in areas where the market value is lower. Sucks to not live in the heart of the city but luckily they have their entire lives ahead of them to accumulate the wealth required.
yes, the fundamental BS assumption of algorithmic social media is that human desire and curiosity is legible, and can be calculated in a straightforward way.
A lot of postmodern maximalist books like IJ try to reckon w/ these problems. Its easy to reflexively hate the meaningless, ironic, media-fried worlds they build, but they are mirrors, not instruction manuals.
One of IJ’s most important topics is the critique of the empty destructive postmodern cynicism that you are referring to. Wallace was in fact trying his best to do more than just reflect. One of IJ’s main and most likeable characters, Mario, is basically a walking sentiment of this. Infinite jest is surprisingly preachy in this regard; “this is water” and some of his other essays are even more so in this regard. that’s why literary critics argue that Infinite Jest doesn’t belong to the postmodern, but rather to the post-postmodern, sometimes also called the metamodern, or the new sincerity. Like a hybrid of modern and postmodern that oscillates between the hip and the banal, between cynicism and hope, fiction and truth.
Here is Wallace on this subject:
“Irony and cynicism were just what the U.S. hypocrisy of the fifties and sixties called for. That’s what made the early postmodernists great artists. The great thing about irony is that it splits things apart, gets up above them so we can see the flaws and hypocrisies and duplicates. The virtuous always triumph? Ward Cleaver is the prototypical fifties father? "Sure." Sarcasm, parody, absurdism and irony are great ways to strip off stuff’s mask and show the unpleasant reality behind it. The problem is that once the rules of art are debunked, and once the unpleasant realities the irony diagnoses are revealed and diagnosed, "then" what do we do? Irony’s useful for debunking illusions, but most of the illusion-debunking in the U.S. has now been done and redone. Once everybody knows that equality of opportunity is bunk and Mike Brady’s bunk and Just Say No is bunk, now what do we do? All we seem to want to do is keep ridiculing the stuff. Postmodern irony and cynicism’s become an end in itself, a measure of hip sophistication and literary savvy. Few artists dare to try to talk about ways of working toward redeeming what’s wrong, because they’ll look sentimental and naive to all the weary ironists. Irony’s gone from liberating to enslaving. There’s some great essay somewhere that has a line about irony being the song of the prisoner who’s come to love his cage.”
I think it’s fair to say that Wallace is postmodern but less self-satisfied and complacent about it. The elements of irony and cynicism are still there, but insufficient. So he synthesizes a “what next?”, which is labeled metamodern.
reply