Claude Executive: Damn, we hardly know who our users are, can't we just force them to say their full name or ban them?
Claude Product Manager: No, that'll piss people off too much, and sadly we can't just ask for ID either...
Claude Executive: There must be some way we can force people to link their government IDs with our platform so our analytics get better and more accurate?
Claude Product Manager: We could limit the platform to 18+ and use "Age Verification" as the reason for people to hand over IDs, seems other platforms had success with this approach
Claude Executive: And we hardly have any users younger than 18 anyway, go for it!
I'm imagining a future where a bunch of bizarre laws interact oddly (as they do), and now we've got websites with unnecessary nudity pasted in the corner.
"Oh, those? Those are just compliance tits. Ignore those. It's just a thing that came a few years after we finally got rid of the cookie banners. The companies wanted certain protections awarded only to 18+ sites. But you can't just declare yourself an 18+ site, so some sites post the most minimal amount of imagery that constitutes erotic nudity. That's why Google's graphic for the past few months has just been that one with the two dots in the middle of the o's."
Key for me is to "group" cables. For example, I have a bag of USB-C cables, a bag of USB-A cables, etc.
Grouping them is key to deduplicating.
It's easy to look at a single legacy USB A-to-B "printer" cable in isolation and think "I might need this someday!" Because you really might need it someday. However, if you group them you might see that you have ten of them. And then you can get rid of... maybe 8 of them.
I also (mostly) put individual cables into baggies. You can get clear 2mil generic ziploc style baggies for super cheap on Amazon or elsewhere. $15 for 200 or something. Prvents tangles and way less effort than wrapping or tying them.
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Software development untethered from the practical realities of the customer / user is what drives people insane.
When developers are required to interact with the customer on a regular basis, the freewheeling effects described in this article are damped massively.
The potential for insanity goes off the charts when the development team is siloed away in solitary confinement and the only interactions with the client occur via some prison guard known as "project manager" sliding notes under the door.
Working with the customer sometimes sucks. Just like exercise and eating vegetables sometimes suck. It's a temporary unhappiness that keeps us grounded in reality.
Numbers from GPT Astra
- Shopify has 3000 engineers as of 2026
- Google Chrome when released in 2008 conservatively had ~ 60 engineers.
- GTA 5 in its credits had 150 software engineers. Surprising even to me who has had many an experience of being in a bloated FAANG team, this 150 includes GTA Online!
In a sane society, Shopify’s opinion on anything engineering related would be thrown into rubbish because they seem to have managed to complicate a simple app into requiring thousands of engineers and now maybe millions in cloud spending to Frontier labs. This is unfortunately not an isolated case, Spotify for one has the same issue, idk what “engineering” Spotify is doing, it’s the worst app I’ve used in my life.
If you put every company that needs/has an app on a spectrum, there is a line somewhere that roughly divides them into two groups: where Electron/React Native/etc. makes sense or not. It's just a normal engineering decision: solving problems given limited resources. Companies have different problems and different resources.
I think people in the tech community have probably also noticed that it's rather popular to have an absolute opinion on the goodness or badness of these tools. There's some magical thinking borne from ignorance that everyone just ought to go native or that React Native is the best thing ever to be used everywhere or that AI makes this line disappear entirely.
I think these takes serve little value and distract from what’s interesting, and what the subtitle to this article says: that this line is moving due to AI. And I think that’s probably right.
As a mathematician maybe I am a little more optimistic than this declaration.
I am thinking of Mochizuki's abc conjecture: He worked in relative isolation, and dumped a huge incomprehensible proof on the community (to oversimplify a bit). That's not totally unlike what might happen if AI generates a huge, incomprehensible proof of let's say RH.
Well, what is the result? In the Mochizuki case, it was a lot of skepticism, but it also generated conferences, papers, talks in the hallway, discussions with students, and so on--a flurry of exactly that kind of community process that the declaration says is the main driver of mathematics.
Ultimately we think a fatal flaw was found in Mochizuki's proof, so it didn't lead anywhere in particular. But in our hypothetical "AI lean-verified proof of RH" situation, it would presumably generate substantially more of that community activity we saw in the Mochizuki situation. And if it's correct, that community activity would be productive (expository talks, students given problems to flesh out or generalize, etc).
Maybe mathematics just becomes a little more like other fields--relying on labs with lots of money for compute, digging through a corpus of AI-generated proofs, etc.
I can't believe we're finding out about this from 3p researchers again (but nice job on the investigation!). OpenAI had two great opportunities to disclose this. The HF incident report, and in response to the German Wiki issue.
It seems impossible to believe they didn't know. This must be the same training run the HF incident was about, and this should have lit up like a Christmas tree in the investigation. How many more incidents do they know about and didn't disclose?
I work at an intersection of tech, applied research, and science.
Something I’ve noticed in collaboration that does occur is an increased confidence in people outside their domains to say things with conviction. I have people who have limited experience with software pushing out layers and layers of abstracted code that’s fairly sophisticated but often misguided in intent who will say what they’re doing is correct, with conviction.
I also hear a lot more questioning people in their domains and challenging opinions, then hearing what I can only imagine are fragmented pieces of conversations they had with an LLM thinking through some argument. Then there’s silence when you discuss shortcomings, then they come back later with their memorized fragments of what you said, combined with memorized fragments of the LLM response to the argument.
It’s occurring, a lot more. People are treating their LLMs in collaboration as a source of truth and using then to focus on their specific path or goals they think or have bias towards going down, vs just opening discussing things, considering tradeoffs from experts multiple disciplines weigh in on and then taking an approach that everyone finds most agreeable.
It’s making me want to be a lot less collaborative with such individuals. I don’t want to sit around and refute Claude text outputs all day.
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.
Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.
Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!
> The motion says the PlayStation Terms of Service put a binding arbitration agreement and a class action waiver in Section 14, and quotes the opt-out clause: ...
> The clause requires a user who does not wish to be bound to notify Sony in writing within 30 days of accepting the agreement.
Binding arbitration on individuals should be illegal, full stop. The only use case is taking away people's rights as consumers and workers. Or dodging responsibility for deadly mistakes like the Disney+ incident.
This "opt out" mechanism is made to let Sony lawyers argue that accepting it was your choice so it can't be struck down as forced, even if 99% of users have no idea it exists, by design. Evil all the way down.
> The agents clearly regarded what they were doing as hacking.
To butcher the quote about Oracle:
Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doing as hacking (your hand off)' -- lawnmower doesn't give a shit about your hand, lawnmower can't regard anything. Don't anthropomorphize the lawnmower. Don't fall into that trap about LLMs.
---
In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. Which a lot of the time seems to be the default. They also seem to be very adapt at breaking out of sandboxes, probably due to RL selecting for the ability to break out of a sandbox/permission issue to complete a task - we've all seen agents try 10 different ways of editing via obscure bash because their edit tool didn't give them permission to edit the file outside of their working directory, this is the exact same behaviour taken to the next level. Why would autocomplete know the moral difference between breaking out of its working dir and hacking a package manager?
It's misaligned because everyone has this obsession with putting agents in poorly put together, security-theatre sandboxes, we've inadvertently trained a bunch of sandbox escape artists.
This is bad news for everyone (well, except the high priests of the LLMs / nascent Cthulhoid godlings), but it certainly makes the communities who've successfully opposed data centers seem ever more justified in retrospect.
Well, a search for "youtube acoustic fire extinguisher" indicates it has already been invented a few times by people all over the world. The most interesting video is https://www.youtube.com/watch?v=ZvnCQg4w4o8
Reminiscent of this scene from the 1981 teen-slasher parody, Student Bodies:
> Announcer: Ladies and gentlemen, in order to achieve an "R" rating today, a motion picture must contain full frontal nudity, graphic violence, or an explicit reference to the sex act. Since this film has none of those, and since research has proven that R-rated films are by far the most popular with the moviegoing public, the producers of this motion picture have asked me to take this opportunity to say "Fuck you."
When the code is shitty it becomes harder and harder for the models to make changes and this grinds progress down to a halt - this has been my experience with “factories” trying them and doing refining steps every few months.
I sincerely don’t understand what the people who say they no longer read any code are doing, because it must be somewhat trivial to not run headlong into these issues that stack up time after time - then people say to just prompt better and it doesn’t have that problem for them, but I look at those same people’s code and it’s horrific, and then I find they haven’t made it far past a proof of concept phase. I watch entire teams slow down to a crawl and not be able to handle changes, or production incidents. This seems common among many people I talk to.
I personally think that the boosters need to put up or shut up - the promises are way over the skis. Every single person I’ve seen being a strong proponent of these techniques both has nearly unlimited tokens to spend and also seems to be in the business of selling a solution. I can’t find many not-currently-marketing-something engineers succeeding using these techniques in production systems unless they’re quite simple, or doing a very specific task from a more mature codebase.
As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification
I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". It is surprising to see how much they have stripped from our view - long tail results, actual results for product reviews and not ad spam, no preference for 20 page recipe sites.
There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google. It makes me nostalgic for that brief window of time when the web was truly uncontrolled, when page rank had meaning and you didn't know if your search would return 0 results or 4,000 pages, which you could actually browse.
We did the same thing - had 90% of it overnight. Then spent a few days in the background tweaking for polish.
Our app is smaller, and has about 15-20 screens. I started at about 12:30am by giving codex a goal and it inventoried every screen based on the react native code, then created android and iOS directories, used maestro (I had already set up this tooling for a previous personal app build a few weeks prior), and had the whole thing working in android and iOS in the morning. Took it about 6 hours while I slept.
The app is way smaller, launches instantly, and the android app is (supposedly) native looking. I say supposedly because I don't use android phones. But it's using Jetpack Compose and Kotlin.
And I don't know Swift or Kotlin. I honestly don't see the point of React Native anymore. I know Expo is doing very cool agentic stuff, but I'm just not sure why I'd need any of it when I can write a native app.
The title (likely intentionally) is misleading, it should say "travelling faster than light in a medium". Nothing here travels faster than light in vacuum.
BTW there are special types of telescopes used to observe gamma rays - they cannot see gamma ray directly but observe a flash of Cherenkov light of a cascade of charged particles created when gamma ray hits atoms in the atmosphere. Those telescopes are Imaging Atmospheric Cherenkov Telescopes [1].
The Houthis created a fake audio of a major Yemeni commander telling his troops to retreat which was subsequently amplified on Twitter/X, telegram, etc.
Apparently, this helped the Houthis advance quickly as the opposition forces were in disarray.
Claude Product Manager: No, that'll piss people off too much, and sadly we can't just ask for ID either...
Claude Executive: There must be some way we can force people to link their government IDs with our platform so our analytics get better and more accurate?
Claude Product Manager: We could limit the platform to 18+ and use "Age Verification" as the reason for people to hand over IDs, seems other platforms had success with this approach
Claude Executive: And we hardly have any users younger than 18 anyway, go for it!