My thought is more like, if OpenAI can't control or even monitor their model in a test of its breakout potential, what about the future of mid-budget companies which will just be deploying agents left and right with vague instructions.
All instructions are vague unless its code. But you can also give llm "code" and expect vague outcomes if you ask it to emulate what the runtime would look like.
Maybe they are using traditional weapons for a completely different reason, but I'm reminded of why India and China have their soldiers on the disputed border carry poles and clubs. It looks old school, but it prevents clashes from turning into a larger shooting war where everyone's getting killed.
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
Let's see what happens. In contrast to this whirlwind of math that's going on right now, the millennium prize rules require publishing in a reputable journal and 2 years of waiting time to establish that the proof has been accepted by the community. So nothing happens in the short term.
Since no one mentioned it - this seems to be a major and real problem in the crypto job space. In their job market it's more believable that a 'stealth startup' is reaching out and doing a code challenge from an unfamiliar email or repo, and crypto devs are likely to have a wallet or passwords accessible on their system. They are willing to go above and beyond the regular spam or AI conversations to get access.
No you don't, you just have to believe there are people more gullible than you. The creators of those monkeys or the people who hacked whatever wallet aren't gullible.
I'll add the category of people with the arrogance to believe that, after 900 previous ventures had all their coin stolen, they'll be the ones to write a secure exchange that can't be hacked.
I have no inside info on this, but I believe that Anthropic's trouble with the feds really started when they delivered a presentation on EA and Constitutional AI and someone's takeaway was "this is very strange and probably 'woke'"
I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
From what I've seen in the bounty-related subreddits, AI is flooding bug bounty inboxes with low-value or meaningless reports, or straight-up hallucinations when people use smaller models (to turn a profit, you make lots of low-value bug reports and see who pays out).
This has a negative effect on humans doing their work with or without LLMs: curl shut down their bounty program, and GitHub just announced they're "restructuring" theirs.
The author of this post also makes a case that HackerOne hasn't been honest about LLM training and use, either to hackers or to their own staff.
Didn't Daniel later report that curl recently started getting mostly high-quality LLM reports on their bounty program? I can imagine that there would definitely be a few "bounty spammers" trying to get hits, but it seems like most of them are doing good work.
I'd say instead that the problem is that a lot of people don't care anymore about the quality of the work being done, and LLMs are accelerating it. Bounty programs have shifted from ways for people to report security bugs to ways for people to try to make money.
An LLM finds a dubious bug, an LLM turns it into a convincing report, and now the proposed solution is to have an LLM triage it? There are a lot of turtles holding up this approach and the circular logic seems hard to miss.
Automated triage can filter obvious spam, which was already fast and easy for humans to do. The hard part is independently reproducing a plausible finding and assessing its actual impact. If LLMs could already do that reliably, then the slop report problem wouldn't exist in the first place.
I tried to build it (kind of). My team is going to use it to alert us to the most critical issues so we can hop on them before waiting for triage. It's decent at figuring out criticality but it's TERRIBLE at actually doing triage and assessing whether the report is plausibly or implausibly true. Security can be really nuanced, and from my experience so far with the model I'm using it's really bad at being skeptical enough to actually figure out if something is a legit issue with impact or not. I agree though, it would be awesome if we could get AI triage that worked.
There's also a scam where they contact the owner from abroad, offering to return it, claiming they're an innocent trying to activate a used phone they spent their last dollar on, or threatening the original user into unlocking it. They will say that they're hackers, connected to criminals in your home city, using your name and recent photos to say that they're going to expose your text messages. Anything to get an unlock.
My apartment building is replacing the elevators, and someone on staff revealed that the first renovated elevator is intended to pick up only from the lobby (i.e. if you summon an elevator from the 10th floor, one of the others will come). This has caused some grumbling or calling it an "express elevator". In an apartment it really is mostly trips to many floors from the lobby. The only time this maybe would be inconvenient might be in the morning commute, when fewer people are re-entering.
You should observe if it is the same elevator all the time and the building people just don’t understand it. Our building elevators (bank of 2) always seems to try to keep one in the lobby. About 75% of the time you get an elevator right away when you call it from the lobby. But it definitely isn’t the same elevator! Allocating an exclusive single elevator to that function would seem to be strictly worse from many angles. Like if you modeled having an algorithm like “elevator 2 only services calls from the lobby” you’d find it would be very inefficient and not every effective. But I can see having one in the bank always returning to lobby right away making sense.
My guess is whoever hinted at that didn’t fully understand what they were taking about.
reply