Hacker Newsnew | past | comments | ask | show | jobs | submit | csomar's commentslogin

How are they going to generate revenue if their old model is Ads on search and the new model is yet to show how revenue can be generated. ChatGPT (on free) is showing ads now but they are not targeted and I highly doubt they are worth anything.

Alphabet just reported another blowout quarter for Google Cloud — revenue jumped 82% year-over-year to $24.8 billion

So after chasing AWS for two decades, they are still dwarved by them.

Embedded evaluator will be a great side/consulting-gig for Karpathy and the likes. How do you think these people will be picked?

Military action against whom? the US can’t even handle Iran let alone China.

Americans live in their own plane of reality, where they play the superhero surrounded by evil villains.

America doesn't want to pay the price of handling Iran. There's a difference. Clearly the US could conquer Iran but it requires a land war with mass casualties, not just sporadic bombing.

Obviously China is in an entirely different league.


> Clearly the US could conquer Iran

Vietnam (and Ukraine) would like a word


Korea & Afghanistan too.

The consequences of the Iraq blunder are also still resonating.


> America doesn't want to pay the price of handling Iran.

And Iran has wisely not provoked the US public by launching terrorist attacks against the “homeland”.


"I can afford ferrari. I just don't want to pay the price."

The security implications of this are very severe when you consider the amount of people who google government websites, banking, crypto and others. And google will happily serve you a phishing website either in ads or results.

Can you elaborate?

When you search Chase Bank and get a link that says Chase Bank you can't know if it's Chase Bank until you click on it

Sure. And why does that matter?

Doorway effect, once you click on it you forget you're not sure it's the real Chase Bank.

You can see the URL address bar.

What do any of these have to do with you handing over you ID to some random web service?

I've got downvoted and flagged for saying similar. It's actually positive you are the top comment. There has been aggressive brigading around reddit/hackernews and also traditional media. These are malicious companies, so not out of their line. So far, AI seems to be useful only for programming. It's not clear it's useful for other professions. It can't even write straight without being recognizable from afar.

If AI is not even useful for programming, its value drops significantly. And some people seem to have dropped hundreds of billions on this.


Shopify supports multiple countries both as a buyer/seller. I assume there is all kind of customizations needed to comply with each country weird rules. Are all of these software engineers doing core engineering? Probably not. But they are probably tagged as software engineers within the organization.

> there is evidence bioengineering is already happening

Mind sharing this evidence with us?



No, it's not. And it's the kind of evidence that doesn't help and rather confuse. It's not even clear from the article whether an LLM or a dedicated model were used for this purpose.

Also

> That is, I'm not sure that anyone needs to deploy a new compound in order to wreak havoc - they can save themselves a lot of trouble by just making Sarin or VX, God help us.

We already have toxic nerve agents that are largely available for state actors and possibly available for individuals. If you have decided, as a human, to make great harm, you can already do that.


Are the models improving? Because I am not seeing it. I have been trying Astra for a few quantifiable tasks in my codebase and performance wise, it's pretty similar to sol 5.6. Now when it comes to expressing the problem/solution, holy Christ, what a mess the writing has become. It is on the level of Opus 5. Now when it comes to burning money, Astra is just insane. With a $100/month subscription, you can easily burn through your weekly "allowance" in a morning.

Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?


Yes?

If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.

And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.


Math problems are highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance. There's a lot of quality material on which to train and it's easy to tell quality apart from crap. The search spaces are a priori much smaller than in other areas and the people using the tools to study them are themselves good mathematicians.

Success in such problems does not automatically extrapolate to other contexts.


Finding a training algorithm that can do recurrent networks and continual learning is also a "highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance"

That's the thing I'm most worried about - LLMs that are super clever at coding and maths, making an actually very very dangerous model that is far more efficient, and clever in a more innate (less brute force) way.


>compared to problems in engineering or finance

Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.


I think it’s more likely that that’s because no one tried to solve such problems with them before (OpenAI apparently started working in Navier-Stokes after a rumour that someone seriously advanced the problem with AI) plus improvements in orchestration. Fair, the latter could be as dangerous as stronger models.

Seems like hundreds or thousands of agents are needed to come up with real breakthroughs. Both with the Navier-Stokes project and in the Hugging Face “project” there were lots of agents co-operating on the tasks.

I doubt the Hugging Face one would take that many if hacking Hugging Face was the direct goal being optimized.

I agree that it could be done with fewer agents. It would take longer though. Seems to me that these agent farms are good at coordinating and co-working in large projects, with the agents using message boards for communication.

Some people claim Astra is significantly better than anything else and significantly more token-efficient, and others (like you) say it's meh and way more expensive to boot. I really don't know what to think.

Kind of a tangent, but one thing I am curious about is to what degree the Navier-Stokes result announced today was primarily a brute-forced result based on the 'program' previously established by researchers to find counterexamples (blowups), or whether the model actually added significant/novel intellectual value beyond its ability to run at arbitrary parallelism. With 10K agents and a staggering $15M in compute (IIRC), I am feeling like a lot of the former may have been involved, but I don't really understand either the problem or the approach (or, indeed, the solution).

Obviously the potential for parallelism and coordination between so many agents is quite scary by itself, but I think brute force by 10K mediocre AI mathematicians is much less scary than ~one AI mathematician reasoning its way through the problem where all human attempts have failed. It seems fairly obvious that massive parallelism lends itself to brute-force counterexample-finding, and I suspect it isn't a coincidence that most of the touted AI math results have been counterexamples.

It's all still quite scary, but coming full circle: I really don't know what to think.


Try GPT5 and you will feel the difference. Not one from 2 months ago, but one from a year ago. And then you can get the idea of what happened in just 1 year and what you can expect in 1 year.

A lot of people don't realize that when a leveraged position goes against you (especially with a regulated broker), you get liquidated once your equity runs out, but you can still owe the shortfall on top of that. And the broker can come after your assets to collect it. So your real exposure here is 380K, not 98K.

That being said, I think the parent AI is using paper money. Though who knows, this is the brave new world of AI.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: