Hacker Newsnew | past | comments | ask | show | jobs | submit | charcircuit's commentslogin

Because people want to participate in activities that cost money. It's easy to see use cases like an AI agent notifying someone that a band they like is coming to town and asking if they want to secure tickets to the show.

« Nothing to hide » evolved to « nothing to think ».

>The demo uses Claude Haiku rather than newer models like Opus or Fable. This is because I'm assuming that shopping agents of the future won't use newer, more expensive models,

Why are we assuming that shopping agents are going to be using the model that is easiest to fall for prompt injections? Only testing a year old, small model is going to lead to a misleading conclusion.


Author doesn’t understand cost dynamics of the industry - that prices keep falling.

Author also needs to make a point, and can’t do so if they are using a powerful model.


this is the issue most people don't understand, or maybe they're incapable of understanding.

next year, astra tier model will be $2/1M or whatever (point is - cheaper) and whatever model is $10/1M will be insane like astra feels today.

also, in any agentic workflow, you use models of various levels together... not just one model of one level...


> or maybe they're incapable of understanding

incapable of understanding speculation?


Speculation of what? There are models today that perform better than Haiku at 47x lower cost.

Excessive skepticism does no good.


speculation of cost decreases and performance increases in a matching trend year-over-year

comparable capability has gotten cheaper.

gpt 6 Luna can be swapped for 5.6 terra or even sol in some cases, and it is much cheaper.

source: i have a lot of evals, and i have ai find the best model for me for those evals over time so i save $ lol


Also the author is deliberately chooses a model that is 1 year old and was supposed to be underpowered even then. If cost was deciding factor, they could’ve chosen Luna.

  Model                       AA Index   Cost/task
  Claude 4.5 Haiku (thinking)    17       $0.21
  GPT-6 Luna (max)               37       $0.07
  GPT-6 Sol (low)                34       $0.13
  GPT-6 Sol (medium)             40       $0.25



Luna (low) scores higher and costs about 47x less per task and performs better.

And Luna (max) scores 37 while still 30% of the cost of Haiku which scored 17.

Why deliberately choose a model that is known to be poor and also costly?


>yet they want to take away the main feature that lets them be an active collaborator and participator in the process.

It's just redundant when you can just collaborate in the "main" mode. There's never been a point to having a separate plan mode.


There is a point, plan mode has harness level behaviour - it's restricted to readonly tools, and it re-reads its plan into context after compaction

>The output of an LLM is useless without a human mind to comprehend it

This is like saying software is useless if the user doesn't read the source code of it. This what happens >99.99999% a person uses software. People want to be entertained or have their problems solved.


Your analogy doesn't work. A "product" in research pure math is not the same as writing code.

>the economically dominant strategy to not verify them and not double check them

In the big scheme of things is it really that expensive to verify it if a lean proof is generated? The agent itself will likely have already verified such Lean code before calling it "done".


I feel like knowing something is true is useful, but if you don’t understand how and why, you won’t understand the implications

>Anthropic opposing the US government

Is it not possible that this would change the risk profile of depending on them. It makes sense to work with people who support you than oppose you for things which are critical.


- Having principles and sticking to them is not opposing.

- Just because you don't do everything someone wants you to do does not make you a supply chain risk.


My point is did everyone know all of those principals and how committed Anthropic was to them from the beginning or did new information come in and now they need to adjust. Just because something is a term in a contract that doesn't mean it's necessarily strongly held belief that will never change.

So you renegotiate the contract. You don't get to retaliate.

Only it turns out that with sufficient corruption, you do.


You are again confusing contract terms with risk. They are independent.

Who runs the government changes every 2 to 4 years… are you suggesting we completely swap every vendor in the US government to align with whatever political party is in office?

There’s nothing indicating Anthropic was not keeping their end of the contract. If they had, there would have been other recourse.


>are you suggesting we completely swap every vendor

I think American companies should be supporting the government as much as possible. Even when it switches.

>There’s nothing indicating Anthropic was not keeping their end of the contract.

Which is why we are talking about risk and not breach of contract.


Because they have to follow the law.

That there is competitive pressure to sell access to AI that would cause mass harm.

Can you say more here?

You don't think that there is competitive pressure between Anthropic, OpenAI, Google, the various Chinese model companies, and others, to advance the capabilities of their AIs?

Or you don't think that sufficiently advanced AI can cause (mass) harms?


I am saying that if they lost alignment and their new model started injecting cryptolockers, dropping all tables of productions databases into the software or if it stated writing poisonous cooking recipes there would be backlash, lawsuits, and more towards such a lab. Consumers don't want to trust such a dangerous model so they won't buy the tokens and the lab will not want to spend money on lawsuits.

>Or you don't think that sufficiently advanced AI can cause (mass) harms?

Even a feather can cause mass harm if it's used to sign a declaration of war. Something merely being capable of causing mass harm is not an issue and doesn't mean that the existence of feathers are an issue.


But this assumes an evil superintelligence. We don't even need that for terrible things to happen; we just need people doing people things and AI doing AI things at scale, and that scale is rapidly exploding as we scramble to deploy AI in the real world.

Case in point, that hallucination (which, if it was an evil superintelligence, could have been "strategic") that almost led to US boarding a Chinese ship over suspicions of nuclear weapons: https://www.msn.com/en-gb/news/other/us-military-ai-failure-...

It doesn't have SkyNet, it just has to be WOPR. And it doesn't have to be people who are motivated to cause harm, it just has to be people making consequential decisions who are careless, distracted, paranoid or anxious -- which everybody is at some point or another.


>that almost led to US boarding a Chinese ship

And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.

>it just has to be WOPR

Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.


> And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.

Maybe not you personally, but many, many other people have killed many, many pedestrians, and when they exhibit a pattern of bad driving -- or other deviant behavior -- we take them off the streets. That is an example of society being robust.

Except, over here, we're rushing to make AI, which we know has many deviant behaviors and we know caused harms in many circumstances (starting with AI-assisted suicides), even more powerful AND deploy it in more and more real world systems!

As history and, literally, current events show us again and again, society has failure modes that lead to widespread harm and destruction. We've had world wars and then literally had multiple close calls with nuclear war right after. And then we have all that's going on out there. (In related news, the Pentagon threw a hissy fit because Anthropic would not let them use Claude for autonomous killing machines. Do I have to even mention what they've been up to these days?)

Most of these failures are caused by misaligned incentives and socioeconomic forces. The incentives and forces around AI have hints of many brand new failure modes that we can't even foresee because things are moving so fast.

> Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.

"We're already going down this mountain pass at 200 miles an hour, why slow down now?"

I think we should not just slow down frontier AI development, we should also slow down where and how that AI gets deployed.


Are these old texts really going to improve benchmark scores compared to other things the labs could invest into?

I don't know, but I imagine the compute spend is less than the marketing spend to get similar headlines.

>security backed loans

Which currently require interest payments of ~6-8% APR. Meaning that you need to be able to invest that money that is being borrowed back into the economy to hopefully get a return more than that. And if your investment fails you will have to realize a different investment. The interest being paid doesn't get hoarded either and is used to make other investments, pay employees, build products, etc.

The idea that a bunch of people are just hoarding their money and not reinvesting it back into the system is flawed. Taxes actually have the opposite effect to contributing to the system. Taxes are like if someone was to come and start hoarding money under their mattress for himself and not contribute back to society.



It's funny that their case study ignores how much wealth you lose from the interest on the loans. If you pay $140k of interest you can avoid $62k of taxes.

I don't know what point you were trying to get across with your link, so I gave my general thoughts on the article.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: