Hacker Newsnew | past | comments | ask | show | jobs | submit | manmal's commentslogin

Astra is such a mixed bag. It makes some amazing reviews and sometimes architecture suggestions that I like. But it’s also lazy and will just make up things.

I do want the messy in-betweens. The earlier you catch problems, the cheaper they are to fix. I hope this will become the default of collaborating.

Have the really good copywriters left the profession already?

I work with Sol and Astra only in my daily work, and occasionally I check out Claude Code so I don't get completely out of touch.

I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?


Claude's use of language is hideous at this point. It is verging on gibberish wrapped in important-sounding prose.

It's definitely worse and getting worse. I'm curious, do they not know this is happening, or not care? I struggle to believe people prefer the way it writes, which is becoming drastically different than its competitors.

I wonder if they’re training heavily on Claude generated content or conversation transcripts

People are also being trained/acculturated to LLM speak, so even if they are trained on human output, they could be getting reinforcement for their LLM-tinged crap-speak.

Fable 5.1 was explicitely supposed to improve that, I used it only a bit so far and it seems at least better.

Fable 5.1 is fairly pleasant to work with, the first in a while. Too bad it's so ridiculously overkill for most tasks. They need to reel in Opus and Sonnet.

I don’t like how LLMs answer and as they are statistical machines I have found out that I need to check every text they produce. From the first sight everything seems cool. But it isn’t. :)

Exception is code that is so huge output you can’t read everything. But I have to try ponytail skill for coding that should shorten the output.


Big empty words are probably cheaper to produce than concise, information rich text of the same length.

Maybe the humans are suffering from model collapse, and don't notice it.

> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.

> [...]

> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”

WSJ interview of Amanda Askell: https://archive.is/rDes9


Agreed. The language consistently triggers visceral negative reactions from me at this point.

Yes, definitely load bearing.

Today, I asked Opus what it’s gibberish actually means. It started with (topic about signing implementation):

> This is your coat token, to my coat hanger in the opera.

On one hand - maybe yes??! On the other who the hell speaks like that and it’s so specific…

I’d expect Alice and Bob with locks or house keys. Is this infamous old book scanning (and destroying) affecting latest models?


> On one hand - maybe yes??!

No.


Sounds like one of those "Is to as Is to" analogy questions from the SAT.

Are you sure you didn't typo signing as singing somewhere?

> > This is your coat token, to my coat hanger in the opera.

that is fucking insane

how on earth is Anthropic allowing this to happen? AGI? kill us all? I refuse to believe any of that until they can get this basic shit sorted out, I mean seriously, it's becoming a joke at this point


> it’s gibberish

ironic


load-bearing!

"Avoid use of "mannered" speech in your responses."

goes a long way for Opus and Fable.


The issue isn't the prompt you give it. The issue is that as context grows, Claude distributes it's "attention weight" over all that context, and quickly reaches the point of "missing" stuff.

You can give it any prompt in the world, but Claude's ability to remember that instruction quickly degrades the more you use it.


Yeah, I'm well aware. I shpild've been clearer I was suggesting a 1-line entry like that in CLAUDE.json, which can pay dividends in keeping context lean - which in turn is essential for avoiding the "dumb zone" (over 125k-150k tokens), precisely the effect you describe. Experts like mattpocock suggest keeping context as lean as possible for this reason.

I tried it for the first time today for stuff I'm well versed in.

It said something along the lines of "remember $USER five runs of Chrome is not enough for high quality benchmarks! What you're doing is called a _trial run_".

To have a useful continuation to the conversation I had to remind it that I was the one that wrote the documentation it was quoting back at me.

It's not actually an intelligent being so I didn't get angry at it but it was a piss poor experience.


You have to use Fable or Opus 4.6 for tolerable language output.

Using Opus 4.7-5 is harmful to your health.

https://x.com/wolframs91/status/2090159644849353058?s=46


5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.

But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.


Opus is the worst at it. Fable 5 still does it a little. Fable 5.1 is much improved, at least.

It's bad enough I stopped paying.

GPT doesn’t do all of that all that much when it itself is the implementer. RL has made implementation and reviewing two different behavior sets.

The thing that changed with 5G is that my iPhone 13 Pro‘s battery has been draining 30% faster. Enable the hotspot and it gets burning hot.

The critical IP that isn’t in the weights yet, is often a few hundred to a few thousand lines of code. All the rest, surrounding code, is infrastructure that is indeed in the weights.

It is a convenient time, just as coding capability seems to reach the upper end of the S curve.

How? Coding capabilities keep improving, there is no break

Even benchmark improvements over Sol are not that great (and they certainly tried). If we exclude the weird ARC-AGI situation.

> How? Coding capabilities keep improving, there is no break

So? The upper end of the S-curve is also a line that still rises.


is there some metric that supports this?

feels and hype

HN comments are just on a different planet lately lol

We're in interesting times. Trust in AI CEOs, just as an example, is likely about zero and so, unsurprisingly, people are increasingly skeptical of anything they say, question their motives, the timing of their announcements… Huge amounts of capital (perhaps entire economies?) are hanging in the balance.

(Even when I see a number of comments I disagree with, I get at least a sense somewhat of the HN Zeitgeist, FWIW.)


I‘m using Sol all day every day. It’s on average doing better than Astra at what I need it to do. Astra is like an absent minded professor - gives good direct responses, but too inconsistent and forgetful for my codebase.

? What world are you living on.

Let me ask you, have you noticed that all the major labs are releasing models very similar in gains and performance? How do you explain that other than juicing the max out of the current architecture and processes.

They have been releasing similar gains/performance for years now... Pretty much never has one company been dominant for a long time (except the initial GPT3/4 release I suppose. That period took a while for others to catch up)

Yeah - that's why I think they are basically squeezing the scaling laws and the current architecture with incremental innovations.

I expect they will continue optimizing and improving for the current use cases/benchmarks. But the core capabilities will stagnant, that's why they will be forced to slow down until another major breakthrough happens.

I personally don't think LLMs have any understanding of the world, nor any imagination or even agency. I think it got very good as recognizing patterns and shapes in human thinking, language and knowledge, but that's pretty much it.


? "everyone is advancing but tied" is completely different than arguing we're at the top of an S curve

I think it's narrowing at the top.

The one where it’s infeasible that a company that needs to live up to a valuation of like 5% of US GDP but is still burning cash and has no moat would willingly decide to slow down the pace of development of their core technology

> has no moat

I would love to see someone try to escape with a copy of openai's raw unaligned base model


There is no need for a copy. If another can do most of what OpenAI/Anthropic can do at a fraction of the cost (like the Chinese models) then the moat evaporates.

the capability gap is underestimated. the raw unaligned base model from a pretrain is like the telemetry recorded from a particle accelerator - it's hugely valuable and not just because of the cost sunk building a collider.

the things are kept under air-gapped national weapons grade security measures not because it's literally going to escape and threaten the world but because if somebody walked out the door with a copy they would have everything. the capabilities we see at the surface are mostly the result of mining the great unknown space and attaching feeble control surfaces, ablating/lobotomizing dangerous areas, and fencing off illegal/secret/embarrassing ones.


> the capability gap is underestimated.

Based off what? There’s pretty precise measurements of the capability gap where open weight models like Kimi K3 score higher than the latest flagship models just 6 months ago.


> the capability gap is underestimated.

The capability gap is in the minds of the users but the frontiers' closed business models run opposite to it.

> the things are kept under air-gapped national weapons grade security measures

That's an enhancement of closed but not of moat

> because if somebody walked out the door with a copy they would have everything.

Including you and me? I'm not sure this is comforting news, despite the "trust me bro" asurances from behind the door where we can't see or verify anything.


> Including you and me?

I don't know about you but I would love to get my hands on a raw frontier base model - and a 16 node galaxy blackhole supercluster to talk to it. the capabilities currently being loboptimized for are just a narrow market-shaped slice that cheaper models can distill a competitive subset of but the untapped power of the base is a true technical moat.

about the benchmarks: the difference between a distilled competitive subset and the untapped base might be the difference that matters in any given task


FWIW, I meant "they would have everything... including you and me". It seems I wasn't clear enough and it sounded like "you and me... walk out the door with a copy" - obviously the latter isn't realistic, but the former is.

> the capabilities currently being loboptimized for are just a narrow market-shaped slice

Offensive capabilities aren't interesting to me except as a risk I have to consider. It might sound surprising but the intersection of offensive and practical-for-life capabilities is an almost empty set.


I expect they are slowing down in an attempt to keep from having those models improve to the point where they can effectively escape unassisted into the wild.

Unless you’re making 3D models of apartments and oneshotting trivial games, I‘m not convinced Astra writes better code than Sol. I‘m way more often disappointed about Astra forgetting some aspect that Sol just nails.

Can you really do HIIT on a treadmill? It doesn't sound safe to sprint on a machine.

Re your observations. Are you referring to elderly people? I'm not sure this study is about them. My mental model of this mechanism is that relatively young people can choose between e.g. long distance running and sprinting, and that the latter will yield better cardiac outcomes.


In my experience HIIT works very well on a treadmill. Many treadmills even have presets for HIIT.

The kind of HIIT I'm doing is only 10-20s bursts, to stay within ATP-PCr energy metabolism (little lactic acid production). A short sprint followed by 2 minutes rest. That's hard to achieve with a treadmill.

https://en.wikipedia.org/wiki/Bioenergetic_systems#ATP%E2%80...


Yeah, it doesn't actually work for "real" sprinting. Standard treadmills take too long to get up to speed.

But if you do 30-60s work intervals treadmills work just fine.


I guess it depends on your age and current fitness level. When I was young and fit, I could outrun the treadmill for the periods of time I needed.

The treadmill also wasn't able to change speeds fast enough to really do the good Sprint simulations


I like sprinting a couple times up our hill, ca 10s each, 4-6 times. It’s full out, and when I‘m untrained (like now), I‘m quite exhausted afterwards.

But the dopamine boost I get from this is immense. It lasts 24-48h. I‘m in a general good mood for this period. Sometimes, a little more competitive than normal - that’s a downside.

Normal jogging or long walks don’t have this effect, at all.


A long race will do it just fine. AKA Runner's High.

I also do hill sprints. I try to do one [HIIT] workout a week and also some longer runs.

Hill sprints tend to be safer then sprinting on flat surfaces. Last time I did my HIIT sprinting on level ground my knee didn't really like that at all. A stationary bike (like the "assault bike") is even safer. Or you can just do burpies. People who don't run and haven't sprinted in a long time shouldn't just jump into sprinting.


I have to disagree on both points. Long races don't do this for me - that's also the point of the paper. And I have had more close calls of spraining an ankle when running uphill. The problem with hills also is that you might want to jog down after, but that's super bad for knees.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: