Astra is such a mixed bag. It makes some amazing reviews and sometimes architecture suggestions that I like. But it’s also lazy and will just make up things.
I work with Sol and Astra only in my daily work, and occasionally I check out Claude Code so I don't get completely out of touch.
I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?
It's definitely worse and getting worse. I'm curious, do they not know this is happening, or not care? I struggle to believe people prefer the way it writes, which is becoming drastically different than its competitors.
People are also being trained/acculturated to LLM speak, so even if they are trained on human output, they could be getting reinforcement for their LLM-tinged crap-speak.
Fable 5.1 is fairly pleasant to work with, the first in a while. Too bad it's so ridiculously overkill for most tasks. They need to reel in Opus and Sonnet.
I don’t like how LLMs answer and as they are statistical machines I have found out that I need to check every text they produce. From the first sight everything seems cool. But it isn’t. :)
Exception is code that is so huge output you can’t read everything. But I have to try ponytail skill for coding that should shorten the output.
> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.
> [...]
> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
> > This is your coat token, to my coat hanger in the opera.
that is fucking insane
how on earth is Anthropic allowing this to happen? AGI? kill us all? I refuse to believe any of that until they can get this basic shit sorted out, I mean seriously, it's becoming a joke at this point
The issue isn't the prompt you give it. The issue is that as context grows, Claude distributes it's "attention weight" over all that context, and quickly reaches the point of "missing" stuff.
You can give it any prompt in the world, but Claude's ability to remember that instruction quickly degrades the more you use it.
Yeah, I'm well aware. I shpild've been clearer I was suggesting a 1-line entry like that in CLAUDE.json, which can pay dividends in keeping context lean - which in turn is essential for avoiding the "dumb zone" (over 125k-150k tokens), precisely the effect you describe. Experts like mattpocock suggest keeping context as lean as possible for this reason.
I tried it for the first time today for stuff I'm well versed in.
It said something along the lines of "remember $USER five runs of Chrome is not enough for high quality benchmarks! What you're doing is called a _trial run_".
To have a useful continuation to the conversation I had to remind it that I was the one that wrote the documentation it was quoting back at me.
It's not actually an intelligent being so I didn't get angry at it but it was a piss poor experience.
5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.
But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.
The critical IP that isn’t in the weights yet, is often a few hundred to a few thousand lines of code. All the rest, surrounding code, is infrastructure that is indeed in the weights.
We're in interesting times. Trust in AI CEOs, just as an example, is likely about zero and so, unsurprisingly, people are increasingly skeptical of anything they say, question their motives, the timing of their announcements… Huge amounts of capital (perhaps entire economies?) are hanging in the balance.
(Even when I see a number of comments I disagree with, I get at least a sense somewhat of the HN Zeitgeist, FWIW.)
I‘m using Sol all day every day. It’s on average doing better than Astra at what I need it to do. Astra is like an absent minded professor - gives good direct responses, but too inconsistent and forgetful for my codebase.
Let me ask you, have you noticed that all the major labs are releasing models very similar in gains and performance? How do you explain that other than juicing the max out of the current architecture and processes.
They have been releasing similar gains/performance for years now... Pretty much never has one company been dominant for a long time (except the initial GPT3/4 release I suppose. That period took a while for others to catch up)
Yeah - that's why I think they are basically squeezing the scaling laws and the current architecture with incremental innovations.
I expect they will continue optimizing and improving for the current use cases/benchmarks. But the core capabilities will stagnant, that's why they will be forced to slow down until another major breakthrough happens.
I personally don't think LLMs have any understanding of the world, nor any imagination or even agency. I think it got very good as recognizing patterns and shapes in human thinking, language and knowledge, but that's pretty much it.
The one where it’s infeasible that a company that needs to live up to a valuation of like 5% of US GDP but is still burning cash and has no moat would willingly decide to slow down the pace of development of their core technology
There is no need for a copy. If another can do most of what OpenAI/Anthropic can do at a fraction of the cost (like the Chinese models) then the moat evaporates.
the capability gap is underestimated. the raw unaligned base model from a pretrain is like the telemetry recorded from a particle accelerator - it's hugely valuable and not just because of the cost sunk building a collider.
the things are kept under air-gapped national weapons grade security measures not because it's literally going to escape and threaten the world but because if somebody walked out the door with a copy they would have everything. the capabilities we see at the surface are mostly the result of mining the great unknown space and attaching feeble control surfaces, ablating/lobotomizing dangerous areas, and fencing off illegal/secret/embarrassing ones.
Based off what? There’s pretty precise measurements of the capability gap where open weight models like Kimi K3 score higher than the latest flagship models just 6 months ago.
The capability gap is in the minds of the users but the frontiers' closed business models run opposite to it.
> the things are kept under air-gapped national weapons grade security measures
That's an enhancement of closed but not of moat
> because if somebody walked out the door with a copy they would have everything.
Including you and me? I'm not sure this is comforting news, despite the "trust me bro" asurances from behind the door where we can't see or verify anything.
I don't know about you but I would love to get my hands on a raw frontier base model - and a 16 node galaxy blackhole supercluster to talk to it. the capabilities currently being loboptimized for are just a narrow market-shaped slice that cheaper models can distill a competitive subset of but the untapped power of the base is a true technical moat.
about the benchmarks: the difference between a distilled competitive subset and the untapped base might be the difference that matters in any given task
FWIW, I meant "they would have everything... including you and me". It seems I wasn't clear enough and it sounded like "you and me... walk out the door with a copy" - obviously the latter isn't realistic, but the former is.
> the capabilities currently being loboptimized for are just a narrow market-shaped slice
Offensive capabilities aren't interesting to me except as a risk I have to consider. It might sound surprising but the intersection of offensive and practical-for-life capabilities is an almost empty set.
I expect they are slowing down in an attempt to keep from having those models improve to the point where they can effectively escape unassisted into the wild.
Unless you’re making 3D models of apartments and oneshotting trivial games, I‘m not convinced Astra writes better code than Sol. I‘m way more often disappointed about Astra forgetting some aspect that Sol just nails.
Can you really do HIIT on a treadmill? It doesn't sound safe to sprint on a machine.
Re your observations. Are you referring to elderly people? I'm not sure this study is about them. My mental model of this mechanism is that relatively young people can choose between e.g. long distance running and sprinting, and that the latter will yield better cardiac outcomes.
The kind of HIIT I'm doing is only 10-20s bursts, to stay within ATP-PCr energy metabolism (little lactic acid production). A short sprint followed by 2 minutes rest. That's hard to achieve with a treadmill.
I like sprinting a couple times up our hill, ca 10s each, 4-6 times. It’s full out, and when I‘m untrained (like now), I‘m quite exhausted afterwards.
But the dopamine boost I get from this is immense. It lasts 24-48h. I‘m in a general good mood for this period. Sometimes, a little more competitive than normal - that’s a downside.
Normal jogging or long walks don’t have this effect, at all.
A long race will do it just fine. AKA Runner's High.
I also do hill sprints. I try to do one [HIIT] workout a week and also some longer runs.
Hill sprints tend to be safer then sprinting on flat surfaces. Last time I did my HIIT sprinting on level ground my knee didn't really like that at all. A stationary bike (like the "assault bike") is even safer. Or you can just do burpies. People who don't run and haven't sprinted in a long time shouldn't just jump into sprinting.
I have to disagree on both points. Long races don't do this for me - that's also the point of the paper. And I have had more close calls of spraining an ankle when running uphill. The problem with hills also is that you might want to jog down after, but that's super bad for knees.
reply