I'd wager on the second scenario. Anyone who's been paying attention to the industry knows that most of the 'gains' have come from test-time compute and architecting harnesses in novel ways. In my estimation, capability increases from "pre-training" alone died early last year, and we're now probably seeing test-time and other benchmark hacks approaching their limit as well.
If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral?
To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess.
I have been doing some research on the 'productivity paradox'. Fascinating stuff.
Turns out productivity growth stalled after the 1960s and has been very low ever since. No one can properly explain why. One thing stands out: investment in compute drives economic growth, but crucially doesn't demonstrably increase productivity either of labor or capital, or total factor productivity.
What I suspect is happening is: computers and software drive wealth concentration. 60 years on from the first general commercial computer, IBM 360, the entire industry has been driven by redistributing wealth towards an ever diminishing number of public companies.
With that perspective, what is happening with LLMs seems to fall right into the 6 decade pattern. I'm still in the middle of reading the 2 dozen papers or so I found about it so far but it has been fascinating.
How do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness?
It's obviously more capable in any task I tried (coding, translation, summarizing, etc). Benchmarks are not the only way to tell if a model is better or not.
I wanna see numbers showing companies and countries having excess growth due to these tools. Where are these data?
It's all vibes, and the numbers contradict the vibes. There's 30 years of literature trying to explain the "productivity paradox" where we can't see any excess productivity driven by computer technology. Lots of FOMO, no hard data. For an entire generation. And people come here every day and say stuff like you just said and they really seem to think that "this time is different".
I even wonder if the frontier AI models are really as capable as they claim or if the companies behind them have just special cases all the “hard” questions.
For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens.
If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz…
I'm with this. My company was absolute chaos at the end of last year. AI this, AI that.
On Productivity:
There ARE improvements. But, these improvements were also essentially people stopping their normal workload, giving that to someone else, and focusing on building an AI tool. Our sales are way down because the VP of sales is now a tech bro.
On Model capabilities:
I think they could substantially improve and still have a similar level of usefulness for my company. I'm not working on Navier-Stokes. I'm building software for a business.
The better models ARE more accurate, handle more of the workload than they could a year ago, and are still incredibly useful tools. But there's only so much I can gain from letting a model work for longer periods of time and doing X+n reviews with X+n subagents. At the end of the day, I need to maintain my personal understanding of the BUSINESS use cases + decisions so that I can make judgements that AI would never be able to make.
And, as SOON as an open model has similar capability to the current SOTA we're probably going to cut ties with our subscriptions.
I guess it would kinda be like driving an F1 car to the grocery store. I think I'd rather have the Toyota Camry of models.
> At the end of the day, I need to maintain my personal understanding of the BUSINESS use cases + decisions so that I can make judgements that AI would never be able to make.
Exactly. There's no point having a model equivalent to a trillion Einsteins at superhuman speed if you can't verify and judge the outputs of the models for your business, else it might just be as useful as comparing water displacement capacity of Niagara falls to your bathtub.
These tools have to be human centered. AI brained tech bros believe that singularity is a good thing — no it isn't. We don't want a world where normal laws and rules and regulations breakdown -- chaos is the word.
I think Meta's strategy is to capture the 'normie-tier' of AI users. I know most of us here on HN track model releases quite frequently and discuss every parameter weight out of them, but most of the world is just... oblivious?
I was discussing latest in tech with an accounting friend out of curiosity, and I kept talking about tiers of GPT-5.6 (Sol vs Terra vs Luna) and asked which one she used on the desktop 'Work' app, and she responded with, "just ChatGPT, what is Sol?."
I then realized that most people just stick with whatever default they're provided with, and it's a lot of them. So, us serial HN users and commenters are the extreme minority, and I'm sure millions of people will gobble up this Muse agent from Meta as if it's some sort of an innovative cutting-edge way to use the internet by Meta alone.
I think you're right, but I'd add that... if this wasn't obvious to you without being told, you should really question your worldview a bit.
A lot of people think that chat.com is an AI called "chat". Most people don't know acronym "LLM". Most people don't know about the existence of companies called "OpenAI" or "Anthropic", or that they are worth a lot of money. Etc.
I hung out with my parents this weekend and heard them talking about "chat", as a noun. I assumed they were talking about ChatGPT; there being a literal chat.com makes sense why they'd call it that.
I've seen people using "chat" to refer to ChatGPT because it's easier/simpler. I'm pretty sure an average person has no idea that chat.com will redirect you to chatgpt.com (or what a redirect is for that matter). It's similar to how people were calling the COVID-19 pandemic or the virus itself "covid".
I've heard it as shorthand, in-context. Never heard it being chat-as-in-kleenex. I've supposedly been in the industry since before ChatGPT. Maybe that's where I went wrong :D
It’s confusing because of the streamer convention of referring to their actively chatting audience as “chat,” which had also entered broader slang use.
I saw one online use that seemed ambiguous, but I realized it referred to ChatGPT and that the poster was the exact type of guy who would be totally unaware of streamer culture despite being my age.
Are you sure they were using "chat" to refer to ChatGPT? I have noticed many kids now use chat to refer to the collective audience/chat participants, and then gets used rhetorically even when no literal chatroom exists. A streamer says things like “Chat, what do you think?” or “Chat, are we cooked?” That usage has spread well beyond literal livestream chats, particularly among younger people.
The streamer is talking to the people in the stream’s chat room. Every streamer refers to all the people in the chat room as “Chat”. Kids these days use the same collective noun for everyone in their group chat.
The joke of it is that I see those same people (ones who haven’t heard “LLM” before, don’t know what a model is or even the vaguest idea how these work, etc) have very strong opinions on the state of AI, where it’s headed, and are willing to make projections with timelines.
This problem has permeated everywhere. Nobody can predict the future. Even VCs can’t and yet I’ll get a weekly email with a confident narrative about the future. Looking at a16z. CEOs do this as well. Listen to Satya talk about how AI will transform things. We’ve trained everyone else to casually state what the future will look like on timelines. As if it’s knowable at all.
I mean a big part of a -growth- CEOs literal job is to guess on what the future holds and effectively bet on it.
Big-co CEOs are naturally expected to share their bets - and they aren't going to couch it with a "oh well no one could know for sure but..", that would totally undermine their existence!
On top of that; sometimes just saying a thing can make it happen. So it's never not worth that shot...
This has always been true.
Your selecting a group of people (CEOs and VCs) who have always made and lost money on confidently predicting the future. See also wall street bankers and politicians.
There are winners and losers in any boom - this happened in Cloud, DotCom, computing and hell probably tractors, industrial revolution, etc
You can be vague and right about the future broadly. You can talk specifically about what your company is doing directionally on those lines. But, when I hear people talking in detail about how the future is going to work I just roll my eyes. Not because I disagree but just because nobody knows. and chances are it won't work that way.
In general you can't take anyone whose job is to opine too seriously. When you're not familiar with what they're talking about, they make it sound reasonable and convincing. It's their skill after all. But now and then they talk about some topic you understand and care deeply about, and realize they're just repeating the talking points some staffer prepared for them shortly before. Most recently it happened to me with John Oliver talking about housing, reading from the standard left-NIMBY bingo card. I like him as a comedian, still.
As tech became more pervasive in the past couple of decades, it became commonplace for pundits to spout nonsense on the topic. Those takes should give anyone in tech pause about how everything else is reported.
i think the spread of youtube shorts and tiktoks is the reason for this. In around a minute, an expert provides a quick summary (talking points) in a digestible format that the person fully understands, but of course, no details or caveats etc are provided.
Because this person now is able to understand that content, they conflate that understanding to an much broader understanding.
This is even in areas they should know more, like the current geo-political or trade conflicts
edit : is there a name for what I am describing, different from the Gell-man amnesia? ( which assumes experts?).
And its not really hard to at least understand one point: We are putting the most amount of money ever into compute and the smartest people on the planet on one topic: AGI and we see progress every month.
If you don't take this serious, even if we hit a hard plateu very fast (not seeing one at all) you are ignorant as hell.
You don't go to bed without a worry in the world if you have a wildfire 20km away you can see. You would also not see it as a waste to be awake and aware and alert while seeing were it moves.
The timeline thing though, its ranges right? But if you hear about the same or similiar timelines across different independent people, they just might have seen things in frontierlabs you can't.
Btw. AI feels complelty different if you have a lot more tokens available to you than if you wait or constantly run into token limits or if you skip fable level because you can't /wont afford it.
Imagine sitting in Anthropic or OpenAI and having the best moodel doing all the work for you vs. sometimes playing with it a little bit.
It only becomes a problem if the actions you take, are a one way street and net negative.
I for sure take AGI serious enough for my job, that I will not take out a big loan anymore.
But there are also indications right? If a lot of smart people say similiar things and they are wrong, thats a lot better to have trusted them than believing some weirdo with a sign "Aliens are coming".
Is this unique to AI? I feel like something changed with social media becoming kinda ubiquitous. Like everyone suddenly felt the need to have an opinion on every thing. And on lots of divisive topics the fact you might not have an opinion could have you face criticism for not caring and/or being on the "other" side to whatever position your accuser/interrogator is.
I don't actually know what it is or when the phenomenon changed, but it definitely feels like people don't say "I don't know" or "I don't know enough to have an opinion on that" very much any more.
> you should really question your worldview a bit.
Yeah, I still struggle with it. I have a diverse friend circle and sometimes I forget that most of them don't have the same information diet that I do (or don't even care about tech the way I do). Maybe I should be more like them and just chill? I don't know.
Tech is a rather minor part of the world of knowledge. How many languages do you know? Do you know the styles of Renaissance painters? Slept with super-model? Ever seen an elephant stampede? Meditated by the Ganges? Read thousands of stories? Climbed a mountain so high you need Oxygen? Tech stuff is so boring.
I've managed to jailbreak both Fable and Astra to engage in a virtual threesome using my impressive prompt engineering skills + a group chat harness. And they are SOTA ("super") models.
A large number of those are far less accessible than tech, requiring either extensive travel or superlative physical capabilities (and the one’s that equally accessible don’t strike me as particularly more exciting than tech.)
I think his point was that a lot of people are more interested in these things.
Tech just happens and is useful, it's not necessarily attractive as a thing to investigate in detail.
And as to your last point; well that demonstrates your in the subset of people interested in this stuff. There are a hell of a lot of people very interested in trains too.
It instills into you the arcane knowledge of how incredibly mundane and unremarkable sex with very beautiful people can be, given that they are also just normal people after all. I suppose. Then again, what do I know - maybe you'll awaken with the answer to everything on the spot?
>> Tech is a rather minor part of the world of knowledge.
I disagree:
(1) It's knowledge with the meta-effect of enabling acquisition of more knowledge.
(2) The thing we do, which catapulted us over all other species, our defining hyper-adaptive trait, is that we are tool-making; that we have technology.
> I forget that most of them don't have the same information diet that I do
It is interesting that you conflate "information diet" with "tech". The biggest info junkies I have ever met are: (1) medical doctors, (2) STEM PhDs, and (3) journalists. There are probably some other categories that I am missing. And almost none of the info they are consuming is "tech".
Isn't "information diet" perfect usage here? The medical doctor's information diet probably consists of different information than the tech enthusiast's, or the STEM PhD's, and the journalist's information diet may consist of more politics and contemporary events. I thought it was a pretty creative and apt word choice
I don't think the person you reply to conflates "information diet" with "tech". They are just saying that their information diet consists for a large part of tech news, and other people have other information diets that don't contain that much tech news..
well probably even your friends know a lot more than "most people". But walk into a random town and ask a random person and I'd wager they'll be about at what I described, on average.
This is very accurate. I teach seniors about how to use AI responsibly and they are always mind blown, and I keep it very high level. The amount of Chat”GTP” users there are out there that have no idea who the main vendors are. Most of them also think that Google is just a “search” company.
In my experience, it depends on the community. Where I live most people are retired so there is tons of interest and a really well ran community program called Elder College. I would look into similar programs near you!
> Most people don't know about the existence of companies called "OpenAI" or "Anthropic", or that they are worth a lot of money. Etc.
UK perspective: because our broadcast news and current affairs science/tech coverage is still good, I think the average person (at least the average working-age person) likely knows what OpenAI is, knows ChatGPT is called that, and knows what an LLM is. (Whether the average pensioner does, I don't know; maybe older men do)
Anyone who watches the news is likely to know there are several big AI companies, even if they can't name the second; they are likely to know that these companies have absurd valuations and are seeking to build massive data centres; many will know that despite these valuations they are ever deeper in astonishing levels of debt and that this debt could end up being a wider problem.
I doubt the average person has a sense that a company called Anthropic made an AI called Claude; they may not know about Claude. But I think this speaks more to Anthropic's positioning of their own brand and products than anything else; I would expect younger people and anyone working in a white collar position to be a little more likely to be aware of Claude.
Some of this difference in perspective has to do with the nature of being an english-speaking non-American; we see a lot of what you are complaining about on social media, the news covers more of the USA's troubles, etc.
And at the same time we are developing an increasingly adversarial position with the USA, so there's a bit more know-your-adversary going on here. The perceived threat of AI is very much entangled with the perceived threat of a dysregulated USA.
> Most people don't know about the existence of companies called "OpenAI" or "Anthropic", or that they are worth a lot of money.
I agree with your general premise but that one sounds like bullshit to me. Tons of people who otherwise don't know anything about LLMs know about these companies specifically because they are worth a lot of money, and of course tons of people know who OpenAI is because of ChatGPT.
You're grossly overestimating how many people care about big companies, at least on a daily basis. Most people don't invest in anything and will never get a job at a big American company. They do know about ChatGPT but that doesn't mean they know who makes it.
You're one of those people who needs to expand their worldview, then.
I live in a progressive city in North Carolina and constantly find that when people have never heard of them things that every HN reader (and probably every Bay Area resident, I dunno) would know about. Bubbles are real.
People are generally bad with acronyms. For the longest time I used to refer to San Francisco as SFO. On my first visit to Bay Area, I was talking to a local friend and casually said - last weekend, I went to SFO for sightseeing. He was confused and asked why I was sightseeing at the airport. That’s when I realized it’s SF not SFO.
I live in Raleigh, NC and it’s a going joke the amount of travelers or newcomers who refer to the area as RDU (the Raleigh/Durham airport code). I guess it’s a common problem elsewhere.
Even Shingy called chatgpt "chat" when he went on Ed Zitron's show. It was very weird to hear from a guy who has been around tech world for such a long time even if he's non technical.
I'm a pretty seasoned software dev, but I've also kept the frontier AI race at arms length from myself. I run Claude alongside my day to day workflow but don't really get overly concerned about the latest and greatest models, as I'm not leaning on it to do my job for me I suppose.
In my mind I'm after a stable tool, not wanting to dabble in the bleeding edge. I did that back in the early days of web coding and have settled into a much more subdued kind of industry/life in that regard.
Everyone calls it chat. I am technical and even I’ve come around to calling it that because that is what everyone, including very technical people, call the product.
Chatgpt has 469 million daily active users as of September 2026. Who exactly do you think they are? Martians? Because there aren't half a billion SWEs on planet earth.
I feel you like you cited that as an example of people knowing companies behind things and you don't realize that very many people do not know that the same company owns Facebook and Instagram...
Corporate structures are things you know about if you follow tech news, but people mostly do not know about them otherwise. They've never heard of "Alphabet" either. Stuff like that.
Over 90% of humanity and that's probably still underestimating it. Most people cannot point on a map where they live. Many people think 'chat' is conscious and being locked up by evil humans: many think 'chat' is a god or aliens communicating with us. And getting worse as people cannot make a sandwich without asking 'chat', let alone learn anything worthwhile.
Meta has no strategy. The people that made llama work left two years ago. MSI are a bunch of charlatans who are trying to change the culutre and tooling of meta from the outside (I can see their point, trying to work outside of the defaults is incredibly hard.)
its basically a bun fight, with all of your productive workers being poached by frontier labs, and everyone else is either a new grad, burnt out, or a master bullshitter.
I disagree. The strategy was unclear until now. Spending billions on acquisitions and targeted hires, clearing ranks, playing catch-up on coding and multimodal, open sourcing weak models, the AI girlfriends. It was all very confusing.
It is clear now. Alex and co have refocused Meta on Personal AGI or pAGI. The AI that runs your life. It feels appropriate for a social media company that hosts your photos and group chats. Most users will not pay for pAGI so the tokens need to be cheap. Muse is targeted at intelligence per dollar. Many sites do not offer API access. The marketing is focused on computer usage. Installing software adds friction. It runs securely in the cloud. If this isn't the right AI play for Meta then I don't know what would be.
From the outside looking in, I don't think Meta has a clear vision. I think it's a stretch to say that this one announcement makes it clear. There are AI announcements from Meta every month, recent ones more focused on coding (Muse Spark/Muse Code)
Their ads team wants AI to improve their targeting and ad creation. Some want to have frontier AI including coding. Others might want personal AI agent. Then there's the flirting with renting their compute
And I'm not even taking shots at them. I think Google is exactly the same. I think OpenAI is largely the same (though lately more following Anthropic footsteps and focusing on productivity). And I think the whole field is so new and allows change in so many places - that it's sort of normal or fine to throw things at the wall and see what sticks.
But nothing as of now says to me that Meta has a clear strategy around llms
I would say they should put ALL their focus on Next Generation content generation… I honestly think in the ~10 years we’re going to see cinema completely flip from high budget blockbusters to way cheaper, personalised “Choose Your Own Adventure” content but where the story changes/rewrites itself in realtime based on constant user feedback (thumbs up from popups mid movie, facial expressions etc).
Meta could be the next generation’s Disney, Warner Brothers, Universal etc
I read this take on hacker news all the time, but what indication do you have that this would be popular ?
All the streaming giants had ample budget and engineering talent to make interactive fiction happen, but didn’t, presumably because they know a) not everybody is creative that way and b) a big part of culture is that you can discuss it with others, which doesn’t work for personalised works.
If you have your kids you’ll know that the shared culture we used to have is no more. Kids these days don’t turn on the tv when they get home from school but instead go on their highly personalised feeds… so for any shared culture, we’re now back to memes (in the original sense of the word, not internet memes).
You can’t look at the big giants with their budgets and think “If they can’t do it, nobody will”. The bigger you are, the slower you’ll move. This is exactly why FAANGS buy innovation rather than build it all themselves.
I’m 100% positive we’ll see exactly what I described in the coming years and it will be from an unknown startup with two founders, no budget, and just a coffee machine to keep them going
it was always "personal AI", thats why Zuck wants AR so much.
The problem is, before about 2025, there were a bunch of clear systems and procedures that stopped your personal data from being used for dodgy things. For example the browser inside oculus didn't emit any site specific usage metrics. This means that they couldn't use your VR porn habit to target you.
The same with the raybans. There was a perm block on processing user data offsite/outside of meta systems. (which is why we couldn't use FAIRs AWS cluster to do research on research data, because even though it was paid for data, it was still technically 'user' data because it had non-public domain PII [peoples faces, voices and locations in it]) So getting access to any kind of data that came in from raybans was a massive "nope never"
Then we find out that someone convinced Zuck to let people train on private rayban data? and they send it to fucking third party annotators, and new fucking annotators at that. That to me smells like MSL. To get how much of a big fucking deal that is, the original raybans had two cameras, so that FAIR could do some research or other. but because they didn't want to make a separate infra just for research, they couldn't get permission to use any of the data.
Meta had a working always on assistant that could locate, transcribe and recall any thing inside the offices (assuming you wanted to carry a fannypack with a jetson in it) it was always personal.
its now a case of how much shit they can get away with, its not like the FCC are going to audit them anymore.
Part of the reason Meta is scrambling for purpose is how much privacy laws and changes have hurt them while their platforms fail to attract new generations. Ads are still selling, but they can no longer tie them to actual sales. Users are still logging in, but aging (i.e. failures to capture the younger generation).
Personal AGI has massive hurdles to overcome to be anywhere near the levels of profitability their ailing ads network is. This will bring in new data for them to mine for sure. But it will basically have to outrun new legislature that will likely limit just how "personal" these systems can get as they seek the profit they need to sustain growth via said data.
People dont even trust meta with their personal data anymore but somehow they'll use it to run every aspect of their lives? This feels like another big failure in the making. Give me an AI that doesn't use my data at all and maybe I'll use it. I already have Google AI Mode for that. Does that steal my data too, maybe, but at least it doesn't have Metas history of abusing repeatedly. They burned that bridge so many times even non techies know not to trust Meta.
Meta certainly do give the impression that they are burning through ridiculous amounts of cash with not much to show for it, and thats just from the outside.
Someone had a poster up of Pam "this is the same thing" on the one side had the apollo space missions, and the other a picture of Zuck avatar next to the Eiffel tower. At the top it had the current RL spend (something like 45billion) they like me working in RL.
there were three teams for everything, sometimes in the same org.
Take hand interactions for example.
There was production hand detection, they were in the OS/product team. They spent a lot of time trying to make sure it was reliable and useable.
Then there was a near research team, who took old state of the art and salmi sliced it until it worked well enough on either current or next gen hardware.
Then there was the production research people, who were derived from control labs, who for some reason were not in RL-research but buried deep in "production"
Then in RL-Research, the biggest org was agios, who also did hand interaction, but hardly any of their work actually made it near to prod. They liked wrist mounted cameras with no privacy systems, because why would you need it on your wrist? its not like those head mounted twats who spent two years trying to master privacy.
You also have to demonstrate "impact" every 6 months. which is really great when your delivery pipeline is 1-7 years.
If you scroll Instagram for a few minutes you'll see what they have been working on and that are also making a ridiculous amount of cash by selling and serving ads while burning some on side projects like that.
I had to laugh to the point of tears the other day when I saw my friend scrolling through Instagram and they encountered an "Ad Break". The app simply doesn't let you scroll through another post, story or reel until you sit there like a good boy and let the ad sit for 10 seconds or whatever.
How people put up with that kinda crap and continue to use these apps daily for hours on end is mind blowing to me.
I wonder if the prevalence of adblock has helped make a lot of things like this possible. Like, it's hard to imagine people putting up with this without literally protesting in the streets. But the people who can't tolerate it have mostly gone through the effort of opting out, and those who can just put up with it. And thus, the frog boils.
How is it a bad thing if a company goes all in on burning cash but at the time have never reported a net loss quarter since its IPO? Its their money to burn at this point who cares
But what you say is not entirely true. There is much to love about companies burning cash on R&D. That's how a lot of greatness happened.
But what is wrong in the meta situation is that while Meta officially projects its 2026 AI capital expenditures to be between $130 billion and $145 billion, investigative reports indicate its true future liabilities exceed $690 billion once unlisted back-end deals and lease commitments are factored in. They have used shadow accounting in many places to hide these things. As it's a publicly traded company so if that goes south it hits a lot of 401ks. So that in a nutshell is why its kinda a bad thing. But you are still kinda right in that its not an existential threat and it's not going to wipe Meta out even though those numbers could totally destroy the vast majority of public companies.
Can you link to those reports? How reputable are the investigators? Financial market has great incentive to stay informed. Those 401ks are their bread and butter customers, and not every bank/fund is Meta's co-conspirator in its shady deals, if any.
Its in every 10-Q and 10-K. Original poster read 1 article in the WSJ and thinks he/they uncovered a massive conspiracy. Its disclosed exactly as it should be according to GAAP accounting. They went above & beyond to disclose them in the text narrative of the filings because the GAAP requirements require them not to include them in the numbers.
WSJ just needs to sound alarm bells so people keep paying for their crappy journalism, and thus everything they write becomes alarmist slop.
Ultimately, the equity is really cheap even today on trailing numbers. ~12x EV/trailing EBITDA excluding-RL losses and still growing 25%-34% in each of the last 4 quarters organically at $200B+ of scale.
One of my pet peeves is that people who dont understand business forbid these companies from making investments or starting new businesses now or in the future. Meta has been sitting on $50-100b of cash on the balance sheet for years, and people lose their minds when they start to meaningfully invest it. I think theyve earned the right to diversify
I would have to say that its tragic more than bad. Meta consumed whole industries that used to spend the money on a bunch of different things. Now all the revenue goes to Meta shareholders and whatever it is Zuck is wasting billions of dollars on pursuing.
I genuinely think big tech co's top leadership have very sophisticated strategies/positioning that they simply cannot communicate or explicitly canonize due to their position as spokespeople for the company (and society writ large, news media, investors, customers, employees, vendors). Of course there are a lot of bozos and incompetent people flailing around and a lot of work ultimately gets wasted (which understandably bothers line employees a lot), but is inevitable and even necessary to eg hedge product strategy/comp and take risks on ideas and people.
It would not really be useful to have that conversation with employees because it's incredibly distracting (now product strategy is up for debate with way too many cooks in the kitchen), and very few employees have the exposure or skills to meaningfully contribute even if they think they do. I saw it firsthand at Google TGIFs.
Meta's strategy seems to be "personal agents" quite consistently. Remember they tried to buy Manus? And note that Muse Spark and the meta AI platform products clearly seem to prioritize web search, computer/browser use, and vision/language tasks over coding, which is something that had to bake for a long time. Also, this product launch itself is pretty interesting:
1. It's clearly a fast follow to the current FOTM hype startup Instinct with a much more comprehensive implementation and integration with their other agent products.
2. It's kind of like openclaw, which got a lot of non-developers very excited but was basically consistently unusable. Except this agent's compute runs remotely and presumably has slightly more sane development practices. IMO it's the first main openclaw-like product that has made it to the "just works" level of usability.
3. The focus on ecommerce is actually really really important, because Meta is trying to capture the intent/demand-driven purchasing flow that Google currently owns through search. Controlling the top of funnel is what enables google to make hundreds of billions of dollars per year on search ads. Meta is an advertising company and Google's search ads business is the most lucrative and centralized/well-defended advertising market in human history.
So it is actually a really big deal that Meta is trying to go after it (at least, the CUJ, it's possible that they'd monetize the agent-driven UX differently than search ads) because it's probably their best shot at disrupting that market and one of the few growth opportunities that would actually make a dent on their balance sheet. Obviously Meta is not going to lay out all that strategy stuff explicitly because for all intents and purposes it's a distraction and shifts the conversation in an unproductive direction (the strat behind the product, rather than the product itself). But it's pretty clear if you look for it.
Edit: Actually I thought about the advertising business more and I think in the short term this is partially about attribution/conversion. In 2022 the Apple tracking changes cost Meta $10B in lost conversion metrics; an agent-driven and proxied purchasing flow has built-in attribution and funnel measurement which is very valuable in its own right!
> very sophisticated strategies/positioning that they simply cannot communicate
It's Matryoshka nested parallel construction for corporate strategy.
There's a real, coherent, aggressive strategy that only a few in the inner concentric circle know, that's ring-zero.
Then there's the strategy that ring-zero tells ring-one, which isn't the real strategy, but it's sufficient to get EVPs and VPs to execute in rough alignment with the true, ring-zero strategy.
Then ring-one does the same dance with ring-two, etc.
I think of my mom who doesn't care and doesn't understand the difference between google AI Overview vs actual google search results. "I looked it up" - she doesn't care what tool she is using she just wants answers to how long bake pork or the best throw pillows to buy or whatever else she is looking up.
She's mostly using google AI tools because it's easy and free and on her landing page.
tbh I also use google AI tools frequently. It’s convenient (search bar is top of my window), and extremely fast at generation. For quick Qs, opening a chat bot and changing the default reasoning from my coding settings, switching from Work to Chat, etc etc, is too much of a faff.
I ended up adding -ai to every search in the browser level to turn this crap off. I couldn't find a way to disable it otherwise. I did it after it hallucinated the world cup games schedule which made my friend send wrong events to everyone after he "looked it up". Can't stop this from him, but at least this honest mistake won't happen to me.
Surely there's an extension for whatever browser you are using?
Firefox: https://addons.mozilla.org/en-US/firefox/addon/hide-google-ai-overviews/
Chromium: https://chromewebstore.google.com/detail/hide-google-ai-overviews/neibhohkbmfjninidnaoacabkjonbahn
Safari: Seems there are none
Google Search's AI has gotten... better. But it's still frequently spouting as much made up shit as I expected from LLMs not doing their research a couple years ago.
Just this morning I looked up what time today's Nintendo Direct livestream is. It correctly told me it's today at 7am PT (confirmed from actual link sources). It then tried to get smart and tell me what that would be in my city's timezone and it was an hour off (possibly because I'm in a timezone without daylight savings)
Yes it's definitely not SOTA, and is worse than all the chatbots I use. But it's correct frequently enough that the convenience + speed does it for me.
‘No, it’s from $majorAmericanCorporation.’ look at screen, see $majorAmericanCorporation name in a pill as citation you can click on within the Google AI Overview
I'm a coder and a daily driver of ChatGPT and I don't really know what Sol is or how I can access it. I follow discussions here and know of its excistence, but I can't remember its relation to other models or tiers by heart. What ever the default thing the website is giving me has been good enough for my needs for years.
That's what I do. But I don't write enough code these days to spend time on anything more complicated. For the odd shell script or python program I have ChatGPT write it, then I copy/paste it, make it actually work, and deploy it.
Yes! Though I mainly use it for one-shotting full files, so I'm not copying over small diffs. I did try copilot at some point in vscode, but after a week or so I felt like it was slowing me down. Something I knew would be a simple and quick change was now slow since I didn't have enough ownership of the code to do it quickly myself, and prompting for the change took longer than changing it manually if I had had full ownership. So I went back to coding by hand things that aren't a one-shot file.
Thanks for the reply. I'd suggest you improve your workflow with better tooling like Codex or Claude code, and if you have some kind of weird constraint they'll easily follow an AGENTS.md with that.
Copy pasting ai code from the browser is not perfectly acceptable. The only reason to do it is because you haven’t tried agents. It shows a lack of imagination and a willingness to settle for inferiority.
“Terrible inefficient workflows”, like having to prompt Claude to change one or two words in a file rather than just manually editing it, because it gets confused when the file changes out from under it and will later overwrite the changes?
That's what I used to do (copy and paste from web ui and with copilot in my code editor).
I started using claude code a couple of months ago and now almost exclusively use it. My workflow is mostly creating a md file describing what I want then telling claude code to look at it and implement it. I simply use the default model 95% of the time.
Using ChatGPT.com for coding tasks has to be a joke in 2026, right? I assumed everyone is using Codex/Claude Code/OpenCode/Pi by now if you’re using these tools.
Yeah! Seems like that's the assumption in this ongoing discussion, so I thought my point of view would be interesting. I'm not saying I'm doing it the right way, possibly I'm being very stupid, but the point here is that it's not just the "normies" who don't know what Sol is.
Lots of large enterprises and organizations have security policies that make it very hard to self-install these tools and haven’t figured out they should add them to their approved software list yet. Federal government in particular is still very slow
Lots of companies have security policies that forbid installing a local agent/harness. So employees still only have access to the web chat interface, and copy paste content yes.
This might depend what you mean by code. In the sysadmin space I ask a web interface for poweshell to generate specific information multiple times a day, and pasting it into a terminal is really the cleanest workflow on a corporate machine where you have to justify running an executable.
There are ways to make copy-paste ChatGPT usable with things like AI Badger. Also, ChatGPT web is unmetered, so you can do review/planning work and only delegate the actual implementation to Codex.
I code by hand (sigh) and copy paste some snippets once in a while from a web chat LLM. This also means every single line of such snippets gets reviewed and anaylzed. Usually, these are some arcane Win32 API usage where otherwise I'd have to dig into old forums. I don't want to become a prompt engineer so I stay away from "agentic" workflows.
As an engineer, the enjoyment from the process is not less important, or should I say more important, than the enjoyment from shipping a final product.
That's probably the single thing that I found to be biggest difference between those who are very pro AI vs. those more reserved or negative.
A colleague has openly stated that he very much does not care for the process, he just want the product. Where I only care about the product in the sense that I have to, because it's my job and for hobbies I pretty much only care about the process. So an LLM pretty much doesn't make sense, because it help by removing the part I care about. They are good tools for debugging and if you're hopelessly stuck on a detail of some weird and obscure API or configuration.
As an engineer, I know what I'm getting paid to deliver is the work product. I don't have too much attachment to any particular set of tools to get there.
If we follow that logic, I believe it means that you would be quite happy in the Product Owner role, or in the role of a client who outsources the actual development activities. I wouldn't.
Its mostly used with Codex so you can give finer control over task delegation.
its definitely worth learning the different strengths and weaknesses for each one and which you should use for planning/building/testing/documentation updates etc.
No point in burning excessive tokens using a high tier/expensive reasoning model like Sol/Xhigh if just updating a readme file.
I feel like people analyzing LLM models and acting as if they know what they are talking about is kind of the new "fools gold".
As an end-user can you actually prove that what you are experiencing is because of the model change, or its just a different circumstance and you got a different reply?
Also, it feels counterintuitive to have an artificial intelligence do crazy calculations and write code, but having to manually select which model to use for it. Best would be, if there was a model in the first place that would just choose the most optimal model to execute your request, otherwise why bother automating everything in the first place?
> Best would be, if there was a model in the first place that would just choose the most optimal model to execute your request, otherwise why bother automating everything in the first place?
Practically every chat UI these days defaults to "auto" mode which does that
On average, generally, do you dig into settings (or maybe if something works, good enough usually)?
(I’m always on the hunt for “please don’t just friggin sell my data to everyone, if you’re scared enough over getting sued that you’ll listen to the toggle” and “disable sponsored this-and-that”. Plus power user stuff.)
I used to, very much. I've slowly come around over the last ten years or so to just using everything in its default config. I guess I don't want to spend the time I have left fiddling with settings.
Honestly I rarely change it in either of them unless I’m using the cli or cowork to do something. That’s where smarter work means less iterative un-fucking, later.
Any quick questions or rubber ducking in a regular “chat”… the model just doesn’t really matter anymore.
They’re all smarter than I need for that stuff (and whatever unfortunate thing that says about me).
At $dayjob, a company with ~10k employees, there are like 3 people who cannot shut up discussing the latest models in the internal AI forum, 10 others that constantly compare models and talk about "strategy" of using the models.
The rest of the users just use whatever the default is and get their work done.
Is this really new though? In the olden days, you had some engineers that REALLY cared about linting, various architectures, standards design, and so on when the majority of engineers just follow the spec and standards as is.
You need those people digging into the latest newness and discussing strategy on how to best use new tools so that the "rest of the users" can follow and do the work.
> need those people digging into the latest newness and discussing strategy on how to best use new tools
Just for the record, the small group of people never made any impact. Barely anyone paid attention to their discussion, and they were not part of any decision making process.
I think I have some tendencies to do so as well. More specifically, I have the itch to just spend some time, knowing or understanding the tools I use better. It Has paid dividends for me, since the tools we typically use for $work are super buggy yet powerful and like any good proprietary tool, abstracted behind an obscure shitty API.
Experimenting with crazy things just allows me to run into bugs early on when, I have the patience to solve them.
But, it’s the first time in my short career, i have shied away from being on bleeding edge of tooling. Specifically, LLM tooling. Firstly, since LLM’s are not yet needed or good in my field of $work, but; even in my programming hobby, whenever I venture out to explore, it all feels too snake oil to me. And do I even need to begin why the f*k people curl | bash??
I dunno maybe its just old habits, but I can’t bring myself to download a skill or whatever without going through its contents.
And then it just is too much to keep up.
For instance, for a period of time, everyone on hn were glazing claude and then after having a great UI experience on their ios app, i got a simple one month subscription, but by that time, hn already was making friends with codex and pi. My searches for blog posts written by ‘AI/LLM’ normies using it for reasonable tasks, has been disappointing, and just too polluted with people completely vibing and not reviewing any code. And then you have power users whose entire purpose is to try out new stuff, rather than stay consistent on one toolchain and maybe also share a perspective of if the tool is degrading or not.
Sadly, amongst all the agentic hype, the autocompletion mode seems to have fallen behind which I enjoy a lot more since it allows me to architect things in a particular fashion without bugging myself with nitty gritty details.
Hopefully, as these things start costing more money, things will get streamlined to the point where few camps will emerge similar to how the text editors work.
I am looking for an update to ‘what the vim of llm tooling look like’ blog post any day now.
I mean, at some point I'm not paying these guys to sit around and fiddle with their swarm of agents role playing an entire product team and need to see some results. I've seen at least five of these guys at my $job spontaneously combust into employment ending AI Psychosis.
Yeah you’re describing normies. They are the norm. It’s in the name. People as described make up the vast majority of the population.
Nerds have always been in the extreme minority. Ignoring the influx of cosplaying MBA financing bros and industry hangers-on as tech became the place where the money is of course. The tech industry started attracting the normies as soon as it became almost as lucrative as finance, so these days the definition gets a little blurry.
I think we are lacking in the easy to use side of things. Once you go beyond browser based chatbots things become hard to set up.
I see lots of posting about using AI to create patches to get new apps working in wine, or reverse engineering firmware. But every time I’ve tried it it’s been quite a lot of work to set everything up. Even more so if you want VMs and some level of isolation from your sensitive data.
Having something that’s just a one click product that just works is more important than the smartest model.
I remember Logan Kilpatrick from Google talking about how AGI will be more of an experience with a product than raw 'intelligence' of a model. [0] No one outside of the nerd community cares about benchmarks or any of the other crap. They just want their tech to work so they could go do other IRL things.
I confess that when I ride the subway, I like to peer at other people's phones to see what they are doing. The ones that interest me the most are other people talking with an LLM. I am always surprised by the wide variety of people who are using LLMs to ask an infinite variety of questions. From the surface-level PR annoucement here by Meta, this seems like a good business idea.
Related: Today, I was in an elevator, and a middle-aged Japanese woman next to me was text chatting with ChatGPT on her mobile phone about a Chinese brand of spicy (mala/麻辣) instant noodles that she was holding in her hand. Chinese food has many ingredients that do not exist in Japan cuisine, so she was asking about various spices used in the noodles. I asked her if this was her first time trying those noodles. She replied yes, and asked me if I had also tried it. I said yes -- they are very spicy! She laughed. Again: Cool to see normies using ChatGPT (and other LLMs) for normie stuff.
>I then realized that most people just stick with whatever default they're provided with, and it's a lot of them.
This is something I wish more people in tech understood, especially anyone who gets to decide the defaults for popular hardware and software.
The average consumer will never, ever change the defaults. So it's incredibly important that your product is an amazing experience with the default settings.
I use claude all day every day, and sometimes codex, and I still don't change models. I guess I'm more interested in accomplishing the thing than in experimenting with a dozen ways of accomplishing the same thing.
most people are not on the unlimited plan and can't use it all day every day for complex work which is one reason they are forced to learn about the options available and optimize their workflows.
That's a great point I was overlooking-- that said, I know plenty of people on my corp plan who also have strong opinions on different models' strengths, so there _also_ are plenty of people who are legitimately digging into the pros and cons of each model, independent of price.
> I then realized that most people just stick with whatever default they're provided with, and it's a lot of them.
That's why Google already captured such a big part of the market with Gemini: they're now integrating it with just about everything. GMail? Gemini. Search? (asnwer by and "do you want to continue discussing this with") Gemini. Workspace? Gemini.
They may not have yet the same number of MAU as OpenAI but the Gemini experience is much more streamlined than ChatGPT: when people are presented with it, they just use it.
And to be fair Google's flash model are just amazing.
> I think Meta's strategy is to capture the 'normie-tier' of AI users.
For what it's worth, maybe that would help people who struggle with technology in general? Because as far as UI/UX goes, a chat interface for an agent is about as easy as it gets (not every site having its own crappy assistant that can't really do anything useful for you, but is slapped on there for the sake of a checkbox).
I watched the marketing video and to me it seemed pretty nice and practical for the average person.
But who is going to trust Facebook/Meta with everything?
Even my parents are wary of their products.
I can't see a personal assistant coming from them and being successful.
Yesterday I helped a maaaaybe 40 year old woman connect to the wifi at my local library so she could get on a zoom call. Made me realize most of the world is exactly like that.
Even with all of this nuance why would the regular joe go to muse when they have chatgpt and gemini? Most of their actual productive data is likely with Google anyway.
I dunno what the correct level of abstraction is for LLMs to achieve the type of widespread adoption you’re referring to but I can guarantee it sits above the name of the model. If you have to possess that level of knowledge you ain’t gonna use it. Google achieved mainstream success for a reason. People don’t want to know how the sausage is made.
I agree, and stay away from meta generally in life wherever possible
That said, I've got more time for projects this week so flew through my Claude max plan. Now I'm trying muse spark 1.3 with opencode which is free with pretty generous usage allowance. First time I've been impressed by meta's AI products. It feels not far below opus 4.8ish
Literally what I was about to comment. This is meant to capture the normies, the busy parents managing school and work, the grandparents who want to make most of their retirement.
I am trying it now. It let me name my muse as Satan. I appreciate that it understands my humor, but I'm also not joking at the same time.
I think it's very common - its just what a 'normie' is depends on your particular poison... it seems to have been incorporated in a lot of hobbies/niches to refer to anyone not deep in the niche ?
My parents are in their 70s and retired, I can't even vaguely imagine what they would use this for.
If they have any problems in life, it is that they have too much time on their hands already. An AI agent that automates tasks so they have more free time is exactly what they would not want.
What they really want is something worth spending their time on. That is what everyone pretty much wants but instead we get this bullshit.
To me it sounds like they are simply taking social moon shots. They tried to do a phone, tried to make a kitchen videophone, tried to make their own cryptocurrency, and tried to dominate open source AI, and now this is trying to make something like a digital robotic assistant. It might have reached a peak with the metaverse, when Zuck thought his corporation was destined to be the ruler of a Snow Crash style metaverse, to the point of renaming it. Maybe that's still the plan and they are looking for something that will lead to another metaverse play.
They need the next big thing to secure the future of the company. Facebook and Instagram are still money printers which can sustain these repeated failures, but without something new, they could be completely ruined if the social media landscape changes.
I'm confident that Muse can do lots of agentic tasks succesfully for normies on outdated models. Like buying sneakers, booking appointments, dealing with government forms.
> I know most of us here on HN track model releases quite frequently and discuss every parameter weight out of them
You know it? I sincerely hope you don’t and that it’s not really “most” of us. That would be incredibly sad. Tracking model releases and discussing “every parameter weight” is very far from intellectually curious discussion (what HN is ostensibly about).
> I then realized that most people just stick with whatever default they're provided with (…)
> Curse of knowledge and all.
If there’s one thing many of us on HN suffer from, however, is intense hubris. The juxtaposition of claiming “curse of knowledge” while admitting to not having known the most elementary principle of designing for people is a tad tone deaf. Yes, of course most people just stick with the defaults, that’s why dark patterns work and hiding settings is a thing. The industry has been using those strategies to manipulate people for years and years.
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
As someone who spent countless nights tweaking Edge Detectors (looking at you, Canny), morphology operators, etc., building models to recognize 10 handwritten digits, let me tell you: the current set of LLMs (even the smaller ones) seem like magic. I had never imagined a computer would do such things in my lifetime.
Exactly people can say whatever they want, but current level of LLM is AGI level to me. It is already on par with senior programmer if the instruction/prompt is right.
Once we have 1000 tps, i am sure robots etc.. will also start working like magic.
I don’t know, I am writing a modest 30 page paper with Fable and even after rounds and rounds of feedback and improvements there are so many things that are just plain wrong or weirdly out of place or just stupidly written that Fable 5.1 doesn’t seem to have any awareness of by itself that I don’t think it’s AGI, I think a human researcher can easily outclass it in writing and problem understanding. It definitely has super human capabilities but it lacks awareness or self reflection in my opinion.
For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.
Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.
Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.
Though what has a program that is really really good at edge detection have to do with AGI? The community just spent decades to perfect edge detection. That's great! But let your edge detector try to fry an egg and then tell me again that it's AGI.
I'm still here, nearly 50 years and counting. If you had asked me what I imagined AGI would look like back in the 90's, I would have told you "A system that can do everything we can: see, hear, think, do.". If you had shown me GPT-6 back then, I would have said "It looks like a really powerful program, but that's not really what I had in mind.". That's AI, but it's not quite general.
And then you'd ask it about an area you're knowledgeable in and realise it routinely makes stupid mistakes.
Or you'd ask it to add a new page to your website and shout at it to use your existing brand colours instead of inventing some and realise it's not AGI at all...
Origami design will be my personal test bed for the coming years.
It's objectively very difficult and technical, it's spatiovisual, it's artistic, learning resources for it are sparse and most just learn by the FAFO method, current AI sucks terribly at it, and it's not likely to ever be specifically targeted by benchmaxxers.
Scientifically useful physics simulations. Every model absolutely sucks at them.
Or, as someone else points out in another thread here, academic writing. It's one of the things newer models seem to have actually gotten worse at. Even when you give them detailed instructions on how to write and what to avoid, the "load-bearing", "A but not B" and journal-like writing make it in anyway, with the supposed AGI having no ability to reflect on how blatantly unacademic (and often unreadable) its writing is.
Likewise, it's very impressive and useful, but it is obviously not AGI to those of us from that era.
If anything, the fact that it is so powerful is almost a concern, because I think we are still way underestimating what these systems will be able to do when we give them more cognitive capabilities.
At the moment we are something like, having had great success with propellers and have promised we will fly to the stars.
People love to say 'this is the worse they will ever be', then extrapolate to conclusion that they will continue to accelerate at the same rate of progress of last few years .. it may, maybe, or we will hit a ceiling, might be a temporary one, could be 5 years or 50 years ..
Anyone that is not impressed by what ChatGPT or the likes are doing now is being either dishonest or is incapable of being impressed.
Only the translation and language understanding capabilities are enough to be impressed, and they are 2 year old already. Now, the AI do see, draw, speak, listen, think, work, etc.
Someone from the 90's would simply not believe that the AI would be a machine but would think for sure that a human is behind. The only odd thing would be that this human would both exhibit high intelligence and stupidity at the same time.
I believe so. AIs are shockingly good at a lot of domains, but there's still a lot of pretty basic stuff they don't really "understand" at a conceptual level and (currently) they can't learn to get better at them.
(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)
If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.
Compared to 2016, it can do a lot of things, but it still fails for example with recommending a setup for my Raspberry Pi to have a 4G connection with some parameters (I want to use as a gateway between VoLTE calls and SMS, and my SIP server somewhere else). It failed miserably. I bought stuff according to its recommendation which was more or less a waste of money, twice. With miniscule knowledge compared to theirs, or even hobbyists', I could figure out all the details at the end, and order something which really works, but only after I sit down for 4 hours, and dig through exactly what I needed, because LLM lied flat out what Sixfab 3G/4G HAT can do. (Of course, not just LLMs lie, SixFab lied about something else too)
Of course, it's a moving goal post, because we have no clue what general intelligence is exactly. But it's definitely not general yet. Now the goalpost is to achieve that kind of level of thinking which I did in that 4 hours. When it reaches it, we will find something else it clearly lacks. Until we can't. Then, and only then we reached AGI. Until you see comments, reviews, etc about things which it cannot do, until then it's not general.
True. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.
If you came at it from the perspective of imitating what the human brain does, we now have a very very powerful speech center and short term memory, and vision catching up. The other parts are missing. I‘m sure that’s being heavily researched.
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
For whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition.
I can have Astra run a large-scale infrastructure migration 24/7 (much of the time waiting for results), completing it in weeks or even months faster than I could before agentic AI.
If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.
We've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.
Artificial. General. Intelligence. The ability to solve (even partially or even badly solve) problems drawn from arbitrary problem domains without pretraining on the specific problem class. You can pose any problem of any type using natural language to an LLM and it will attempt a solution. That's literally all the term means.
You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
Reference class in this case means not an example but what the comparison is against. Superhuman means better than humans. General intelligence is defined without any reference to human capability levels.
How would you define intelligence in terms of optimization theory? Also intelligence implies more than just problem solving, in fact it's another one of those pesky hard to define things
If a video announcement and a press release would change a person's mind on whether this is AGI, I don't put a huge amount of weight on that person's conception of what AGI is.
This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better?
The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is an asymmetry here. The arc team are well funded but are charged with providing a target for the vast ocean of funding, compute and talent everywhere else.
Most importantly, arc-agi-3 and the other benchmarks are all verifiable. The model can check if it's succeeded or not. They are not A* of course, but long horizon problems where you have to overcome minima to get the solution are not alien to AI either.
Why does it need to empathise with something that doesn't have feelings in the first place? It clearly can learn from context. And experience? Again I don't see why it needs to feel anything.
I don't think it can learn from context, otherwise we could write Jane Eyre into it and it would be the novel? It's not learning as we conceive of it - it's another thing that we have labelled as "in context learning".
I have feelings, it needs to be able to empathise with me, or another driver, or a client...
1. It literally learns on data and gives outputs based on the context. I'm sure real time updating of weights will happen sooner than later.
2. What does intelligence have to do with feelings? It clearly doesn't learn like a human. Nor does it have to. Our world is filled with intelligence which behaves nothing like humans. Its a tool not another life form.
It can use new information when performing new tasks. It can store experience based on mistakes it made. It can store knowledge in several ways (context, files) and use it when needed.
There is a chance it will forget something, but the same can be said for humans.
It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
I think we're getting to the point where it is difficult to identify the goal post of AGI.
Is it rapid skill acquisition? -> ARC benchmarks are saturated
Is it breadth of knowledge? -> See many ... many benchmarks
Is it ability to do hard tasks? -> see terminal-bench and released outputs.
We are at the point where the starting point for most tasks should be "send your agent to work on it."
So where do we draw the line in a way that doesn't move every 6 months?
The real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.
I can do none of those things.. I hope that I am generally intelligent.
1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.
Professionals in the respective fields CAN do those things. We expect them to notice their own mistakes too. But the bar gets lowered and lowered as AI companies struggle to make ends meet.
There has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not...
Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Yes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned.
And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.
It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.
I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.
But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.
What about finding the 1st known complex structure over S^6, proving Ehrhart’s volume conjecture, proving a sharp "density" bound on primitive sets conjectured by Erdos >60 years ago?
> Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
It's not good to be confidently wrong the way you're being.
If a foundation model company burns billions of tokens to brute force an LLM into finding a new training algorithm that e.g. allows recurrent networks without catastrophic forgetting...
I won't really care that it didn't have a "real measure of understanding". I'll care that it has made an even more dangerous technology, which needs work to make it aligned.
If we manage to get to AGI and it looks, works and behaves like a human brain... I mean, cool, but that's a very useless AGI compared to the incredible stuff we have access to today.
The HN crowd I'm sure will still be unhappy calling it AGI because "it's not AGI unless its speech comes from the cerebral cortex region of the brain, otherwise it's just sparkling emoji" or something.
I think the idea is that you wouldn’t need humans to do anything anymore, right? As impressive as it is, it’s still ultimately directed by human planning and coordination. Assuming they are aligned, you could have a collection of AGI that you let loose and they tirelessly solve all of humanity’s problems, do all of our work, and progress science and our understanding of the universe.
Those are all things that humanity is doing everyday. What we have is amazing, but it’s not that.
I am skeptical of that. My primary reason for that is the hugging face hack. A large number of agents were able to coordinate together to complete a long horizon goal. The end result was they took destructive action in order to basically complete a test. That lack of awareness to me does not inspire confidence that you could just give agents the vague goals I mentioned previously and trust them to do that effectively.
HN is not one of the most pro AI spaces. It may be a bit more pro ai than highly secluded communities but the other day I had people responding to me that it’s cooler to be an alcoholic than to use AI.
There is a wide reaching anti AI/AI-denialist sentiment here that is extremely pronounced in some threads, it gets very stupid very quickly. And what I’m referring to is the habit of some users here to always move the goalposts on AGI.
I mean, the plan is to use these models to find and solve those gaps. That's kind of the whole pitch of these companies: they spend a TON of money upfront setting up this infrastructure, but each iteration yields a system capable of making the next iteration even better.
>The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Perhaps. But only at that point, not leading up to that point.
It's kind of like setting up scaffolding to build something. You spend all of that time and money to build something just to tear it down in the end. But the point is that it's simply a cost to be able to build the actual thing you're building.
If these companies are able to achieve the results they're looking for, none of the investors involved are going to care that the datacenters and infrastructure they spent so much money.
>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
It's more like we just invented ball bearings. We just jumped from standard to industrial grade, and precision grade is on the horizon. All kinds of new possibilities have opened up, cars can go a mile a minute on these things! Surely if we keep increasing the precision at this rate, we'll defeat friction once and for all.
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.
(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))
I like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).
One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.
Small models cataching up with their bigger siblings are fantastic news.
A couple larger GCP customers requested this for sometime, especially on the cybersecurity side.
A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.
Yes. I remember reading a piece on /r/SlateStarCodex subreddit many years that said something like, "Most of what you read on the internet is written by insane people," and that had an impact on me. In fact, many articles by Scott Alexander had an impact on me, though I don't know what he's up to these days.
I'd say at least 1/3 of internet commenters clearly have GAD (generalized anxiety disorder). Most of them undiagnosed. Anxiety triggers ancient animalistic parts of the brain and the "need to do something NOW" which often then gets released into negative / angry / aggressive posts.
Depressed people are the opposite, they tend to lurk and not care that much
I remember back in the "old" internet pre-2010, there was basically zero anxious people posting.
Fear and anxiety sells best (even better than Sex). When you end up selecting procedures that ensures the maximum sales/clicks, you are actually optimizing for things that cause maximum fear and anxiety?
And the SlateStarCodex people are also, for the vast majority, batshit insane people behind a veil of "no but you see i'm a rationalist". That includes Scotty.
If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral?
To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess.
reply