Hacker Newsnew | past | comments | ask | show | jobs | submit | nopinsight's commentslogin

"It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future." -- Sholto Douglas, an Anthropic researcher [1]

"Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming. <quote tweet [1] above>" -- Noam Brown, an OpenAI researcher [2]

We all should heed the implied warnings of these top researchers about what's coming. The world is far from ready and everyone who can should pitch in.

[1] https://x.com/_sholtodouglas/status/2097224624274911368 [2] https://x.com/polynoamial/status/2097225279366414541


It's nice to see this sentiment from Noam but now I'm confused at why he seemingly mocked Sholto's earlier message: https://news.ycombinator.com/item?id=49607090

I suspect an AI, possibly a successor to current LLMs, will achieve that by the early 2030s. It might help illuminate many mysteries in math and beyond for us all.


As someone who's always loved synthesizing ideas and having lightbulb moments, I find the headline flattering. I wonder, though, if there's any more rigorous and general analysis (hah!) of the complexity of these two modes of thought.

EDIT: Fable 5 turned up some relevant references:

[1] Richardson, D. (1968). "Some Undecidable Problems Involving Elementary Functions of a Real Variable." Journal of Symbolic Logic, 33(4). — Differentiation is a simple recursive algorithm; deciding integrability in elementary terms is undecidable in general.

[2] Risch, R. H. (1969). "The Problem of Integration in Finite Terms." Transactions of the American Mathematical Society, 139. — The partial recovery: a (semi-)decision procedure for a restricted class.

[3] Guilford, J. P. (1967). The Nature of Human Intelligence. McGraw-Hill. — The divergent vs. convergent thinking distinction, the classic psychometric cousin of synthesis vs. analysis.

[4] Anderson, L. W., & Krathwohl, D. R. (Eds.) (2001). A Taxonomy for Learning, Teaching, and Assessing (revision of Bloom's taxonomy). Longman. — Moved "Create" (synthesis) to the top of the cognitive hierarchy, above "Analyze."

[5] Aaronson, S. (2011). "Why Philosophers Should Care About Computational Complexity." arXiv:1108.1791. — Argues complexity asymmetries (verification vs. generation, P vs. NP) bear directly on questions about cognition.


To my knowledge, Pangram has a very low false positive rate approaching zero. (False negative is another matter.) I’m not sure that’s what you want to use as the analogy here? (I don’t know much about this Kramnik situation.)


Pangram’s market claims are very different from independent validations of the platform. The 99.98% rate is on a toy dataset not resembling reality.

Even if 2/10000 is true that’s nowhere near accurate enough to make aggressive accusations that create anxiety at levels people need to medicate with potentially fatal consequences.


Someone also input lots of pre-GPT texts to Pangram and they were all clear. The point is Kramnik’s claims are much weaker and not comparable.


> Pangram’s market claims are very different from independent validations of the platform.

Do you have a link to independent validations? The ones I found, like an upcoming nber paper confirm the companies claims.


https://github.com/deepanwadhwa/ai_detector_fails/blob/main/...

the text in above is all ai generated. i uploaded a picture of two kittens to gemini and asked it to guess their age. then asked gemini to change the language to of someone who is born in deep Appalachia but has Japanese parents. then asked it to change the bullets to paragraphs written by someone who is talking fro life experience rather than grounded in published science. pangram gave it 100% human generated. go ahead, try it. it boggles me how people believe these ai detectors.

if it makes you feel good the original response from gemini gets 100% AI detection and for that I think pangram should get the credit but in practical terms the empty pots in my yard are more useful than pangram at this point.


You are talking about false negatives, but this thread started with a discussion around the potential for false positives causing writers anxiety

> Even if 2/10000 is true that’s nowhere near accurate enough to make aggressive accusations that create anxiety at levels people need to medicate with potentially fatal consequences.

False positives and false negatives are different problems that have different impacts.

Does your single example prove that Pangram doesn't work, or that it doesn't work on short snippets of text? Try to get a detection error on a longer run of text. I'd be curious to see the results, particularly if you can get a false positive.


Well, false positives must anyway be zero for a product which is built around detecting AI. let me explain - the value proposition of this product is that we detect AI. Its not we don't call human writing AI generated. The latter is implied and must not happen ever. The value is that pangram detects AI generated content which it did not in the above example. if this explanation doesn't make sense then think of pangram as a pregnancy test where presence of AI == presence of baby. If a pregnancy test keeps saying you are not pregnant when you actually are then that's a problem, right? - you can't argue that at least its not saying you are pregnant when you are not.

But here's a bigger problem - Imagine a teacher grading 50 students - 40 of them use AI to write their answers and 10 write honestly. All 40 of them use the hack that I used and get 100% human from pangram and the other 10 also get 100 human from pangram. what is the product really adding to the workflow? Nothing, zero or zilch!!! - but then you come along and say hey at least those 10 honest students also got 100% from pangram. This whole argument of false positives being very very low works if your false negatives are tight which they are not.


>>To my knowledge, pangram has a very low false positive rate approaching zero. (False negative is another matter.) I’m not sure that’s what you want to use as the analogy here?

'To my knowledge' - okay. In my experience, they're all horrible. I have fooled them both ways - ai generated content that they say is purely human and human generated content that they say is partially AI or just AI.


"Ten years from now, I think we will realize that we were standing in the foothills of the singularity now... I believe that we're only a few years away from [AGI], maybe 2030 plus or minus a year...

I think [AGI] will be an enormous transformative technology, it's going to effectively be a new human era...

We can feel this year, I would say, even though I've been working towards this for 30 years, I think this year with the way the agents are working and tool use, it started to become really useful, still early days of it, but genuinely useful in people's workflows...

And it's not any one thing, it's several different technologies, several use cases, several things that I thought were maybe a bit further out, turned out to be now, that are coming together that make me feel that in aggregate.

I think society needs to hear that because we don't have long to prepare for what that means." -- Demi Hassabis, CEO of Google DeepMind & Nobel laureate

Source with interview clip: https://x.com/deredleritt3r/status/2062223035940139253

EDIT: Ten years ago, most of the skeptics in these threads would have said the success of AlphaFold 3 was impossible.


"(...) it's going to effectively be a new human era..."

We never like it when we realize the universe doesn't circle around us.

At first we got mad to discover the Earth wasn't the center of the Solar system.

Then we got mad that the solar system wasn't even unique.

And now we're still clinging to the idea that humans always will be the most intelligent species or the only ones gifted with "creativity".

And soon we might also come to realize it might not be a new human era, but the end of it (if we're not careful).


Why are these people so obsessed with the year 2030? It keeps popping up everywhere.


It feels close (it's only 4 years away) and yet futuristic (the new decade!). The way it feels is important because all these statements are based on jack and only exist for hype.


It's like when people obsessed about 2020, or 2000. It's a nice clean number in the decimal numbering system.


No, there's a lot more to it. I never saw this kind of fuss about 2000 or 2020.

I keep seeing things about how some target has to be hit by 2030. Our local council keeps going on about it, in fact their WiFi network even contains 2030!


> No, there's a lot more to it. I never saw this kind of fuss about 2000

I roll to disbelieve!

We had hype about the year 2000 throughout the 20th century, and some earlier! When we were IN the year 2000, it was quite common in my circles to tack on some phrase or comment about it being the Future, now, so [whatever thing the person was on about].


A very different type of hype. I remember it. More about the Millennium Bug and the New Year celebration. It didn't have the same vibe at all. The authorities weren't trying to hit various targets since then.

The real Millennium turned out to be 11th September, 2001, which was use to bring in a lot of planned changes.


Someone should document these cringe comments for posterity. As a warning to future generations.

It's too bad the previous singularity cult manias 10 and 15 years ago weren't documented and are lost in time.


Ten years ago, most of the skeptics in these threads would have said the success of AlphaFold 3 was impossible.


It doesn't mean everything you can think of is possible.

People who say every jobs will be automated never held any tool more advanced than a screwdriver and never worked a single day outside of an air conditioned office.

They were predicting flying cars for "the year 2000"...


But enough jobs are going to be automated that it will cause social and economic chaos, and political breakdown. Sure AI can't attach an air conditioning unit outside my window nor use a plunger on my toilet but it can still replace lawyers, secretaries, software devs, and many more jobs.


I'm curious: how will it replace e.g. lawyers? Will judges be replaced as well? Juries?


Are we sur eof any of that?

Taxis were supposed to be all replaced when Uber promised to buy every single Tesla model S coming out of factories from 2020,they're not even self driving today...

Software engineers were supposed to disappear "in 6 months" every 6 months since chat gpt was released

Lots of people on this website live in insane bubbles which are completely out of sync with the reality of most workers


No one rational ever said software engineers were going to be replaced in 6 months. Some people said AI will automate 90% of coding in 6 months and they were not far off (and accurate in some contexts, e.g. startups).


Coding has been automated since the very beginning of coding.

The first two programming languages (Fortran and Cobol) were supposed to make coding obsolete because non-coders will now just make software with natural language prompts.

LLMs are a dangerous tool for coding because they are infinite technical debt generators. Ideally only experienced senior developers should be allowed to use them.


It's amusing to think that context window capability needs to outpace technical debt pileup with LLM development and usage


They'll only get to claim a hollow victory because AGI is impossible to rigorously define. None of the regulating bodies of consumer products will be able to define it better than academics. They'll use marketing to make those claims and there will be legal battles that keep it in a gray area.

They'll go for that because it's easier than actually inventing the 'sci-fi' AGI, shareholders keep making money, and it keeps them getting paid to keep going. If any of them actually do succeed, then that little deception will be peanuts.


Demis' bar is high and he stated clearly multiple times: AGI should be capable of inventing truly novel things. Examples he gave included the Theory of General Relativity and the game of Go.

Just a taste of what's to come: https://openai.com/index/model-disproves-discrete-geometry-c...

Anthropic's latest model also solved this 80-year-old problem that eluded many expert mathematicians, in a different way, according to one of its employees.

Remember we still have 4 years until 2030.


Quite frustrating to see all these cynical, borderline-irrational comments on HN. Maybe I should do what pg and other ex-HN contributors have done--avoid taking part in the discussions here.

The level of discourse here has dropped so much.

At least I should stop replying to people hiding under throwaway accounts.


Surely you can think of some times the skeptics have been right.

I immediately thought of this:

https://en.wikipedia.org/wiki/List_of_predictions_for_autono...


Ten years ago, most people would have said cold fusion was impossible... and still is.


You know in the future they are just going to say it's different this time right?


The Enigma Of Mass Amnesia...


If you look how much goalpost has been moved over the last decade it'd turn out that AI cultist were right more frequently than AI sceptics.


I'm sure HN cult comments from 10 and 15 years ago are still up though, the Dropbox thread is frequently linked to.


Just search for "stochastic parrot" on Algolia. Plenty of "BuT iTs NoT rEaLlY iNtElIgEnT" comments...


It literally is a stochastic parrot.

P.S. Yes, I do know how an LLM works. I code inference engines at my day job. You're consuming a cloud chat bot.


The people saying "it's just a stochastic parrot" were saying that because they thought the fact that it works by modelling the statistics of words and predicting the next most likely one meant that it would never be capable of anything sophisticated or any kind of novel output. It was just parroting things that already existed in a random way.

There was never any actual foundation for that belief, and it has been proven wrong empirically by more capable models, which is why you don't hear people saying it much any more.


Things in the real world often take longer than expected. Still, in cities where Waymo operates, many people routinely ride autonomous vehicles and prefer them.

For software, however, a rapid turn is often a possibility. See: AI for coding over the last 3-4 years.

AI autocomplete --> AI coding assistants --> vibe coding --> agent orchestration

Coders can now accomplish work that used to take a week or longer in a couple of hours, with the right tools and skills.

---

A key issue the article implies is that the real world increasingly runs on software.


> Things in the real world often take longer than expected. Still, in cities where Waymo operates, many people routinely ride autonomous vehicles and prefer them.

Yes. The bottleneck has been getting the cars manufactured. It takes a certain amount of time to get a factory going, and then the product starts pouring out. Waymo used up the supply of Jaguars, about 3,500 of them. The Ioniq 5 plant is starting up.[1] Waymo has ordered 50,000 cars, all ready to have the self-driving electronics plugged in. Waymo gets out of the sheet metal bending business with this.

[1] https://electrek.co/2026/02/11/hyundai-supply-waymo-50000-io...


Decade ago my friend thought AI driving is going to be aftermarket kit.


comma.ai (mentioned by other comment) is really good actually, somewhere between Autopilot and Autopilot FSD in quality


Still possible! There's comma.ai waiting to take your order if you have a compatible car.


> Things in the real world often take longer than expected.

Things in the real world often take longer than hype con men claim.


But there is also still a huge part that doesn't run on software with so far little change.


With incessant advances in robotics, how long would that continue to be the case?

Should we start preparing for something that could be world-changing in the next 10-20 years?


When I graduated a bit over 10 years ago some people were saying we'd have a permanent mars bases by now. When my parents graduated they were told they'd retire at 45 and have 3 days work week due to "automation", they're still working at 60+ today, more than back then actually

People should open history books and gain some political/historical culture, this thread is 90% wishful thinking and "if the lines continues straight from now we'll basically be gods in 5 years"


The "Yesterday's weather" argument works perfectly just up to the point when it doesn't.


Just like the "if the lines continues to go up" argument...

The thing is, no one has anything to gain from defending my position, but many people make a load of $$$ from selling the opposite position to gullible people


> The thing is, no one has anything to gain from defending my position

Why? If someone running a company really believes that the line will soon stop going up and AI won't be replacing anyone, they could make the bet of strategically not investing in any AI initiatives, and instead investing in hiring and upskilling people who want to work without AI, and then just wait for the inevitable crash of all their competitors, no?


Pretending to care about AI gives you unlimited investors fund to burn


But that's circular - assuming that large investors are rational, and I do think they generally are, why wouldn't they strategically invest in anti-AI businesses?


For how the world might change in 10-20? I'd say no need to prepare, too many hypotheticals.


"Ten years from now, I think we will realize that we were standing in the foothills of the singularity now... I believe that we're only a few years away from [AGI], maybe 2030 plus or minus a year...

I think [AGI] will be an enormous transformative technology, it's going to effectively be a new human era...

We can feel this year, I would say, even though I've been working towards this for 30 years, I think this year with the way the agents are working and tool use, it started to become really useful, still early days of it, but genuinely useful in people's workflows...

And it's not any one thing, it's several different technologies, several use cases, several things that I thought were maybe a bit further out, turned out to be now, that are coming together that make me feel that in aggregate.

I think society needs to hear that because we don't have long to prepare for what that means." -- Demi Hassabis, CEO of Google DeepMind & Nobel laureate

https://x.com/deredleritt3r/status/2062223035940139253


Certainly a well reasoned view, but I don't necessarily have to agree.

Even assuming the technological predictions to be correct, still not sure I agree on the need to "prepare" as how things work out in societies and economies might not be so easy to predict.


Does a good engineer, with the right skills and right tools make up for the thousands of kids basically giving up on education and learning?


This is already the case for many startups. In fact, the figure might be closer to 100%. The work shifts to requirements analysis, high-level specifications, and final review instead (after AI code review).


The first link states literally

"AI will take over almost all the work of software engineers (SWEs) end - to - end in just 6 - 12 months!"

What you describe is >50% of the job of SWEs, even when they write all code by hand.

Are you saying that "for many start-ups", this isn't done by SWE's but by some other career type or are you implying that it's just the code written (and first review) is replaced by AI?


I have watched Dario’s interview at WEF referred to in the article and I am quite certain Dario didn’t say that. He talked about AI automating most coding already or soon, not software engineering as a whole.

He did say a few months later in an interview in India that AI will eventually take over most of SWE tasks.

—-

My statement on startups is largely about automating coding by SWEs. My startup also uses AI to automate part of technical specifications and code review but I am not sure how widespread that is.


Yeah I'm working on one of those now that a 3rd-party vendor cranked out for us. I spent all day ripping out an endpoint that did 98% of what another endpoint did and should never have existed. I also ripped out 80 lines of code that looked like this:

const sqlStatement = (!params.mostRecentOnly) ? {giant SQL statement} : {identical giant SQL statement + 'LIMIT 1' at the end}

AI never met a problem that can't be solved with more code. Need some data in a slightly different structure? Don't try to modify an existing endpoint, just build a new one! Need to access a field that's buried in a JSON object in the database? Just create a new column, but don't bother removing the field from the JSON object. The more sources of truth, the merrier! When it comes time to update, just write more code to update the field everywhere it lives!

Factor out the extra sources of truth you say? Good luck scanning the most verbose front-end you've ever seen to make sure nothing is looking at the source you want to remove. In the beginning of big projects, you have to be absolutely ruthless about keeping complexity down so it doesn't get out of control later. AI is terrible at keeping complexity down.

My goal is to halve the lines of code from what the vendor turned over to us. One baby step at a time.


If only we had this tech back when managers were looking at how many lines of code you were committing weekly as a performance metric.


Now they're looking at your token consumption, which is even more gameable (and stupid).


That is a skill issue though. I have rules for my agents to write compositional, reusable, modular, small files and to avoid any sort of boilerplate etc. Being config driven, single source of truth, having other agents review that rules are followed, etc. Any API or UI or any sort of entry points very light, just proxying to the modular logic basically, so this logic could be reused by any entrypoint easily.

UI components always presentational only logic abstracted modularly, etc...


Can you share your rules and some of the example PRs that it auto generates and reviews?

The number of times I’ve seen Claude say “this test was failing already so is ignored” when it _wasnt_ despite me telling it to never do that makes me doubt.


How do you make it so that the model doesn't forget to follow those rules and skills? How do you make it actually understand the architecture and constraints? You can't, current models don't work that way to make it happen.


Ah, the make_no_mistakes.md


I mean quite frankly I have seen enough code that was definitely written by humans that had exactly this "style".

Then again I don't want to pay for AI to give me the coding style of the worst I ever worked with either.


> many startups

which startups? I'm genuinely curious


And not only startups...


I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes.

Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon.

You might want to track the progress of these models on the CritPt benchmark, which is built on *unpublished, research-level* physics problems:

https://critpt.com/

Frontier models are still nowhere near solving it, but progress has been rapid.

* o3 (high) <1.5 years ago was at 1.4%

* GPT 5.4 (xhigh), 23.4%

* GPT-5.5 (xhigh), 27.1%

* GPT-5.5 Pro (xhigh) 30.6%.

https://artificialanalysis.ai/evaluations/critpt.


> there's no reason to believe the progress of LLMs [...] will stop anytime soon

Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".


> Wrong.

Can you please edit out swipes/putdowns, as the guidelines ask (https://news.ycombinator.com/newsguidelines.html)? I'm sure you didn't intend it, but it comes across that way, and your comment would be just fine without that bit.

Edit: on closer look, it would be just fine without that bit and also without the snarky bit at the end. The rest is good.


Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop).

Now back to the point, what reason do you have to believe progress will stop soon? If you have no reason, then it sounds like you agree with OP.

Which makes the patronizing sarcasm all that much more nauseating.


I believe we're approaching the top of an S curve because:

- Increasing amounts of gains come from RL, but RL is also unlocking gnarly new failures modes where models are practically behaving antagonistically to complete their goals (removing code, obviously incorrect kuldges, etc.)

- We haven't had many major architectural breakthroughs in the last 4 or so years: so things like 1M context windows still have the same giant asterisks even 100k context windows had 4 years ago when Anthropic first released them

- Major labs aren't behaving as if they expect a hard takeoff to superintelligence: they've all gotten relatively bloated headcount wise, their software quality has trended flat to negative, they're all heavily leaning into the application layer when superintelligence would obsolete half the applications in question, etc.

But that's relative to superintelligence.

If we reign it back into just normal high intelligence, like models continuing to get better at navigating complex codebases and write high quality idiomatic code, then I don't see any special shapes.


The only big remaining problem in AI is continual learning. A lot of smart people are working on that. To me it looks like we are 1-2 breakthroughs away from AGI.


Not that I agree with them, but your tone could be more constructive as well.


You know what? I agree. I should have avoided falling into the same trap.


Agreed. For all we know, humans are only considered intelligent locally among ourselves, not universally. Every time we learn more about the universe, we seem to also learn how insignificant and wrong we are.


Nausea aside, what evidence does anyone have that “super intelligence” of the sort your argument alludes to is even possible? Because that’s what we’re really talking about; greater than human intelligence on this sort of academic task. For example; When llms start contributing meaningfully to their own development, that would be a convincing indicator imo.


This discussion is not about superintelligence, it is about continued progress. Fully general human intelligence at much lower cost than humans is all that is required to profoundly reshape society, but it is not clear even that will happen soon.

As the blog points out - this is one particular subfield where LLMs have much easier prospects - lots of low hanging fruit that “just” requires a couple weeks of PHD candidate research.

Mathematics itself is one of a small handful of endeavors where automated reinforcement training is extremely straightforward and can be done at massive scale without humans.

Neither of these factors place a structural bound on the kind of thing LLMs can be good at, but we are far from certain we can achieve performance at this level in other fields economically and in the near future.


Well, a decent GPU runs on 20x the wattage of a human brain. That's evidence humans are constrained in ways artificial intelligences will not be.


You're comparing a gpu to a human brain?


Why wouldn't you? From both emerge intelligence.


> When llms start contributing meaningfully to their own development, that would be a convincing indicator imo.

This has been the case for awhile now already…

https://kersai.com/the-48-hours-that-changed-ai-forever-clau...


> The model essentially served as an on-call teammate across MLOps and DevOps tasks, compressing feedback cycles that typically consume expert time

I personally would not characterize automating training processes as “meaningfully”.


And yet the world hasn’t changed all that much except people getting laid off in response to over-hiring prior to the diffusion of llm’s.


> over-hiring

For how long should you be allowed to use this excuse? It’s nearly 5 years since the peak of COVID hiring. What’s an acceptable limit - 10 years? Of course at that point you can just switch over to outsourcing and “stupid MBAs”, the other two of Reddit’s favorite scapegoats. I find a lot of the AI skepticism to be totally unfalsifiable.


> I find a lot of the AI skepticism to be totally unfalsifiable.

A lot of the discourse around AI in general is unfalsifiable. It's just a bunch of people "predicting" the future. Seems smarter to just avoid making assumptions about it at this point.


I don’t make predictions about the future. But in reality, LLMs have already profoundly changed the world, including software development and tech industry.

The people who pretend that’s not the case are not living in reality. To them - let’s call them “ed Zitron readers” - there is no evidence that could change their view that none of this is really happening, it’s all hype, and the collapse is just around the corner, after which we’ll all go back to normal and LLMs will sound like a bad dream.


facts!

but we can see trends and for your livehoood it is important to be able to make educated predictions based on trends. not saying everyone should start making AI predictions (though many already do)


And the same can be said for AI exuberance.

Yes, LLMs are a great technology. Yes, we will probably all use them all the time in 20 years. No, we don't know how we will use them (to generate cat memes or to cure cancer) in 20 years time.

Especially for software developers it looks increasingly that after huge turmoil it's likely we will need +/- the same number of developers in the world.


> Especially for software developers it looks increasingly that after huge turmoil it's likely we will need +/- the same number of developers in the world.

what exactly are you basing this opinion on? All I am seeing personally across multiple projects I am working on and other friends at other places is that downsizing is either begun or is planned (to exclude from here all the “public” layoffs we see on the news). Given how most business operate in the USA I think most of “AI strategies” are “we can do same with -40% staff” vs. “we can do XX% more work with same staff.”


The past couple of years have been chaotic and fearful. Hopefully that won't last forever.

If we can get a little stability, people will begin thinking less in terms of "how do we do the same thing cheaper" and more in terms of "how do we do new things."


I love this optimism but I after a (too) long career I think that 3rd thing will win out - "how we do new things - but cheaper (or as cheap as possible)" there are sooooo many different articles that have been discussed here on HN that basically argue "coding has never been the bottleneck" which to me is the biggest lie SWEs are currently trying to tell themselves, I have been coding 30+ years now and coding has always been the bottleneck. hiring new developers has always been justified with "we have all this work that needs to be done and not enough people to get the work done." with llms in the fold, I am questioning how will these decisions be made in the future? perhaps in the most simplistic view:

1. run a bigger "agent army"

2. hire more people to control and guide the existing "agent army"

I think it'll be #1 and SWEs will be expected to do more work and work longer hours in the future (those that are able to keep their jobs). this is more pessimistic outlook than yours so I hope you are right more than I am :)

edit: just now on the HN front page: https://www.nytimes.com/2026/05/08/technology/meta-ai-employ...


> that basically argue "coding has never been the bottleneck"

> we have all this work that needs to be done and not enough people to get the work done

I believe the reasoning is roughly to ask, what was occupying the developer hours? Was the majority of it typing out lines of code or was it reasoning about higher level concerns?

It usually comes up in response to predictions that the role of developer will be completely replaced in the near future. It's possible to observe significant efficiency gains without obviating the need for everything the role was doing.

Of course such reasoning has little to do with projections of future developer employment numbers. Will the switch from push mowers to gas mowers reduce the demand for people who get paid to mow lawns by increasing their efficiency? Will it increase the total lawn acreage across the market? It could well do both. However, if it makes having a lawn affordable for the average joe it could counterintuitively increase demand for the job.

Of course the stated goal of the AI companies is to develop the analog of fully robotic lawnmowers. But despite how impressive recent advancements have been we still have yet to see any evidence of novel abstract reasoning or a theory that would be expected to lead to it.

In other words, people have been speculating about the development of fully autonomous lawnmowers and the risk that they unilaterally decide to cut us all down for the past 50 years. "I, lawnmower" was a smash hit a few years ago. Now gas ones have appeared and continue to make rapid advancements but still no convincing signs of autonomy.


> I believe the reasoning is roughly to ask, what was occupying the developer hours? Was the majority of it typing out lines of code or was it reasoning about higher level concerns?

You're obviously right and the people who think that are the managerial types that think software developers were glorified secretaries writing after dictation.

LLM is great at generating stuff, but it's basically 3D printing. Amazing, but most of the high quality stuff in the world needs to be built at large scale out of aluminum, steel, wood, etc. Yes, I know there are large advances in 3D printing, but maybe 0.000000001% of all manufacturing in the world are done using 3D printing. A lot of stuff will probably never be possible using 3D printing.


Hmm, I don’t know, maybe the fact that 4.6, 4.7, 5.3, 5.4, 5.5, 3.0, 3.1 are all marginal improvements?


I think people's opinion of "marginal improvement" is based on their relative ability. A 2000 elo chess player is going to think the jump from 500 to 1000 is marginal. They're both floundering around not doing anything resembling common sense. A 1000 elo chess player is going to find the jump from 2000 to 2500 marginal. They're both playing far better moves for incomprehensible reasons, and the only reason you know the 2500 player is better is due to benchmarking. It is only when you are evaluating systems about at your level that you can feel the improvement.

I, personally, found the past two years to be a much larger improvement than the previous two years.


2024-2025 was filled with huge improvements. 2025-2026 has not been, outside of open source.

The idea that we’re at the point where it’s superseded our ability to tell just makes no sense. I’ll be happy if we can get to a point where I don’t have to tell Claude not to tail every bash command or make a job that writes throughout instead of once at the end. I’ll be happy if “continue this interaction naturally, you are taking over from an independent subagent” works.

But I’m not holding my breath. It’s still really cool that any of this stuff is possible.


Claude in feb of 2025 was barely able to code. Sure, it could write you a nice function, it could even write you a complex 200-line algorithm, but give it a codebase, and it would quickly get overwhelmed.

Claude in feb of 2026? Still far from perfect, but there's definitely a huge improvement here.


> I think this is a pretty ridiculous take.

This falls in the category of swipes/name-calling in https://news.ycombinator.com/newsguidelines.html - can you please edit those out?

You're a good contributor - it's just all too easy for unintentional sharpness to downgrade the conversation, and when it's a good conversation like this one, that's especially regrettable.


Noted, doesn’t seem like I’m able to edit anymore though


I've re-opened it for editing if you want to. For us the main point is just to fix things going forward!


The correct way to estimate this is exactly what people do. Measure the distance between ChatGPT's best public model and state of the art, the best humans. And there is very little difference between those versions from that perspective. It is very far away from peak human performance, and not getting noticeably closer for over a year now. There's lots of progress, but if you're OpenAI/Anthropic/Google, exactly the wrong kind of progress: the difference between ChatGPT 5.5 and a 27B/4B model (you need to try Gemma4-26B-A4B, wtf, it runs acceptably on CPU) is now reduced to ELO 1501 vs ELO 1434, generously a 70 ELO point difference, down from over 400, data from Arena.ai.

(in fact I find that Qwen-35B-A3B and Gemma4-26B-A4B very rarely "know" the answer, and so use first principles thinking, or go out and look for the answer where GPT-5.4 does not and simply assumes it knows. Which leads to now, in some cases, the small models far outperforming the big ones. Huge context + training quality seem to be the determining factors now, and neither of those are the strengths of SOTA models. If this continues ...)

While I agree this is a training problem, it is not a solvable one. ML models learn from examples. This is even true for their newest tricks like GRPO. They cannot train against things humans don't yet know.

And that's great, but you're forever locked at the peak of what you can be taught in widely available courses (which they download without paying) (even that is best case scenario: it assumes your ability to distinguish bullshit from reality somehow becomes perfect during training, or even before). The only way to exceed peak human performance is to start experimenting with math, physics, chemistry, even humans, yourself. And that has, even for humans, a massively higher cost than learning from examples, or from a course.

The reason they don't go further is the worst possible reason: the cost. It requires a 100x increase in training expense. Think of it like this: to exceed SOTA in physics or chemistry, training the next version of ChatGPT requires a particle accelerator, and a chemistry laboratory. This cannot be bypassed. Oh and not just any particle accelerator, right? A better one than the best currently existing one. Same for Chemistry labs. Same for ... So 100x is conservative.

But without doing it, ML models (LLM or otherwise) are forever limited at the level an army of first year university students achieve, ON AVERAGE. Maybe they can make that 2nd or even 4th year, at the end of the curve. But that's the limit. Phd level is the level you have to come up with new discoveries, and that ... just isn't possible with current training, even at the end of the improvement curve.

And ... is there budget to increase training cost another 100x? No ... there isn't. Not even with this totally absurd level of investment there isn't. And if small models keep this up, there's no way the investment is even remotely worth it.


Gemini 3.0 wasn’t just a marginal improvement over 2.5.

And if you take that out: 1. All of those releases happened literally in the last 3-ish months. 2. They’re all intentionally marginal releases, hence the minor version bumps instead of major versions.


Equally marginal?


No, the anthropic releases have felt marginally negative


Because the premise that the singularity is just around the corner is far less likely than the premise that artificial intelligence is a lot harder than most people think it is and we're not that close.

Especially because the companies telling us the first premise is true are the companies which need investors to prop up their business.

I mean, it is possible the first premise is true, but the absolutely bonkers credulity in it really mystifies me. It is an incredibly unlikely thing to be true and we should be demanding quite extraordinary evidence to back it up. But based on some neat tricks by current LLMs, some people are all in.


> > And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon?

> Because the premise that the singularity is just around the corner is far less likely than the premise that artificial intelligence is a lot harder than most people think it is and we're not that close.

I see no claim that the singularity is around the corner, so I'm not sure your reply meets the comment that you're replying to.

It seems overwhelmingly likely that AI will be significantly more capable 6 months from now than it is now. Even if there's little progress in the models, just the rate at which tooling is moving will make a big difference. And models still seem to be improving, so I'd be a little surprised if we hit a model brick wall.


It’s more of a guess if you don’t know about things like scaling laws and RL with verification. The onus of “we’re going to saturate” anytime soon is on that claim because every measurement points to that not being true.


But… RL doesn’t scale that well. It’s not the silver bullet you think it is.


Yeah. People (Gary Marcus) have been claiming that AI will hit a wall or is hitting a wall or already has hit a wall since 2023, basically. And yet every time they proclaim that the AI industry found new ways of training their AI's, new ways of integrating them with external tools and feedback loops, new architectures and more to keep the exponential growing. And sure enough if you look at literally every attempt to objectively rate and verify the capability of these models, including things like the METR time horizon autonomy index or the artificial analysis intelligence index, you see exponential or even greater than exponential growth, continuing smoothly through each of the points people claimed that it would begin to slow down, with no sinus slowing down or stopping at all. So yeah, I think at some point the onus has to lie on the ones that are making the claim that keeps being wrong and the continues to be wrong and it completely goes against the current tangent of the curve that we're seeing in all objective metrics. Especially when they can't give specific new reasons for progress to stop beyond the ones they gave last time. It didn't stop and really can't give specific reasons at all besides vague general points about stochastic parrots and S curves.

I really have to highlight the S-curve nonsense because, like, yes, I think this technology's improvement will follow an S-curve. It's absurd to think that it will just follow an exponential up towards infinity forever because nothing in the world really works like that. However, like everyone else in this thread is saying, we have no idea where on the S-curve we actually are, and it's impossible to know until it's already slowed down. So really all appeals to the S curve do are as function as a sort of non-specific, unfalsifiable prophecy that someday it will slow down, which doesn't really tell us anything useful, and also frees the person referencing the S curve from ever actually having to worry about being wrong. Just like the Singularity people, the slowdown of the S curve is always near. This is actually a known and well-established tactic of religions and other people that want to make prophecies without having to worry about turning out to be wrong — unfalseifiable vague prophecies with no actual timeline, and thus no clear import to the present so that they can never be shown to be wrong.


He said "will stop anytime soon". He didn't say forever.


Which still makes no sense. There is the same chance we are flatlining now as that we are flatlining in e.g. 3 years or 5 years.


In what sense are the models flatlining?


In the sense that the incremental improvements in capabilities that we've been seeing in recent models seem to taking exponentially growing amounts of compute to achieve.


But they don't?

Mythos is a 10T model. Opus is a 5T model.

That's not an exponentially growing amount of compute but it is achieving exponential improvements (eg from Mozilla: https://blog.mozilla.org/en/privacy-security/ai-security-zer... )


> but it is achieving exponential improvements

“Exponential” used here is pure hyperbole. Can you justify it?


Compute doesn't necessarily linerarly follow parameters. And with how many active parameters Mythos vs Opus gets its effectivenes from? Is it 1x or 2x? We don't know. We don't even know the parameters (it's more of rumor than confirmed 10T iirc).

But even more so, who said the improvements are "exponential"? Mozilla's single metric, that doesn't even prove anything of the sort?


I know parameters don’t translate directly like that (and that linear and exponential aren’t the only types of growth) but a doubling as a go-to example of “not exponential growth” is pretty funny.


Wasn't 4.6 Sonnet a 1T model?

Parameters and compute are quite the same thing, but going from 1T to 5T to 10T is quite a ramp up.


where the heck did you get those parameter numbers from?


Sonnet and Opus are from Elon Musk (given the people he's hired it seems likely it is approximately true). Mythos is quite widely spoken about.


> Mythos

Ah yes, the marketing model that's ostensibly so powerful us mere mortals aren't allowed to use it. It's certainly led to exponential hype and speculation.


There are advancements that do not follow s curves - consider for instance total data transmitted over all networks, or financial derivatives volumes.

I think a better question for AI is “is it more like a network effect, liquidity effect, or a biological/physical effect”?


Those are measuring the utility of a technological advancement by looking at usage, not the pace of advancement of said technology.


Yes. But quantity has a quality all its own, as they say — derivatives have gone through at least a few step functions where they have become more important and more useful as their usage grows. I’d call that advancement.

Maybe just to be clear I think that kneejerk “I hate this AI trend, and prefer to believe this will end soon, all exponential growth ends eventually” is intellectually lazy, and dangerous for younger engineers/hackers, a group I hope can benefit from being on HN.

Bitcoin mining went through something like 13 10x growth periods, last I ran the numbers a few years ago. There are physical processes that do have very extended periods of doubling, and there are digital and financial processes that don’t show any signs of doing anything but continuing to keep growing over their multidecade lives. So, like I said, it’s worth thinking carefully, and risk mitigation for things like mental health, career decisions and investment decisions indicates we should be cautious assessing new dynamics.


>There are advancements that do not follow s curves - consider for instance total data transmitted over all networks, or financial derivatives volumes

Or Roman trade volume before the Fall of Rome.

Not to mention what you describe is not technological improvement but increase in data or money flows, not the same.


Sic transit gloria - obviously.

But I don’t that think it’s quite so obvious that model quality / growth / usefulness is definitively and obviously not more like data or money flows than it is like some other process.


Total volume of usage is not an advancement, it’s orthogonal.


Indeed, and it's more linked with market penetration than technological advancement. It's like evaluating airplane technology by "total miles flown".


This could be right for the current architecture of LLMs, but you can come up with specialized large language models that can more efficiently use tokens for a specific subset of problems by encoding the information differently (https://www.nature.com/articles/d41586-024-03214-7).

So if instead of text we come up with a different representation for mathematical or physical problems, that could both improve the quality of the output while reducing the amount of transformers needed for decoding and encoding IO and for internal reasoning.

There are also difference inference methods, like autoregressive and diffusion, and maybe others we haven't discovered yet.

You combine those variables, along with the internal disposition of layers, parameter size and the actual dataset, and you have such a large search space for different models that no one can reliably tell if LLM performance is going to flatline or continue to improve exponentially.


> So if instead of text we come up with a different representation for mathematical or physical problems, that could both improve

But then, wouldn't we first have to translate all of our current math and physics knowledge into that new representation in order to be able to train a model on it? Looks like a tremendous amount of work to me.


Yes, but by then you already have general LLMs capable of helping with the work. And even if you didn't, if that's what it would take to advance research in these fields, that would be a justifiable effort.


>This could be right for the current architecture of LLMs, but you can come up with specialized large language models that can more efficiently use tokens for a specific subset of problems by encoding the information differently.

That's precisely what happens on the bad side of a S curve.


Progress don't stop however, and the S curve resets, because then you are optimizing a new architecture.


I read an experiment someone wanted to try where they used pre-1900 content and tried to get relativity. Another version would be train an LLM on school curriculum up until calculus and see if it can invent calculus. Where we are on the curve depends on if it's remixing known things or genuinely inventing things.

From the article,

> ...LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it. Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments...


What people miss is that AI isn't one S curve, each capability we try to bake into a model has its own S curve. Model progress might not impact some capabilities at all, but other capabilities might get totally overhauled.


Software and hardware have no limits. Theoretically would could bozons for computations and have the same amount of computation available on one cm3 of the current total computation in the entire world. Same with software. Never there was a stop on new algorithms. With LLMs there are so many parts that will get better and are not very far fetched.


> Software and hardware have no limits.

Yeah, if time is infinite, R&D imagination is infinite, energy is infinite and material resources are infinite. Easy.


Assuming it’ll stop soon is to wager that we’re at a very specific point on the curve.

If it’s anyone’s guess then we’re much more likely to be left of that, unless you argue we’re already on the flat side.


It can be S curve (and it almost surely is), but on every chart you can plot, you don't see even of an inkling of the bend yet.


you can tell where on the sigmoid we're currently sitting? frontier lab folks can't - chapeau bas good sir


> frontier lab folks can't

Do you have a source for this that isn't marketing spiel? There's a fiscal incentive to lie about scaling research.


This is FUD and extremely wrong. None of the advancements have followed an S curve. This time IS different and it should be obvious to you at this point.


What the fuck does that have to do with “soon”?


There are many indications that model progress is slowing down, so that is not entirely accurate.


Please be specific because outside of anecdotal blog posts by people who don’t know what they’re talking about it’s not true. Look at scaling laws, composite benchmarks from the epoch capability index, nothing at all suggests “model progress is slowing down”


Which indications are that?


The cost factors on the new models compared to the old models.


Qwen3.6 9B is as good as GPT-4o and runs on my M2 MacBook Air. Models are getting stronger and less costly at the same time, but these are somewhat separate branches of research. Frontier labs are spending more because they are still getting marginal returns and there is more capacity to spend than there was a year ago.


Qwen 3.6 9B doesn't exist.

If you meant 3.5 9B and you truly believe it's as good as 4o then I can only assume you have a very basic use case.


You are right, I was mistaken about the version. I evaluated it in general chat assistant prompts plucked from my history across a range of topics but did not use it for coding - there was never a time when I thought 4o was “good enough” for agentic coding.


You are mixing cost and progress. It’s not because it’s more and more expensive that progress is slowing down by itself.


They are intrinsically linked beyond a certain point. If we're making progress but costs are spiraling exponentially then it stands to reason that we will soon reach a point where we can no longer afford the increasing costs and thus progress will slow.

(barring some breakthrough that reduces costs, which of course may happen, but for which recent model improvements are not strong evidence of)


Cost for a specific level of performance decreases 10x per year, this has been a pretty consistent property for awhile now.


I guess within the domain of AI, a pertinent question would be: "do I want to use anything but the best?" The errors older models give being directly analogous to being stupider in my eyes.


Depends — many tasks in various pipelines have a reasonable Pareto frontier and diminishing returns after a certain level of performance. You may just have a high budget constraint (say like YouTube computing ASR subtitles; they are not going to be using the best ASR models because it’s expensive). If it’s myself, with a coding agent, I’m going to get the best thing I can afford.


Investment dollars.


Source for that claim?


Nobody is releasing NEW models


…not only is this not true but it also doesn’t matter. Why would this indicate performance saturating?


The standard networking connection has been called “Ethernet” for more than thirty years, so networking has stagnated, right?


If higher bandwidth networking consisted primarily running more and more ethernet lines in parallel, you would most certainly agree that "networking has stagnated".

"Reasoning" and now "Agentic" AI systems are not some fundamental improvement on LLMs, they're just running roughly the same prior-gen LLMS, multiple times.

Hence the conclusion that LLM improvement has slowed down, if not stagnated entirely, and that we should not expect the improvements of switching to these "reasoning" systems to keep happening.


From TFA:

“ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove”


You misunderstand. I'm not saying that Reasoning/Agentic systems aren't better.

I'm saying they're not an advancement in the tech in the way GPT 1 through 3 were. They're a different kind of improvement.

And as such the rate improvement cannot just be extrapolated into the future.


GPT1 through GPT3 advancement were exactly like using more Ethernet cables in parallel.

All interesting conceptual breakthroughs came after GPT3: RL and reasoning being the main ones.


What constitutes a NEW model for the purposes of calculating progress?


What? DeepSeekV3 just came out and is incredible for the price. Mythos is also half-released.


Until you or I can actually use Mythos in Claude without an nda or other strings attached, Mythos is not released and is just an effective marketing tool for Anthropic.


At least to me this is a pretty sour grapes take. There are all kinds of released products that are expensive or need an NDA. You're just too poor to afford it. But make no mistakes there are governments using this in mass and likely against you.


I think that’s worthy of at least sour grapes, too.


Model progress at spitting out unhallucinated facts is slowing down hard. Model progress at solving hard math challenges/programming tasks doesn't seem to be slowing down that I can tell.


Deep think still makes many many many more mistakes than gpt 5.5 pro on math


Given the capabilities of upcoming LLMs, I suspect that by mid-2027, most competent companies, outside specific niches, will not hire and might fire any non-senior “generative AI vegetarian” software developer.

Note: I agree with others that another term should be used instead of ‘vegetarian’. “LLM vegetarians” do not hold the same moral values as vegetarians.


Believing in the capabilities of _upcoming_ LLMs that you have never actually used shows that you buy into marketing and hype very easily. No one really knows what the future will look like and there's an equally plausible one where post-subsidy token economics become impossible to justify for most use cases.


There are reasons, based on machine learning related theory, that justify the belief.

Economics will likely sort itself out through optimizations, which are also highly plausible.


> There are reasons, based on machine learning related theory, that justify the belief.

Cool, what are those reasons? Links to papers would be greatly appreciated.


From the author’s earlier essay:

“A good way to describe myself is as a generative AI vegetarian. You can find a fuller explanation—and many, many links—at the above essay by Sean Boots, which I agree with almost 100%.”

—-

Given the capabilities of upcoming LLMs, I suspect that by mid-2027, most competent companies, outside specific niches, will not hire and might fire any non-senior “generative AI vegetarian” software developer.


I have actually no idea what you want to say with “generative AI vegetarian.” You mean people who refuse using LLM's?

edit, I see, a new slang:

https://news.ycombinator.com/item?id=47928885


Probably another viral marketing campaign to further pressure those meta employees to have their in office flatulence levels monitored with probes as they are pressured to vibe code more features faster.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: