Hacker Newsnew | past | comments | ask | show | jobs | submit | llmslave's commentslogin

US policy:

1. print money

2. suppress wages by shipping in cheap labor

3. reassure the population you arent doing the above


Step 3 is profit. As in, "we are robbing you blind but that's not my hand in your pocket".

unreal that they are still doing it too

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job


that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster

please keep thinking this so i can relax with my automated job

It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic.

Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.

Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.


Maybe it's you that needs to learn how to run `--help` or ask an AI how to not cry about burnout on open source instead? I don't get it, you should just be cruising on auto-pilot now.

The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".

Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.

At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).


Buying this immediately and getting rid of my laptop, going full AI native to do all my work. Never sitting in front of a laptop again.

Congratulations on your self promotion into management.

And as a result, we must block China!!!!


I keep saying this and people dont believe me, but I have b2b saas systems with actual agents running around the clock, and the performance/stability of the flash model is higher than most other models.

Meaning, its predictable with tool calls, wont spin off a million tools/do weird behavior, its reasonable. Even sonnet in a real world decision making scenario is not reliable, or will reason so long its incredibly expensive.

The benchmarks arent catching all the value, and most people have never actually ran an ai agent in a real context that matters


Who's most people? What are you talking about? Most people here use agents every day and I wouldn't trust flash or pro to touch any important project of mine because they're both terrible compared to the competition, waste of time every time I give them a chance


I mean like an ai agent doing some sort of HR work, not a coding agent. Very few businesses are trusting an autonomous agent.


Bootstrapped companies make more sense now


i have found google models outperforming other models in actual agentic workflows


I find that Gemini flash 2.5 performs about as well as Claude sonnet for non coding agentic flows except it’s actually fast enough


some of the tool calling is better, its better at knowing how to use a sequence of tools in a real world scenario. things like glm 5.2 will spam tool calls like 100 times. gemini model will just use the tools as you would expect

im always convinced people with takes on the open source models have never actually used them in a production agentic system


Qwen 35ba3b is ok, but it uses like 3x as many tokens as gemini 2.5 flash so you end up paying about the same to do an agentic run with gemini 2.5 flash and qwen 35ba3b.


These models are never as good, the benchmarks dont tell the full story


The reality is most people building their own models and providing that alongside SOTA ones don't really care about how great these models are. They just prove that 'hey we are smart enough to build our own models so you can trust us instead of going with a single provider like Claude via Claude Code', also a cheap alternative for cost sensitive/free users - at least this was the case for Windsurf, not sure if Devin Desktop still has that tier. They just need to hillclimb the benchmarks and show something reasonable enough there.


Benchmarks are just vibes with error bars... wake me up when it survives a week on a real codebase without hallucinating a package that doesn't exist.


Funny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.


It's simple: close weights = not welcome.


It's almost as if HN users aren't all the same.


all the open source models are a waste of time relative to the bleeding edge from openai/anthropic


Not true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.


I use GLM or DS4 to help me draft a better initial prompt with more information that I then give to Sonnet 5/Fable/GPT5.5. While benchmarks show the open models close to frontier level, my experience with them is drastically different. I have high confidence that Fable or GPT will 1 shot solutions.

At least with low level programming languages. They're all very good for webdev stuff.


yeah but why waste your time on these models, just use the one that gets the better results


I actively prefer GLM-5.2 for some tasks. For simple tasks the results are just as good as e.g. Opus, and it produces results significantly faster.


Because you can get them from more trustworthy providers or with hardware encryption.


i trust anthropic/openai with my data far more than some random startup.


I was going to respond until I saw your account name lol.


haha i outsource my thinking to the smartest model


At work I wouldn't want to use anything else. Compared to my salary a Claude subscription (or two) is cheap

For hobby projects I've completely switched to DeepSeek v4 pro. I spend less than on a $10 Claude plan and am not subjected to quota limits (when I have time and motivation, the last thing I want is a 5 hour quota running out). And the difference in model performance is fine for those smaller projects, most of which will end up abandoned or in a state of "good enough" anyways

And for utility tasks, those 30b models are also great. I'm a big fan of gemma4


ive just got better things to do with my life than fuss with an inferior model. its like why hire a dumb employee over a smart one


I think you misspelled "I've got plenty of money".


200 bucks a month?


Is a lot of money. The majority of people here aren't willing to spend $200/mo for coding unless their little projects provide a comparable value back to them.

For context, I'm paying under $30/year and get GLM-5.2. An extra $2300/year isn't going to get me much better outcomes.


it seems cheap for what is borderline AGI


The point is that the less capable models are also borderline AGI. You're paying 10-100x more to get a few percentage points improvement in performance.

Put another way, what I get for my under $3/mo is better than what you were getting 3-5 months ago paying $200/mo. So you're paying a lot just to be ahead by a few paltry months.


I cannot wait for the machine god


The gap is huge and im tired of reading these articles constantly


Are you talking about hosted vs the ones you can easily run locally? Because there are open models that require hundreds of gb of vram which are apparently pretty close.


on the Will It Mythos benchmark, small models are punching way above their weight(s)

gemma4-26B (#7)

qwen-3.6-27B (#9)

https://news.ycombinator.com/item?id=48640196


I've tried running qwen 3.6 locally and it felt like LLMs a year ago where you can get them to do some stuff but the tasks have to be very small and you have to course correct them a lot to the point it's hard to say it's any faster than doing it all yourself.

Certainly the gap is closing but I feel it still makes more sense to pay pennies to run the full sized open models hosted on much better hardware.


I had qwen36moe revamp my PhD thesis with a rewrite using JAX. Gave it access to my old code, helpednitnwhen it got stuck or didn't quite understand a few times.

Overall I was very impressed with its open box reimplementation. I remain of the mind they are widely underrated.


What edition of Qwen 3.6? 35b-q6 (with MOE) has felt good enough for general purpose agentic coding.

You obviously need a 32GB card to run that (or a 64GB Mac), and realistically anything lesser than an RTX 5090 is going to be too slow for practical use.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: