I'm generally extremely skeptical about a lot of the model hype that show up in comments. Except when there is an extreme mismatch the performance, quirks and quality of these things are difficult to nail down. You wouldn't know that from the comment section of every single release.
I think some of these are excited, eager users always ready to hype up the new thing. The same crowd that previously would constantly push for a rewrite from angular->react->svelt->god knows what. Instead now it is on a 6 week cycle and about models/harnesses.
I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.
Ultimately, I think some of it is legitimate shifts in who's in lead and what is the best. You gotta dig through a lot of crap to get to that, and I don't really know how to to.
Ultimately I'm saying is that I always applied a fair amount of skepticism about what I see in comment sections but these day it is extreme amounts.
> I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.
I've assumed the same as well.
I also assume that many of the companies developing these models engage in benchmaxxing.
At my company we've developed our own internal benchmarks for evaluating LLM models as they become available. The benchmarks are tailored to our particular use cases but the utility and knowledge our benchmarks assess is still fairly universally applicable. I see wide differences between what our internal benchmarks report and what the major benchmarks do.
There was a whole lot of fanfare about how amazing GLM 5.2 was when it was released, but it was pure rubbish on our internal benchmark -- far behind OpenAI, Anthropic, Gemini, DeepSeek, etc. I don't know how to reconcile the fact that GLM 5.2 performed very well on some of the major public benchmarks, but consistently performs so poorly on ours. OpenAI models tend to dominate our internal benchmarks.
It actually is showing in public benchmark if you know how to look for it. For example, in Terminal-Bench 2.1, GLM 5.2 received 78%, while GPT 5.6 Sol received 88%.
Then Terminal-Bench 3.0 came out (where the questions are new), and GPT 5.6 Sol received 34.6%, while GLM 5.2 dropped to a whopping 4.6%.
> OpenAI models tend to dominate our internal benchmarks.
That's odd, since Fable seems to be the leader for the industry. Not cost-effective, but if Anthropic models get dominated by OpenAI in your internal benchmarks, this calls their validity into question. Separately, see the jagged frontier effect. [1]
I've been using GLM 5.2 at my day job (mostly Rust backend work ATM). Nothing that blows away the models from OpenAI and Anthropic, but solidly good enough to get it done. A lot of people have experienced this and the fact that an open weights model can do so is where most of the excitement comes from. Optimizing for benchmarks can only get you so far, and people are quick to criticize models that fall into it (like DeepSeek Pro V4 recently).
I think most of it is bot driven spam by the various labs. It's hard not to notice 3 month old accounts with very strong opinions about various frontier labs and little else.
Really no different than when there was suddenly online personas everywhere hyping up TSLA out of the blue. You can see the same thing going on with the BoringCompany subreddit. Crazy that the botnet master isn't able to convince us that Grok is also the best model. I don't think buying twitter was an accident it was probably just literally covering up the evidence.
IMO this is getting hyped because the 27b version runs on a decent gaming GPU. This is NOT a thread for their largest model, this model will run on a mid-high end gaming PC, which you probably have in your household. Mine is 6 years old and it runs quite well.
I strongly suspect that many of these accounts you think might be bots from the labs are just people who only have one interest.
One thing I have noted a lot more of is that there is comparative fanboying going on. Like "this model has done badly, my favoured competition has a model out soon that will beat this in every way" — comparing a released product to unverifiable hopey claims about an unreleased product.
It's tempting to assume that is bot stuff, but if you've been around any other "hot" technical hobby online (cameras, phones, 3d printers, whatever) you will know it's not. It's just fans aligning into teams, some of them laconic and amusing, some of them overkeen and toxic.
Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?
Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…
Anything from ML pipelines for language specific pruning over a Rust/JS/CSS mix codebase to assistance in motorcycle maintenance and different canvas coatings. Most of my evals build on those requirements and especially past failures, whether in pure information, task execution and coding or tool calling beyond the overfitted mainstream. All stuff derived from actual failures encountered, some still only few models come even close to passing. With such a mix, it just takes a while to get any serious opinion on a model. Doubt anyone can do that in such short time, unless their tasks are so simple that most modern models not only succeed but could themselves accurately rate output. If even Fable still confuses PU or wax coated cotton canvas with a nylon shell, or tells me with a straight face to adjust the valves on a bike that has hydraulic lifters that needs experience for human assessment and the time that comes with it. Anyone with less knowledge either wouldn’t see the mistakes staring them in the face and just go by vibes, any model rating these equally can not tell what is accurate and will just go by the output sounding accurate over being. Gives sometimes very interesting results far different to public benchmarks. Inkling, e.g. is more accurate in not telling you to adjust valves that are simply not adjustable then Fable or Sol, which just tell you to adjust every 5000km. Sometimes even when their reasoning and search includes sections about the fact this is not necessary or possible. The beauty of overfitting and unbalanced training data…
Most of us aren’t deploying these in general purpose use-cases. E.g. I use Qwen mostly for vision in my personal assistant. I have an eval set for that. Pretty much each of my use cases has a pre-computed problem set.
Likewise I have a tool-use eval set and a browser-use eval set. I use frontier models for most interactive tasks that are not home assistant or background agents.
I can imagine someone could build evals for that but I have never done so.
Deepseek v4 Pro or GLM 5.3 for software architecture I feel are only deployed for general, rather doubtful those are for narrow stuff, if we are sticking with this threads mentioned models. For small models, sure, narrowly targeted sets which can be self evaluated are amazingly valuable, but I feel beyond 500B we are in a different dimension. Rating any model the size of GLM or V4 Pro in hours I doubt is done beyond pure vibes.
Dude, GLM-5.3 released _today_.
The phrasing "I've settled on" is incorrect for this context.