Cap the number, sort by salaries (perhaps by coarse domains of skills and not very narrow . Nurses vs Doctors. But s/w engineers are s/w engineers).
Testing abilities is a gamable proxy.
How much a company is willing to spend on the person is the best barometer, if we stay true to the intent of the visa (scarce talent)
They cannot yet outmatch humans in critical thinking. Heck, they cannot outmatch a mediocre person like me yet. How many times have we awen "You are right...." Even on sota models?
They are great tools for research and tasks as of now.
I think they can match a human in critical thinking, but they can't match a human in coherence and meta-thinking. You can prompt an LLM to critically think of any idea you want and it's often able to do so, but it's rarely going to do that by itself.
Yep. And the difference is clear as da for anyone using them both. And in spite of that advantage, OAI is now trying out ads. I can only imagine that even they are getting constrained to compute and are trying to find other ways to plug it
In a world where AI advancement depended only on human ingenuity this would make sense. In that world each political power block would be in an existential race for AI supremacy. In our world compute is the limiting resource. Since the US can control who gets compute, the US already has a defacto supremacy so far as frontier model development. Now if it comes about via human (with AI assist?) ingenuity that compute is no longer a restraint, then the situation is much more dire.
When corporations are involved, it is always a good bet to err towards cynisim.
From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads.
So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.
Because you basically get a discount to use claude code via subscription when using an anthropic model, compared to what you pay via api billing with another harness
Personally that's actually another good reason to boycott Anthropic: beside the fact I perceive their models as (at best) marginally better than the ones I'm used to (Z.ai glm-5.3-flash, DeepSeek Flash v4.1), they even force me to use their bloated harness. They are not even open weights and iirc they're even encrypting chain of thoughts now? Litterally, from my perspective there seems to be no reason whatsoever to choose any of the leading US providers, they're not even competing on price.
I too am using GLM-5.3-flash in Pi and I've yet to encounter a scenario it couldn't handle. And the pricing is just incredible, I've handed it a previously unseen codebase, asked it to analyse it and build a new feature, came back after it had done so and the API cost was a fraction of a cent. It's $0.5/1M output tokens on OpenRouter.
If I really need to, I can escalate a task to Opus at $25/1M, and the results are good, but not 5000% as good.
Indeed! It's ludicrous how much I can still squeze out of a 9USD/month lite sub with Z.ai, it's beyond me how these US LLM providers are still managing to keep their evaluations so high... they seem to have the highest prices for the poorest UX, e.g. security gates which don't seem to benefit anyone (see HuggingFace falling back to GLM-5.x for troubleshooting OpenAI attack), less visibility in the name of anti-distillation protectionism, no harness use flexibility to protect their walled garden, etc.
Doesn't look like pi.dev has anything like auto-mode. They tout running in a container, which you should do regardless, but the scope of soundness that has as a security plan is limited. If some information is in the container, and there's any way for it to get out, eventually it will. The scope of usage that can be covered without that being a problem is leaves a lot uncovered.
You are incorrectly cynical. They are telling you things are bad, and because you refuse to countenance they could be worse, you assume they must be better to comply with your mandate to disbelieve.
A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
You're asserting correctness with no facts to offer of your own, just speculation and your own biased assumptions.
What if consolidating AI into a highly regulated cartel, with no chance of upstart competition ruining their position, is the scenario that leads to the worst possible outcome?
Is it really hard for you to imagine that enshrining a cartel and wedding the government to it, centralizing AI even more than it already is, is actually the path that leads to the doomsday scenario you imagine?
The cynicism is about motivations and not that they are inherently not bad. Perhaps they are as bad as they claim. Or perhaps they're worse. All we have a couple of run of the mill breach examples and some people inside the talking about how dangerous it is. Yes, they are far more qualified than I am (or most people here), and perhaps there is a grain of truth. It is the motivation - and it is always money with corporations.
It is always money — but it isn’t always only money. They are not asking for anything that will prevent them from making money in the future, but they are asking for help stopping the runaway train they’re on. These are compatible requests.
Yes. This one is a major issue. It keeps track of various changes it made in the same session and tries to be backward compatible. Have to repeatedly tell not to be backward compatible.
Things that are automatic for humans and aren't even consciously registered, have to be explicitly stated and even fought for, with LLMs.
Well, I prefer the polars version. And now if I want to reuse the CTEs elsewhere, I have to reach out to another tool like DBT or hand roll something to do string manipulation.
Testing abilities is a gamable proxy. How much a company is willing to spend on the person is the best barometer, if we stay true to the intent of the visa (scarce talent)
reply