With today’s technology I’d still want a GPU driver developer guiding the LLM rather than some rando who is out of their element. But cutting down the exploration cycle time and giving the developer massive parallelism (have 10x agents exploring different hypotheses or features) is the real win. We don’t need to skip all the way to slop just to squeak out a little more effort savings.
Never had any issues using it for personal small scale development. I think the premise of OpenRouter as a magical way to fall back on error or select providers based on price/tok-per-sec/uptime isn’t delivered upon. But it’s an excellent way to simplify the experience of hopping between model providers with a single unified billing system. I mostly only use first party model hosts. Theoretically there are alternatives with higher throughput or lower prices, but it’s simplest to not worry about that optimization.
I want them to be sterile and inhuman as much as you do. But I don't draw the line at "I". I'd rather not read through even more awkward English as it tries to work around how all of the training data has something or someone refer to itself.
Depends on the actual audience, I guess. My stated “wisdom” comes in part from
/r/LocalLLaMa, and my impression is that the tasks that users there give their models to try them out lean towards rather simplistic, on the reasoning side.
But there I literally did read “you don’t need anything better than 4 bpw” a bunch of times.
reply