Sub 1300 that's my rating in the singular official tournament I participated at.
But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).
I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.
I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.
If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.
700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.
Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.
Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
You can take LLMs out of opening knowledge by playing chess960, and their performance degrades significantly. I just tried playing Claude Sonnet 5 (high), and it made its first illegal move on move 5.
They played 4...c6, followed by 5...Nc6, somehow forgetting about the pawn the just put on c6. (My move in between was 5. Nc3, and apparently they were trying to mirror me.)
So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.
In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.
smart rings aren't meant to last more than a couple years though - the rate of progress on these means that you're more likely going to be paying for a smart ring subscription than a physical object
Got tired of trying to learn Spanish with Duolingo.
This uses the parts of AI that can actually help one learn efficiently, namely:
- generate stories that are slightly beyond your understanding (i+1) - i.e mostly words you know or are learning, and a few new words
- tracks your vocab and how often you lookup définitions
- auto generates a vocab review deck (using fsfr Anki style)
Some of the models claim to be audio to audio, like one of the Gemini models. But I've tested and it does seem that you're right, it's not getting all the nuance at all
Yeah, sadly "audio to audio" seems to mean "we transcript it automatically for you internally which gets passed to the model", otherwise we'd be seeing models that are able to hear nuance in the input voice and pronunciation, which AFAIK, no model does yet.
the latest Google Translate features based on the all voice real-time 3.5 Live model is as good as I have tried in Live Modes it's not perfect but I can put it down on a table of four or five people conversing and get a reasonable amount of it translated into my earpiece.
I was talking to a friend recently, and i think i aligned with a similar definition.
I've seen people who talk nonstop about AI, and how it's changed everything, but then there's no quantifiable output that supports their claim.
On the other hand, I've had a friend start a business, and be able to create an app much much faster than what would have been possible in the past. He has real customers paying real money, which makes it easy to validate his claims.
Back in the 2010s, the best, most efficient software engineers were characterized by 3 things
* end to end rough map of the entire space of compute in their head
* ability to search the internet for the right things
* ability to quickly experiment and try things to figure out how to do things
AI hasn't change that, it just made 2 and 3 into a very efficient thing.
The characterization of psychosis is best described by believing AI can do the first thing. No modern LLM can "reason" - otherwise you could give it a task like "make me money", it would ask you all the questions it needs about information that it doesn't know about and needs to know to make you money, then it would set whatever it needs to set up to make you money.
As such, you still need to know the domain entirely to be effective. When you do that, AI is fantastic at getting you to the right solution. Furthermore, its still in large part actually cheaper to higher a developer who then can use AI to build you the product that you need long term.
> No modern LLM can "reason" - otherwise you could give it a task like "make me money", it would ask you all the questions it needs about information that it doesn't know about and needs to know to make you money, then it would set whatever it needs to set up to make you money.
If that's your bar for reasoning, then most people can't reason either.
> I've seen people who talk nonstop about ________, and how it's changed everything, but then there's no quantifiable output that supports their claim.
I agree that there are people who can see the potential of these tools, and are enthusiastic boosters, yet are still unable to realize the benefits for themselves for whatever reason. I feel like these people are somewhat rare, and labeling them has having a kind of psychosis strikes me as sneering elitism.
I don't think I elaborated on the type of person that I would consider having AI psychosis. It isn't just being enthusiastic. What I've seen (and it's not common, but I've seen it a few times), is this sort of arrogance around using AI, this "better than thou" attitude, and lack of consideration for things outside of AI, in their real life.
How would you feel about applying the same rule to artists?
I've seen people who work on art nonstop and think their art is really meaningful, get really into it, but there's little output that anyone cares about or that changes the world or that anyone would pay for.
On the other hand I've had friends who created some art, and have real customers paying real money, ....so it's easy to tell their art project is meaningful. (is that the conclusion?)
Maybe something (AI, code, an AI-coded project) doesn't need to have value to others or be worth money for it to be meaningful to the person who created it.
I think a lot of artists want to have psychosis because it would improve their art ahaha.
On a more serious note - real art is not validated by making money. If you want to be a full time artist, then you also have to run a business. But if all you care about is the art, then money has no relevance.
Business on the other hand, is all about making money.
I have literal quantifiable output that supports the claim when I make it. I've literally been building tools for a customer of mine since January, and while it started out kind of slow, it's now paying my bills entirely every month. I am so busy baby sitting the robots and adding new projects on that I don't really have any time to find new customers right now? But I'm thinking I should have most of the stuff they're going to need done here by the end of the year, so next up, I have to find more people and maintain their tooling...
I don't want to get into the nitty gritty, but I'm kind of at the nexus of tech and aviation with my experience in life. The result has been that I can build tools that directly apply to my customers' requirements and I know wtf I'm doing in the industry. I think if I was "just a dev" or "just a former pilot" I'd be kind of screwed because I wouldn't know what I was doing? But, the reality is that because I know how to do both I'm able to make money. I don't have to look up "what an IAP is" means or "what is the 1-4-1 rule and 2-2-1/2 rule" means or "what is an OpSpec?" frantically when I am figuring out what they need. That's quite helpful.
If you know what the hell you're doing, this is a game changer. Like, "oh, I need a custom CRM spooled up that covers things that aren't in typical CRM software" - if you're a developer you probably don't know what people actually want; if you're just an industry expert you don't know wtf sql is. But if you're kind of blend of both you can add real value.
Generalist > specialist right now. I doubt that holds for much more than about a dozen months or so.
No, my friend, if you are deeper than "just a pilot" and also capable of defining and building software (even with claude) that works and works well as extensions to industry tools like CRM, then you are not a generalist, you are an intersectional specialist of the kind Scott Adams used to talk about as "skill stacks" here:
I don't know, I was a good pilot, I'm a mediocre programmer compared to some Leet code heros or whatever, but I'm making money? And it's working? I literally see the evidence every day... that was why I responded.
I am so busy baby sitting the robots and adding new projects on that I don't really have any time to find new customers right now?
That statement seems to somewhat fly against the statement you have unparalleled productivity.
I feel like the way LLMs work involves them allowing you to do things you previously simply couldn't do but doing those tasks involves really a lot of effort at controlling the things. It's compressing "one kind of hard part" (actually writing code is another a hard part that's taken care but articulate/evaluate).
Which is to say, it seems like you're "working harder than you say"
I'd say that I am the main problem. If I was a bit cleverer, or better at time management, I'd be better at finding customers? I don't think that's an AI thing, that's a "me" thing.
Still, I'm making a living doing this, working less, and have way more control over my own life. My life kind of rules now compared to what it was a year ago.
I personally am doing some LLM work right now where, if successful, I will have accomplished what was years worth of effort in a couple of weeks.
But I feel "totally disorganized" in the way I'm doing it. I think that may be inevitable since becoming organized is a natural part of doing a project the slow, "old fashioned way" and just orienting oneself to the process can be a significant and inevitable part of the effort.
I'm not sure if there is actually any evidence of this? Criminals do very well without crypto. If you look at percent of the economy that is fraudulent, it is quite large. If you look at percentage of crypto economy that is fraudulent, it is surprisingly similar
Plenty, search for cases of ransomware for example, you will find hundreds of instances where they demand payment by cryptocurrency, at least 1B per year.
Sometimes Opus 5 (high/xhigh) feels like I'm dealing with the programmer equivalent of Zeno of Elea.
Every time, without fail, it would get me 90% of the way there and then leave a small note, exception, or deferral. When instructed to address that, Opus would somehow take nearly the same amount of time as the first 90%. And then it would finish with yet another deferral. Repeat ad infinitum.
You can sometimes get around it using the `goal` directive provided you are not subject to the constraints of mortality.
They got that from Anime seasons. Every prompt has yet another cliffhanger to keep you hooked. But the Season II story arc where Claude-chan fights the NsPasteboard boss battle on the journey to the UIViewMainController, I thought that was pretty intense. I guess I just gotta keep watching my terminal to see what happens to the main character input - rooting for him to survive the next season, but you know they always kill off the good input characters early.
Yes and the last bit is always mysterious and inscrutable. I have to think way too hard to figure out what the actual problem is. I’ve noticed it does a lot of explaining the mechanics of the problem it found, but almost never explains why it’s important until I ask.
And the worst part is that this little problem will keep sneaking into the context of future sessions, unless you spend the time to fix it. Even if it isn’t important, I’ll sometimes have Claude fix it so it will shut the F up about it going forward.
i think they took a huge bet that speaking like a ted talk was going to be a vast popular differentiator in their offering, i don't think they anticipated that people were going to make fun of it, that it could become a meme..that it could get in the way of getting stuff done and result in cancellations.
it's downright exhausting to read claude, the language style was a regression imo.
I wonder if I can make a tool for it to write messages back to me, say that it can only speak to the user through tool use, and then put a hook on that tool to prevent any of the Claude-isms
Me too. And it does it so often, that I've added a stop hook that detects "honest*" in its response and forces it to regenerate without the banned word.
reply