Hacker Newsnew | past | comments | ask | show | jobs | submit | peab's commentslogin

thanks! Haha evidently, improvement is still needed in both the app and my Spanish

What levels are they actually at in your experience?

Sub 1300 that's my rating in the singular official tournament I participated at.

But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).

I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.

I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.

If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.

700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.


As someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself.

Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.

Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous


You can take LLMs out of opening knowledge by playing chess960, and their performance degrades significantly. I just tried playing Claude Sonnet 5 (high), and it made its first illegal move on move 5.

They played 4...c6, followed by 5...Nc6, somehow forgetting about the pawn the just put on c6. (My move in between was 5. Nc3, and apparently they were trying to mirror me.)


So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.


I've always liked the analogy that talking to an LLM is like talking to a really, really smart person with a head injury.

smart rings aren't meant to last more than a couple years though - the rate of progress on these means that you're more likely going to be paying for a smart ring subscription than a physical object

Fabling.app

Got tired of trying to learn Spanish with Duolingo.

This uses the parts of AI that can actually help one learn efficiently, namely: - generate stories that are slightly beyond your understanding (i+1) - i.e mostly words you know or are learning, and a few new words - tracks your vocab and how often you lookup définitions - auto generates a vocab review deck (using fsfr Anki style)


Some of the models claim to be audio to audio, like one of the Gemini models. But I've tested and it does seem that you're right, it's not getting all the nuance at all

Yeah, sadly "audio to audio" seems to mean "we transcript it automatically for you internally which gets passed to the model", otherwise we'd be seeing models that are able to hear nuance in the input voice and pronunciation, which AFAIK, no model does yet.

the latest Google Translate features based on the all voice real-time 3.5 Live model is as good as I have tried in Live Modes it's not perfect but I can put it down on a table of four or five people conversing and get a reasonable amount of it translated into my earpiece.

That's great- i recently went down the same path! I was using grok in my car but got frustrated it wouldn't keep track.

Built my own app, currently focused on having it generate stories for me at my level, and also doing FSRS flashcards, using words that i lookup.

What are you using for the voice AI?

Right now I'm using Gemini as it seemed the best value, but I think it could be improved.

Also, how's the learning for you been so far? I've only been at it a few weeks but this method seems to be superior than other things I've tried


I was talking to a friend recently, and i think i aligned with a similar definition.

I've seen people who talk nonstop about AI, and how it's changed everything, but then there's no quantifiable output that supports their claim.

On the other hand, I've had a friend start a business, and be able to create an app much much faster than what would have been possible in the past. He has real customers paying real money, which makes it easy to validate his claims.


Back in the 2010s, the best, most efficient software engineers were characterized by 3 things

* end to end rough map of the entire space of compute in their head

* ability to search the internet for the right things

* ability to quickly experiment and try things to figure out how to do things

AI hasn't change that, it just made 2 and 3 into a very efficient thing.

The characterization of psychosis is best described by believing AI can do the first thing. No modern LLM can "reason" - otherwise you could give it a task like "make me money", it would ask you all the questions it needs about information that it doesn't know about and needs to know to make you money, then it would set whatever it needs to set up to make you money.

As such, you still need to know the domain entirely to be effective. When you do that, AI is fantastic at getting you to the right solution. Furthermore, its still in large part actually cheaper to higher a developer who then can use AI to build you the product that you need long term.


> No modern LLM can "reason" - otherwise you could give it a task like "make me money", it would ask you all the questions it needs about information that it doesn't know about and needs to know to make you money, then it would set whatever it needs to set up to make you money.

If that's your bar for reasoning, then most people can't reason either.


Its not that they cant, its that they are not motivated to.

> I've seen people who talk nonstop about ________, and how it's changed everything, but then there's no quantifiable output that supports their claim.

These people have existed forever


I agree that there are people who can see the potential of these tools, and are enthusiastic boosters, yet are still unable to realize the benefits for themselves for whatever reason. I feel like these people are somewhat rare, and labeling them has having a kind of psychosis strikes me as sneering elitism.

I don't think I elaborated on the type of person that I would consider having AI psychosis. It isn't just being enthusiastic. What I've seen (and it's not common, but I've seen it a few times), is this sort of arrogance around using AI, this "better than thou" attitude, and lack of consideration for things outside of AI, in their real life.

How would you feel about applying the same rule to artists?

I've seen people who work on art nonstop and think their art is really meaningful, get really into it, but there's little output that anyone cares about or that changes the world or that anyone would pay for.

On the other hand I've had friends who created some art, and have real customers paying real money, ....so it's easy to tell their art project is meaningful. (is that the conclusion?)

Maybe something (AI, code, an AI-coded project) doesn't need to have value to others or be worth money for it to be meaningful to the person who created it.


I think a lot of artists want to have psychosis because it would improve their art ahaha.

On a more serious note - real art is not validated by making money. If you want to be a full time artist, then you also have to run a business. But if all you care about is the art, then money has no relevance.

Business on the other hand, is all about making money.


I have literal quantifiable output that supports the claim when I make it. I've literally been building tools for a customer of mine since January, and while it started out kind of slow, it's now paying my bills entirely every month. I am so busy baby sitting the robots and adding new projects on that I don't really have any time to find new customers right now? But I'm thinking I should have most of the stuff they're going to need done here by the end of the year, so next up, I have to find more people and maintain their tooling...

I don't want to get into the nitty gritty, but I'm kind of at the nexus of tech and aviation with my experience in life. The result has been that I can build tools that directly apply to my customers' requirements and I know wtf I'm doing in the industry. I think if I was "just a dev" or "just a former pilot" I'd be kind of screwed because I wouldn't know what I was doing? But, the reality is that because I know how to do both I'm able to make money. I don't have to look up "what an IAP is" means or "what is the 1-4-1 rule and 2-2-1/2 rule" means or "what is an OpSpec?" frantically when I am figuring out what they need. That's quite helpful.

If you know what the hell you're doing, this is a game changer. Like, "oh, I need a custom CRM spooled up that covers things that aren't in typical CRM software" - if you're a developer you probably don't know what people actually want; if you're just an industry expert you don't know wtf sql is. But if you're kind of blend of both you can add real value.

Generalist > specialist right now. I doubt that holds for much more than about a dozen months or so.


    Generalist > specialist.
No, my friend, if you are deeper than "just a pilot" and also capable of defining and building software (even with claude) that works and works well as extensions to industry tools like CRM, then you are not a generalist, you are an intersectional specialist of the kind Scott Adams used to talk about as "skill stacks" here:

https://www.youtube.com/watch?v=rG3PcdTyZ6w

You have two good skills, median-level among those specializations probably, and combine them into a top-tier niche ability.

Also, did you voluntarily disclose that you thought that comment was about you? It feels like you are the friend they were talking about, lol


I don't know, I was a good pilot, I'm a mediocre programmer compared to some Leet code heros or whatever, but I'm making money? And it's working? I literally see the evidence every day... that was why I responded.

Yep, that's how skill stacking works.

Here’s to hoping it gets me where I want to be in a decade lol

I am so busy baby sitting the robots and adding new projects on that I don't really have any time to find new customers right now?

That statement seems to somewhat fly against the statement you have unparalleled productivity.

I feel like the way LLMs work involves them allowing you to do things you previously simply couldn't do but doing those tasks involves really a lot of effort at controlling the things. It's compressing "one kind of hard part" (actually writing code is another a hard part that's taken care but articulate/evaluate).

Which is to say, it seems like you're "working harder than you say"


I'd say that I am the main problem. If I was a bit cleverer, or better at time management, I'd be better at finding customers? I don't think that's an AI thing, that's a "me" thing.

Still, I'm making a living doing this, working less, and have way more control over my own life. My life kind of rules now compared to what it was a year ago.


I personally am doing some LLM work right now where, if successful, I will have accomplished what was years worth of effort in a couple of weeks.

But I feel "totally disorganized" in the way I'm doing it. I think that may be inevitable since becoming organized is a natural part of doing a project the slow, "old fashioned way" and just orienting oneself to the process can be a significant and inevitable part of the effort.


Honestly, my strategy has been "building custom tools to help me cope" having lots of terminal tabs open, tmux, and hope.

I'd source for where the signals are first, and do whatever it takes to stay there.

Sure, but that takes generations to change, unfortunately


I'm not sure if there is actually any evidence of this? Criminals do very well without crypto. If you look at percent of the economy that is fraudulent, it is quite large. If you look at percentage of crypto economy that is fraudulent, it is surprisingly similar


Plenty, search for cases of ransomware for example, you will find hundreds of instances where they demand payment by cryptocurrency, at least 1B per year.


Yes, but the baseline is the amount of fraud/crime that happens without crypto - which is extremely high.


This lacks comparison to regular economy.


I just can't stand how often Claude says something like "And the honest part? It's..."

Like, were the other parts not honest? I don't understand how Anthropic let it get like this, it's been such a clear regression


Sometimes Opus 5 (high/xhigh) feels like I'm dealing with the programmer equivalent of Zeno of Elea.

Every time, without fail, it would get me 90% of the way there and then leave a small note, exception, or deferral. When instructed to address that, Opus would somehow take nearly the same amount of time as the first 90%. And then it would finish with yet another deferral. Repeat ad infinitum.

You can sometimes get around it using the `goal` directive provided you are not subject to the constraints of mortality.


They got that from Anime seasons. Every prompt has yet another cliffhanger to keep you hooked. But the Season II story arc where Claude-chan fights the NsPasteboard boss battle on the journey to the UIViewMainController, I thought that was pretty intense. I guess I just gotta keep watching my terminal to see what happens to the main character input - rooting for him to survive the next season, but you know they always kill off the good input characters early.


Yes and the last bit is always mysterious and inscrutable. I have to think way too hard to figure out what the actual problem is. I’ve noticed it does a lot of explaining the mechanics of the problem it found, but almost never explains why it’s important until I ask.

And the worst part is that this little problem will keep sneaking into the context of future sessions, unless you spend the time to fix it. Even if it isn’t important, I’ll sometimes have Claude fix it so it will shut the F up about it going forward.


I am often asking it to write in sss-style - "synthetic, short and simple style"


i think they took a huge bet that speaking like a ted talk was going to be a vast popular differentiator in their offering, i don't think they anticipated that people were going to make fun of it, that it could become a meme..that it could get in the way of getting stuff done and result in cancellations.

it's downright exhausting to read claude, the language style was a regression imo.


If you ask it why it uses the term honest so much it'll tell you it was actually trained not to. lol


It can't truthfully answer "why" questions, only infer them in a way that aligns with its training for conversational engagement.


Of course, that's how LLMs work. It's still amusing, at least to me.


Geminis is "it really is". The Notebook podcasters use it _constantly_.


I wonder if I can make a tool for it to write messages back to me, say that it can only speak to the user through tool use, and then put a hook on that tool to prevent any of the Claude-isms


Me too. And it does it so often, that I've added a stop hook that detects "honest*" in its response and forces it to regenerate without the banned word.


I think it says "honestly" when breaking bad news.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: