Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Smaller models are overconfident and have a hard time to self-correct.

If it’s stuck, usually that’s it.

Bigger models “understand” better, both the prompt and the contents. If you will try to read a paper together with a smaller model, the difference is immediately obvious.

Bigger models will “forget” and drift much less.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: