Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Is it just me or is 3.6 27B Q8 K XL (Unsloth) holding up better in sustained token/s rate as the context fill increases over time? The token/s rate seems to be much higher for a time period deeper into context than previously seen.

At least as compared to 3.6 27B in the same quantization.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: