Hacker Newsnew | past | comments | ask | show | jobs | submit | meatmanek's commentslogin

Assuming the f(ciphertext, key) can produce any plain text (of the same length) from a given ciphertext with the right key, and the keys are chosen sufficiently randomly, you are correct.

Classic XOR encryption is like this. You can make a given ciphertext decrypt to anything you want by XORing the ciphertext with the desired plaintext to get the key.

Therefore, just because you've found a way to decode something to plausible-looking text doesn't mean you've found the correct key.


I just want to rant about these Artificial Analysis charts that you see everywhere:

The "most attractive quadrant" is completely meaningless. The whole point of a Pareto curve is that each point on the curve is better than everything else on at least one dimension, and that you can make these comparisons without placing a value judgement on the relative importance of the different metrics. If you make a composite score of the two metrics (any monotonically non-decreasing function, e.g. a weighted sum with non-negative weights), that score will always be maximized by one of the points on the Pareto frontier.

So going by the numbers in the 2nd chart (1st AA chart) from TFA alone:

   - there's no reason one would choose Deepseek V4 Pro 0813 (max) even though it's in the "most attractive quadrant", because GLM-5.3-Flash is both cheaper and scores better.
   - Claude Fable 5.1 (max with fallback) on the top right* could be your most attractive option if you need the best scoring model and don't care about cost, even though it isn't in the "most attractive quadrant"
   - The un-shown model off the left side of the chart could be your most attractive option if you just need lots of cheap tokens and don't care about quality.
(Obviously if you start including other factors in your score that aren't represented on the chart, then you might choose differently.)

* I also dislike the way they place the labels, and that grey line connecting the label to the point is way too subtle.


Sound critique. I'll add that the Artificial Analysis intelligence index is not considered a good metric for intelligence anymore. Most of the benchmarks that it comprises are saturated or considered low signal today.

What is considered a good metric?

Personally, a combination of low tech and semi scientific tests based on what you normally do (redo the same task you used an older model for with a newer one).

So far, Simon Willison's pelican bike "benchmark" is the only one I've found that shows Fable 5.1 beating Opus 5.5. My personal experience has been Opus is unusable on design work it's so terrible. Evaluating whether we should consolidate AWS DMS tasks (Postgres full load and change data capture) into fewer tasks with more tables, Opus 5.5 was factually wrong and needed correction roughly every other turn.

On a "help me find a sandbox solution for agents embedded in a web app to run untrusted code" research project it kept misrepresenting security boundaries and ended up recommending DuckDB which ironically specifically says it does not provide a strong security boundary in its own documentation. GPT 6 (can't remember if it was Sol or Astra) and Fable 5.1 both recommended FaaS like Cloudflare Workers and AWS Lambda which fit fairly well with the requirements.

I switched from Opus to Fable in the session going badly sideways and told it to "Review the previous conversation and come up with a correct comparison table and corrected recommendations grounded in objectivity supported by citations. Do research as necessary to understand the current ecosystem" and that was a full 180 back to coherency...


Would it be better to have a shaded region parallel to the Pareto curve that gets darker away from it that is labeled "better value"?

> This is in Australia where the hole in the ozone layer often results in 10x ‘normal’ UV levels, you can get noticeably burned in a few minutes.

By 10x you mean 10% more, and by ozone layer you mean the fact that the southern hemisphere's summer coincides with Earth's perihelion while the northern hemisphere's summer coincides with aphelion.

https://www.abc.net.au/news/science/2025-02-04/sun-summer-uv...


I expect they just meant UV 10, which is 10x the energy of UV 1.

> and Qwen3-ASR

Is the ASR inference engine open source as well?


Yes, and it is very good one. Leading position on private leaderboard on HF: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard

I meant the Nari inference engine for Qwen3-ASR. I'm aware that Qwen3-ASR is open source, but I don't see a repo under https://github.com/nari-labs for nari-qwen3-asr or similar.

The Huggingface link on https://narilabs.com/product/stt/ links to https://huggingface.co/Qwen/Qwen3-ASR-1.7B , not anything under https://huggingface.co/nari-labs


the qwen3-asr inference repo is not OSSed as of now. we're planning to write a paper or tech report on it as it contains some general techniques for ASR inference.

how do I follow you? I have a small 5090 doing inference all the time and I barely use tts but a lot of asr, mostly whisper, I ported your tech report for tts and implemented some improvements on my whisper inference based on your tech report as well!

would love to talk sometime!


They have a number of demos and examples in their HF space

https://huggingface.co/Qwen/spaces

I saw a local-ai demo (something + gemma), where the person used ASR to get text and gemma to clean it up (like turning "question mark" into a literal "?", bullet points another one). The presenter also showed a gemma only option, that did both in one go, but had a higher WER on average, and even though the formatting statements were handled without a multi-stage pipeline, they preferred the multi-stage overall


I assume it means changing the settings like fan speed or target temperature.


Flagging is for content that breaks the guidelines, not for content you disagree with.


Clickbait is a form of Spam. Flagging is expressly recommended for spam.


I thought the Kagi API was entirely pay-as-you-go.

EDIT: I see I have $5 of "pre-paid API credit". Is this a one-time thing or is that included with my plan?


I also see $5 API credits. I would also like to know if this is part of my $10 subscription each month.

Edit: this is a one-time promotional credit. One must pay for credits thereafter, in addition to their subscription.


I also confirm the $5 you get is only promotional credit.

I tested on my local LLM chat via Open WebUI, 2 requests from that cost $0.02 (so $4.98 left), and after my monthly Kagi renewal an hour ago, it did not reset, I am still on $4.98.

You are probably better off using Kagi Assistant anyways, as you get $10 there on the $10 Professional plan.

For local LLM web searches, to save on cost I will stick with the Brave Search API which gives you $5 free every month.


they recently made some changes from their v0 API to this new v1 with a more complicated model


What quant, what runtime?


AWQ 5bit, oQ5. oMLX.


When processing multiple users in parallel, don't you end up having to load in multiple experts? Not every session is going to use each expert at exactly the same time.


Yeah, that seems suboptimal, unless you're using a much cheaper model to do compaction.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: