> They may want to rethink their reward structure a bit to avoid paying out to people like me.
Maybe, but who is Waymo to decide whether your trip was gratuitous? Perhaps you had brief business in Embarcadero and chose the particular mix of modalities for one reason or another (timing; or had to carry something heavy one way but not the other etc).
"Investors are questioning whether Anthropic can sustain its extraordinary growth as [...] the potential for AI to destroy humanity cloud[s] its blockbuster IPO."
I just finished listening to A Man for All Markets (narrated by Thorp himself!)
Bitcoin derivatives arbitrage doesn't sound like something I'd personally be comfortable getting into, but I'd be curious to hear more about the types of things you're doing (to the extent you're happy to share).
I can see how it can be useful to start with a broad type, e.g. a union, and narrow it down in a block. However, I don't quite get the opposite direction shown in their example (first an int, then a string, then a union).
Same rationale as flow valuing. Some people like values being reassigned, and some people like types being reassigned.
You might be reading too much into the union example. The checker just doesn't know if the middle block ran, so maybe it remained an int, or maybe it became a string.
Would love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).
It's fairly easy to describe a task that is slightly harder than an existing one.
For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"
One thing that jumped out at me is a slower-than-I'd-expect rate of adoption of Opus 5. If you're running Opus 4.8 and choosing not to migrate to 5, what's behind that decision?
Maybe, but who is Waymo to decide whether your trip was gratuitous? Perhaps you had brief business in Embarcadero and chose the particular mix of modalities for one reason or another (timing; or had to carry something heavy one way but not the other etc).
reply