Yeah, it's interesting how much worse they are than the Astra pelicans. I think that reflects a tiny bit of genuine value still left in the benchmark, to be honest.
Tons of value left, especially for open source models. I would say the benchmark is yet to be truly saturated (just look at the legs and seat to see what I am talking about) and I always look forward to seeing them. Thank you!
To me the main upshot of this benchmark is precisely that the pelicans still usually look a bit wonky. It's bizarre, since this definitely has a good solution, but it's in line with my experience that memorization of the training set just... isn't happening very much? As in, whether a model fails or not doesn't have much to do with whether that exact question was likely posed many times before.
I think some additional value would be had by seeing how well it can modify the pelican.
Like, "now facing left", "sitting on the handlebars", or "with green spokes" to see if it can break out of some pretty obvious statistics in the training data!
And, there's always asking for an STL rather than an SVG!
I couldn't help myself. Here is the game, one-shot on claude.ai, Opus 5 Extra. There could certainly be play-ability improvements made with 1 or 2 prompts, but neato for one-shot.
"Create a fun and cool game from this" - and pasted the code from above.
Sure, but technically proficient users often have other things they want to spend their time on, outside of exercising technical proficiency on every single thing.
Heck, I don’t even do my own oil changes anymore despite it being easy. Life gets busy, you know?
Many systems out there have specific exceptions from their VPN policy for iCloud emails and iCloud private relay. Places that would immediately block fastmail because of its alias feature will not block iCloud because too many people use it. Market share is real power. You can ban 0.1% of your customers, you can't ban 30% of your customers.
You're just returning the name of the tool, the rest of the information (description/input schema) is definitely lost. Cut to the LLM making mistakes in calling the tool with incorrect schema or calling the wrong tools altogether, recovering, wasting tokens and cycles.
The format preserves all fields — name, description, and input schema are all there, just encoded with pipes instead of braces and quotes. It's lossless, not a truncation. I should have made that clearer in the post.
> You gotta admit the timing looks very suspicious.
Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?
reply