Hacker Newsnew | past | comments | ask | show | jobs | submit | saretup's commentslogin

Not that this benchmark is super relevant anymore but these look worse than I expected.

Yeah, it's interesting how much worse they are than the Astra pelicans. I think that reflects a tiny bit of genuine value still left in the benchmark, to be honest.

Tons of value left, especially for open source models. I would say the benchmark is yet to be truly saturated (just look at the legs and seat to see what I am talking about) and I always look forward to seeing them. Thank you!

To me the main upshot of this benchmark is precisely that the pelicans still usually look a bit wonky. It's bizarre, since this definitely has a good solution, but it's in line with my experience that memorization of the training set just... isn't happening very much? As in, whether a model fails or not doesn't have much to do with whether that exact question was likely posed many times before.

I think some additional value would be had by seeing how well it can modify the pelican.

Like, "now facing left", "sitting on the handlebars", or "with green spokes" to see if it can break out of some pretty obvious statistics in the training data!

And, there's always asking for an STL rather than an SVG!


It would be extremely funny if the explosion in SVG generation capability in particular was a result of this benchmark

Now that it's a solved benchmark, can we get the 3d animated version?


I can't waste Fable for play time, but I was curius, so I used plain claude.ai with Opus 5 High. It's pretty cool.

I asked "create a 3d animation", and pasted the Simon’s animation code.

I must admit I used a second prompt to adjust the opening camera angle, but that is all. Originally the camera was above, about 30 degrees.

https://claude.ai/public/artifacts/b37a9ee2-f5bc-4ff9-ae90-a...


I couldn't help myself. Here is the game, one-shot on claude.ai, Opus 5 Extra. There could certainly be play-ability improvements made with 1 or 2 prompts, but neato for one-shot.

"Create a fun and cool game from this" - and pasted the code from above.

https://claude.ai/public/artifacts/e723244b-fa52-4e63-9819-4...

If anyone has the spare Fable tokens, would love to see the difference in the 3D animation and maybe even game.


Jesus. It's not the world's greatest game, but still. What a world we live in now.


That's seriously impressive.


I have to admit, these are not directly comparable to Simon's tests though, as I was using claude.ai, and he is using raw API correct?

I suppose I could use Claude Code, and disable system prompts. That should be the same, right?

If I have spare tokens at the end of the week, I will try real tests.


Are they? Looked correct to me. At high enough speeds even real ones can look rotating either way in videos depending on fps.


It’s an additional layer of wall for their walled garden to keep the technically proficient users that are more likely to hop walls.


Technically proficient users can figure out their own solutions for throwaway emails e.g. aforementioned Fastmail.


Sure, but technically proficient users often have other things they want to spend their time on, outside of exercising technical proficiency on every single thing.

Heck, I don’t even do my own oil changes anymore despite it being easy. Life gets busy, you know?


Many systems out there have specific exceptions from their VPN policy for iCloud emails and iCloud private relay. Places that would immediately block fastmail because of its alias feature will not block iCloud because too many people use it. Market share is real power. You can ban 0.1% of your customers, you can't ban 30% of your customers.


> zero information lost

You're just returning the name of the tool, the rest of the information (description/input schema) is definitely lost. Cut to the LLM making mistakes in calling the tool with incorrect schema or calling the wrong tools altogether, recovering, wasting tokens and cycles.


The format preserves all fields — name, description, and input schema are all there, just encoded with pipes instead of braces and quotes. It's lossless, not a truncation. I should have made that clearer in the post.


Agreed. Are there any open source alternatives for this?


You gotta admit the timing looks very suspicious.


Luna pricing was just cut by 80% https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog post states is a more accurate judge than the previous methodology.


> You gotta admit the timing looks very suspicious.

Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?

That's indeed a bit fishy.


They could just have avoided all of this by not publishing the benchmark until the new methodology update.


Article written by GPT sol. I’ll have my AI agent read it.


Good call. There’s no need to devote time to reading what nobody took the time to write.


> AI = sparkles and rainbow colors — it’s funny when you think about it

Never seen anyone use rainbow emoji to convey AI


Not the emoji, but the text, borders, or other elements for AI features are often in rainbow colours.


Looks pretty blurred


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: