Hacker Newsnew | past | comments | ask | show | jobs | submit | PedroBatista's commentslogin

I sometimes have that feeling too, then ask another LLM to do a code and vulnerability review and OMG: rookie mistakes, over complications and security gaps even a 1st year student would not make regularly.

So.. one more year of untreated bipolar AI psychosis I guess..


I think these kinds of comments really need to say which LLM that is. There's an enormous difference in skill between the frontier ones and say the Google search AI.


Codex Luna, Terra and Sol. Claude Opus, Sonnet and sometime Fable.

They all work, they all are "good", they all are both "smart" and commit incredible basic mistakes a fair amount of times.

Then there's the cost situation..


Also true for human developers.


At least we are at a point where we can have AI review code and reliably find real problems. That alone is incredibly valuable.


This post and comment makes me believe "science" is the new "code" for Anthropic now that the code advantage is mostly gone and lost for OpenAI, ie. they got much better and Claude become significantly worse over these months.


IMO, Codex is worse than Claude with Fable. At least at Rust.

That said, the open source models are not bad and I'm looking forward to more tools and products built on top of them. Code review, security review, etc.

Anthropic needs to change how it treats users though. I'm increasingly put off by Dario, the rug pulling, the lies, and the attempts to regulate open weights. I'm going to bail if this doesn't change. There's plenty enough that's good enough, and those things are hackable and extensible.

If Fable isn't available at subscription price via third party harnesses soon, I'm also going to bail.


The big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, not ours

"Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."


Its "pure capabilities" are definitely worse than Fable, but I find codex has a much more pleasant style and is in comparison much more generous with its limits.


Ugh, people are still saying the Codex limits are more generous. They're not, Claude's are over 2x higher, have been for months! [1] It's just that Claude uses far more tokens, 2-3x is common. Except sometimes GPT will use just as many or even go into a compact loop and then your quota is gone, little headroom for hard tasks.

[1] https://devforth.io/agents-for-code/?sortby=monthly-value And I can confirm the numbers, I subscribe to both and watch the numbers


That website seems to suggest that Opus 5 spends ~57 cents per task, while GPT 5.6 Sol spends ~49 cents per task? That ratio doesn't feel quite right to me. Artificial Analysis says Opus 5 High costs nearly ~3x as much as GPT 5.6 Sol High for a given task: https://artificialanalysis.ai/models/comparisons/claude-opus...


No no our coffee is not more expensive! The serving sizes are just smaller!


Yeah I subscribe to both and watch the numbers too, and it drives me nuts

Who wants to actually watch anyways, rather than worry about it my team just created our own harness that prioritize usage + intelligence and assigns work out (and records token usage..)

https://go7workhorse.com

Still beta, please try it and give me feedback.


I don't get it. It's the same result.


> people are still saying the Codex limits are more generous. They're not

They are if you follow Tibo on the resets.


Well at the end of the day, I can finish more work with the codex limits.


Codex (+Sol) feels a lot more human for sure. Fable 5 is so, so wordy.


It used to be true up to 2w ago, but with the new/reinstated 5h limits I wouldn't be so sure anymore...


that's news to me, I'm still getting weekly limits, no hourly limits.


Claude has a better 5hr limit?


> IMO, Codex is worse than Claude with Fable.

Fable easily trips its safe guards. You can be 95% complete with the plan for it to trip and then lose it all. Anything is better than nothing.


>> Fable easily trips its safe guards.

Maybe it depends on the type of work you do, because for me it almost never happens.

>> You can be 95% complete with the plan for it to trip and then lose it all.

That's... not what happens though. The session will either seamlessly downgrade to another model mid-session, or it will stop with an alert and you can just re-prompt it. It will still have access to the context.


Web apps are where I have this trouble.

Making a web app secure is literally just finding and patching vulnerabilities, instead of finding and exploiting them. You could have the AI "try to make this app secure", find what it patches, and use it for exploits, and the AI can't know if that's what you're trying to do or not. I don't know how you can get around this. I get around it by not using Anthropic products, at present.


Not to endorse OpenAI's particular guardrails, but unless you're doing something groundbreaking, security best practices should be more than enough for web development.


OpenAI is what I use most. Sol 5.6 still rejects a few requests a day when I'm working on web apps, but, overall, it's not too bad. I wish it'd auto-resume and try again, instead of waiting for me to intervene, but it's rare enough that it's not a huge deal.

It probably doesn't help that I'm using frameworkless PHP - I imagine a lot triggers could be avoided if I was using a framework where secure features were baked in.


With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).


When doing basic CRUD apps I can count on fingers the amount of times guard rails haven't tripped and ended the session


Small reminder that the US government rug-pulled Fable, not Dario. Lots of the safety guards that users find annoying/objectionable were the results of negotiations to get the model back online after the US government forced them to take it down.

Maybe Dario should have just "donated" $1M to Trump's inauguration fund like Altman, Meta, Amazon, Microsoft, Tim Cook, Elon, and Google. There's a reason they are the odd man out with this current Administration.


Maybe Dario shouldn’t have tried for regulatory capture. He was constantly on the news talking about how these models are so dangerous and that we need regulation to keep China from releasing open-source models without guardrails.


The US government didn't make the choices to release the worst version of Opus and label it 5.0, and then isolate portions of their subscribers to limited usage of Fable.

They may have been unfairly targeted by the US government, but they are doing more damage to themselves without government help as well.


Fable only being temporarily included in cheaper subscriptions was because anthropic is severely GPU constrained. They still are, and it impacts almost all of those unpopular decisions. They did announce from the beginning it was temporary.


Horrifying excuse, gpu constraint can be used by all of these companies to justify a shit user experience. If the user isn't properly weighed in their priorities, they have their priorities setup wrong.

Their 20$ tier currently isn't serving their best model, and they insulted their users by putting out an ill tested opus 5.0, which is the worst experience ive personally had using a model in probably 2 years(obviously adjusting for expectations at the time of release).


Yes, as a user you pick what works for you. But it is a reality for them that growth has been huge, and GPU manufacturing is bottlenecked.

People were very skeptical about how much investment most companies put into hardware/data centers two years ago, and anthropic was more conservative than OpenAI here, so it's potentially hurting them now.

(Opus is a separate story: it does seem to have improved in coding in my experience, most weirdness seems to be its human communication)


I write a lot of Rust and Lean, Fable 5 is in my experience better at both. Cost/performance is a different story.


Fable has become my go to in Agda as well. It just crunches hard technical tasks!

I find Fable 5 still lacking in library design. But I guess there is no accounting for taste…


This is really it imo. Fable 5 is better then Sol. But Fable is just of the table for anything even remotely long running. Unless you have very deep pockets. And the difference between Fable and Sol is not world shattering if you ask me. I also find codex a ton better than claude.


So when do you use Fable? For difficult singular tasks?


Yes, I use it when Sol Ultra fails to find a solution.


Yes, both of which are domains for which a verifier is readily available.

You can generalize from them to "science".


Don’t forget to mention part of those players are the clients.

“Clients” usually are even worse offenders than these extractive scam companies, mostly because the “clients” are playing with someone else’s money, usually the taxpayer.


Just because a thing is true doesn't mean you can apply it. There are vast numbers of active ingredients discovered or developed that are effective in curing many deceases but there aren't "drugs"/pills with it because we haven't found a delivery mechanism for them to be effective where they should in the human body.

Going back to the subject: Have you ever been clinically depressed or dealt with people in that predicament? Getting up from the bed is already a huge win for that day, let alone exercising and doing it consistently and regularly. Even taking medication regularly can be a hit or miss.

So, yeah the prescription is easy to write and thank god it's real and effective, but that's only ~5% of the work.


Exercise is great on the days I can do it, I love going for a walk. But sometimes when I hear that suggestion - I wonder if those people have ever lain in a totally uncomfortable position in bed for 45 minutes because they felt paralyzed, unable to will their muscles to reposition themselves into something more comfortable. From that perspective, even light exercise seems like walking on the moon; merely a thought, not real.


I appreciate your difficulties but have always wondered at our lack of support for preventative measures as a society here in Australia and possibly worldwide.

Would a personal trainer or paid or not training buddy or something of the sort help your motivation?


It's a good question. I could see it working for some people (accountability, companionship) but at least for me, one of my parents is a certified PT and while she has helped a bit, the problem is just more deeply rooted than that unfortunately, at least for me.

But, again, everyone's experiences are different.


Hear, hear. I've fortunately not had to personally deal with depression, but have been around those who have, and "just go to the gym" is right up there with "have you just tried not being sad?" on the list of unactionable plans.

See also: "Why don't more people with ADHD get treatment?" "Because if they were able to initiate and follow through a complex plan of meeting with care providers, arguing with insurance, and having prescriptions filled, they probably wouldn't need it."


Bingo. It took years to get an ADHD diagnosis after severe issues with productivity and executive function.

I finally get a medication to try, and on the way home it falls out of my pocket. Don’t realize till I’m home and go back and find it… smashed in the street, run over by a car. Lol. So yeah, it’s a real struggle.


That just about broke my heart. I can imagine how awful and hopeless I'd feel in that situation, at least as a first reaction. Great. Just I'll just plan on not planning until I can get the doctor to send a new prescription, which they might not do because FDA.

There's a reason I pick up my refills exactly on time, even if I still have some of the old prescription left. That little overlapping buffer gives me great peace of mind.


The healthcare system in the United States is deeply, horribly broken.

Recently, my wife was going to be out of the country about a week-and-a-half. Her prescription for a medicine to manage arthritis would run out at the very beginning of that time. But insurance company policies refused to allow her to fill it even a little bit early. As a result, her only option was to go without the medicine for several days.

Policies should be set by organizations whose primary goal is to provide effective health solutions, not organizations whose primary goal is to minimize costs in order to make a profit.


> her only option was to go without the medicine

She could've gone to her provider to discuss adverse side effects. Sometimes they will prescribe a new medication that could be filled immediately.

She could've got a pill cutter and cut the tablets in half, or space out the doses, or otherwise take less than the specified dose for a few weeks.

She could've asked the pharmacist about OTC treatments which could be safely substituted during the gap time.

She could've gone to the emergency room or urgent care in the destination country, and tell the provider "I have run out of medication!" This would be covered by the traveller's insurance that you purchased for the trip.


You sound like you would be a great patient advocate. His wife would have been lucky to have someone like you help navigate that situation, but it would be even better if people didn't need to waste their brainpower coming up with these tips and tricks to work around this fucked up adversarial system in the first place.


Or we could have healthcare that doesn't profit off human misery such that his wife could take her full, prescribed medical dose.

The idea of someone cutting pills in half because their insurer is reporting record profits drives me nuts.


The article says that 41% of mental health professionals never recommend exercise. Is it really the case that these 41% have only patients with the severest forms of depression who struggle to get out of bed at all? Or do they have many patients who get out of bed for work and such, and could force themselves to exercise, but don't know that it's worth the hurdle because they haven't been told about the mental health benefits?


I don’t know a single person in my life who thinks exercise isn’t worth it. I think it’s widely known that exercising is one of the best things you can do for your health, but actually making it a habit is extremely hard.


Perhaps I’m atypical, but when my doctor gave me an exercise regimen and told me I must do it in order to treat a joint issue I was having, I found that much easier to maintain than any workout schedule before or since. Knowing it’s good in general is different than confidence it’ll treat some specific problem.


Linus better put his money where mouth is and fire up those AI tokens ASAP.

I don't consider this a tragedy, it's just unearthing the reality.


What does any of this have to do with AI?


There has been a long discussion and "conflict" of opinions regarding the use of AI in kernel dev.

Linus is pro-AI as a tool ( and I mostly agree, there are consequences however that at least need to be talked about ). That's why I was talking about firing up those AI machines and fix those CVEs ( "make no mistakes" ).

But it seems every Linux topic related to security or languages is the ultimate mine field for knee-jerk reactions. :)


Old samurai lost his marbles and his heir to the throne is busy making kingdom ending deals pretending to be an Hollywood mogul.

Life is good, I guess.. Java people should start sweating.


> Petty theft, snatching, pickpockets, scams, etc are relatively uncommon compared to e.g. many popular places in Europe.

Yes, in non-popular places in Europe those are also quite uncommon, even more then in the US on average..

So the lesson here is that those type of crimes are common in tourist heavy places, like.. Times Square in NYC for example.


[flagged]


I don’t think they were making that comparison, rather that touristy cities have more pickpockets, which is obviously true and expected.

You seem to be very sensitive when it comes to anyone that might deign to question the supremacy of the US and very quick to disparage those outside of it.


It can be.

Because most automations never capture the complete scope of the job/task ( not even close ). Just like neurons, if you don't use it you lose it and when the inevitable problems come, nobody knows the why, the how and the what. At that point someone smart would incorporate all those real costs and opportunity loses on the "automating everything" equation. But they usually don't.

Of course automating tasks is a must, but it's very far from being a black and white situation. These dynamics have been happening for centuries by now, nothing new.


<< Because most automations never capture the complete scope of the job/task ( not even close ).

Management likes to think otherwise for a variety of reasons and peons such as myself know it is rarely that simple. Case in point, our team uses Jira, but Jira is not great at.. capturing efforts that are umm.. less code oriented. In fact, there are times when it is genuinely better to leave some details out for considerations that may simply not be part of coding considerations.

In other words, automation assumes proper capture of the entirety of the task, but, I have seldom seen it capturing anything but simple workflows well.


I agree with that, and that’s why it’s better to ask an AI and improve your prompt, instead of hiring a human that will disappear and you will loose all institutional knowledge


The risk of hiring a hunan is easy enough to manage. Ensure that the process is documented and new hires are trained. Any disturbance by someone leavingg is then smoothed out in the long term.

Can you show us how it’s better to use an AI to ensure steadiness over a long period of time?


Ah yes of course, all companies can afford to hire 5 people for one job to smoothen out the loss of information by employee churn


You only have as much institutional knowledge as you're willing to cultivate and be liable for. Many companies already don't give a shit about institutional knowledge, indicated by how little they're willing to invest in keeping a strong team together, caused in-part by long-standing toxic incentive structures.

Ask an AI to solve a problem and it may do that, but if you don't understand why or how it works, or what to do with the information in order to keep it useful for even the medium term, then all you've done is taken away the opportunity for someone else to be responsible for something you shouldn't be.

It's not necessarily a mystery how to make good food. You can ask an AI how to make good food, follow the instructions, and you're off to the races. The question then is whether you want to be in that race.

Would you have gone to chef school? Would you work in a kitchen? Are you willing to deal with customers, or risk RSI from so many repeated kitchen movements? Are you willing to practice and be tested?

If the answer to any of those is no, then get the hell out of that kitchen and let the people who have more grit than you do their job. Do what you can to make it easier for you to pay them consistently and well on the back-end.


^^;

I chuckled. Institutional prompt will disappear so much faster than you imagine.


Good for you, just try to remember those old days when you complained about bosses and "whine" about things to your friends and some work colleges about day to day stuff. Now think what you would say about a situation when you were fed up and had to quit because you couldn't take it anymore and every day you had these tasks going against your values ( doesn't matter if they are "right" or "wrong", they are yours ).

Also, Google is a multi-billion dreadnought with hundreds of millions of dollars for PR, lawyers and lobbying every year. I'm sure they can take a post about someone "whining" and quitting their job in disagreement. Something tells me Google be fine...


The cheapest EV model Renault sells is around €20K, the cheapest BMW EV is around €65K.

It's safe to say the companies are not in the market bracket, no?


The bit the gets me more than the sale price is servicing.

BMWs have a terrible record for needing expensive repairs.

I know you shouldn’t rely on anecdote, but it seems I do.


The only way I would buy a BMW is if it were an EV. I’m just not brave (or rich) enough to buy their ICEs.


The BMW inline 6 were the best engines ever. Their inline 4 and other are a strong contender for the worst engines ever.


I assume you’re talking about the n52 vs n20. I’ve had both engines as daily drivers, and they’re both fine. The n20 has a bad rap due to early models failing from a timing chain guides breaking.


I miss my E39 530 every time I drive. My next 90s Jag is also going to be a straight-six; the V12 is glorious but heavy.


    > BMWs have a terrible record for needing expensive repairs.
EVs? That makes no sense. EVs are so much simpler to maintain compared to ICEs.


They suffer from some of the same problem your likely modern fridge does, and then kick it up two notches.

In the name of "safety", they have made design decisions such as integrating fuses directly into the very large and expensive control boards and making them non-replacable. Just in case this wasn't enough, they also tend to blow an OTP so that in the event that you have the know how to replace the fuses anyways, nothing will work. Naturally you also cannot just swap in a replacement board, as it needs to go through the same pairing process to the ECU as things like the car doors, which in most cases requires an active certificate/license on the ecu programmer that only dealerships/oem have.


This is a company intentionally making sure EVs didn’t erode service revenues


Wow, I stand corrected. Hat tip for the excellent reply. Assuming what you wrote is true, then, yes, I agree.


In theory they should be, but EVs also tend to be more computerised, proprietary and locked down than ICE cars, so in practice I think it's not as simple as that.

For example there was that case of the car that needed an entire new sealed €5k battery controller because it was in a minor crash and blew a fuse.

My garage charges 50% more for labour on EVs. I'm sure part of that is price discrimination but I bet part is also because working on them is more difficult. I would not be surprised if they need to pay more for access to the manufacturer's diagnostic tools too, which are becoming increasingly required.


Simpler != more reliable. Electronics fail quite often too. Just ask SSDs.

Also new EVs fail often too due to being cost cut to the extreme with the "move fast and break things".


If you take care of the car it’s just brake pads, tires, rotors. Pads and rotors are really simple to DIY. Tires are more expensive than like… an Elantra, but if you’re buying a 60k car you can afford 1.2k in tires… otherwise don’t buy the car.

If you get into an accident or let the bmw get into disrepair via neglect, yeah it’s not cheap to clean up. Body work is expensive on any car though, and I don’t have sympathy for people who own higher-end cars and don’t take care of them, they deserve to pay the price for it.


It's more than that though. Any repairs due to wear and tear or whatever, ends up being really expensive. Although you can probably reduce the costs a bit if you get the non-branded OEM part or potentially the same part from another manufacturer (e.g. the toyota supra uses a lot of bmw parts so if the toyota part might be cheaper than the same bmw part).


That was my whole point actually, the wear and tear is really minimal if you get regular oil changes. Things don’t just break and need replacing. Tires, rotors, brakes, those wear out. The tires are not cheap, rotors and pads aren’t crazy expensive and super easy to DIY.

What other wear and tear things are expensive?


After 22 years, my z4 has needed batteries and a starter.

Recently, there was a problem with the engine misfiring but it was $200.

LA, California


If you had bought a 7 or 5 Series at that time, you would not have had that experience. The 2001 7 Series had something like a 25% roadside breakdown rate.


25% every journey, or 25% over the lifetime of the car? Neither seems really believable but I don't understand how else you would measure this.


25% of cars. It was... not good.


So like... One in four cars would break down at the side of the road before it was otherwise EOL? One roadside breakdown every 800,000 miles or so? That really doesn't sound bad.


Not in the first 800,000. Maybe in the first 8,000. They really struggled with reliability early on the E65, they introduced a lot of new (to them) technology all at once.


It wasn’t/isn’t. The reactions in this subthread surprised me. I guess it’s an anti-ICE thing?


Owned a BMW. Had the audacity to use non-BMW windshield washer fluid. The fluid sensor broke; because in a BMW it’s a fancy sensor that is only compatible with specific washing fluids. Sums up my experience with that car. It was nice to sit in, though.


So you didn’t RTFM? Or ignored it?


I admittedly did not RTFM to learn "how to replace the windshield fluid." Lesson learned. Lol.


Mostly just tires and minor maintenance. You're unlikely to need pad and rotor replacements unless you're driving as if you were on a racetrack every single day.

With daily EV driving you have the opposite problem - regen means you rarely, if ever, actually activate the brakes, so you get rust on them that you need to clean out.


It's still good to know that SOTA is further, and we can expect the more advanced designs to seep into more affordable segments.


They share the same OEMs, and both are following the same ex-China automotive strategy.

Renault has also been thumbing China recently for undermining EU manufacturing as well [0] while China has returned to using Wolf Warrior diplomacy against Europe [1][2][3][4] using the same rhetoric that the Trump admin uses.

Of course, under the Xi admin China's foreign policy has always viewed the EU as inferior and a has-been [5] and has become an active participant in the Ukraine War [6][7].

Europe might not be able to trust the US, but it can't trust China either.

[0] - https://www.reuters.com/world/china/renault-ceo-asks-eu-enco...

[1] - https://www.globaltimes.cn/page/202605/1361926.shtml

[2] - https://www.chinausfocus.com/finance-economy/dear-brussels-d...

[3] - https://www.globaltimes.cn/page/202605/1362161.shtml

[4] - http://news.china.com.cn/2026-06/10/content_118541873.shtml

[5] - https://fddi.fudan.edu.cn/_t2515/57/f8/c21257a743416/page.ht...

[6] - https://www.reuters.com/business/aerospace-defense/russians-...

[7] - https://www.pravda.com.ua/eng/news/2026/06/12/8039041/


> following the same ex-China automotive strategy

Is that why Renault EVs (R5, Twingo) are wholesale developed in China? Doesn't seem very ex-to me, more an in- type of strategy.


The EV batteries are sourced from Ampere and LG (in the EU) and the EESM from Valeo (in the EU).

Sharing platforms isn't something EU manufacturers are opposed to, but they do not want to be dependent on Chinese supply chains. That is the crux of ExChina, especially as the majority of an EV's value is derived from the battery and powertrain.


Why do you think the R5 was developed in China? Renault have been quite open about all the improvements they had to make to their processes, development centres and factories in France to make it. The Twingo was partially developed in China.


only replying to the first link: isn't sourcing (buying or manufacturing locally) parts for Chinese cars made in Europe a good thing?


It is, but the PRC has been pushing back against sourcing from within Europe and only intends to use CDKs to assemble EVs. This is what the EU is pushing back against.

What EU states are now lobbying for is if BYD wants to sell an EV in the EU, it should include European originated parts. Just assembling a knockdown kit in Hungary whose parts were all manufactured in China is not "Made in Europe". If BYD or MG wants to sell a BYD or MG car in the EU, they should source the battery pack and powertrain from the EU.

Alternatively, the PRC can drop similar origination requirements from it's domestic market.

The reality is the PRC won't back down, so they will be tariffed by the EU, especially as the EU has lost patience with the PRC due to their active involvement in the Russia-Ukraine War [0], attempting to use diplomatic immunity to kidnap a French national [1], and attempting to embargo the EU's rare earth imports [2].

Additionally, it's easier for the EU to push back against China versus the US while also winning brownie points in the US.

[0] - https://www.reuters.com/business/aerospace-defense/russians-...

[1] - https://www.lemonde.fr/societe/article/2024/07/02/deux-espio...

[2] - https://www.reuters.com/business/autos-transportation/china-...


> Alternatively, the PRC can drop similar origination requirements from it's domestic market.

Can you share any details on this? Is something I've rarely seen discussed


BMW also produces Mini EVs, which start at £26,840


The cheapest Minis are made by GWM in China, and are using different motors and batteries.

However, comparing prices between cars nowadays is a complicated matter. BMW's iX1 and iX2 (they use the BMW EESM motors) theoretically cost about €55k, but they have been very recently available to lease for about €250 euro per month - so pretty much for the same price as the cheapest electric Renault if leased.


same order of magnitude :)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: