Hacker Newsnew | past | comments | ask | show | jobs | submit | raesene9's commentslogin

It's great that they're making the auto usage tokens free by default and I guess auto mode will be a good default for a lot of workloads, but recent changes to the auto mode classifier just moved me to either use YOLO mode or use a different harness.

I've been using Opus 4.6 for some security related work (it has much looser guardails that later opus models) and last week, all of a sudden, the processes started to fail. It wasn't the main model blocking commands but the auto mode classifier changed how it worked and it started blocking the main models commands.

That's one specific incident, but it does have a wider potential problem which is, if you use Anthropic's harness you'll always be at the risk of sudden breakage from server-side changes that are opaque to the end user, which is a tricky one for building long lasting processes.


I had the same issue with Codex, where one day it just became unusable.

Weirdly, this would occur with very common commands like git push, where the classifier would start to yap about not being able to verify the repository's privacy settings.

My bug reports went ignored, the issue continued, and I eventually just switched to Full access.

I have not yet encountered an issue with Claude's Auto mode.


What works in that case is adding a message saying "I authorise you to do $thing", which the classifier treats as explicit consent to do something and execute commands toward that goal.


Same Story as it ever was. The first time I encountered what I thought was a phishing attack at the bank I worked at 25 years ago, it turned out to be a marketing campaign, with URLs that put our company name as a user before the domain name (back in the day when creds could go in the URL).


Fun fact: still can in Chromium-based browsers. https://textslashplain.com/2023/03/22/attack-techniques-spoo...


Rootless helps, but less now that it used to (pre-2026). There have been a lot of local privilege escalation vulnerabilities in the Linux kernel (dirtyfrag, fragnesia, CIFSwitch et al) and several of those can be repurposed as container breakouts.

As a result, if you're looking for good security isolation, I'd say a (Micro)VM is a better option. The other route is hardening down your container runtime with seccomp/AppArmor/SELinux but that can be a tricky game.


That is what I have been thinking too about recent linux vulnerabilities in context of container escape, but upon brief research I am not convinced it's all that straightforward. For example here https://github.com/Percivalll/Dirty-Frag-Kubernetes-PoC relies on sharing same container layers with other privileged workloads, which is quite a stretch to find in the wild and moreso it says that having a seccomp enabled breaks the exploit - "The default seccomp policy disables the unshare syscall." Other thing is that temporary remedy to lot of these exploits is to blacklist esp4, esp6, algif_aead modules, but how on earth are they going to be loaded in host kernel, which they are not by default, from unprivileged container in first place?


So one of the factors in this is that Kubernetes disables the default seccomp policy provided by the container runtime, by default (you can re-enable it ofc, but you have to know to do that).

As a result I reckon there's more vulnerable containers that you might expect.

Also depending on the environment there's things like dirtyclone https://github.com/raesene/vuln_pocs/tree/main/CVE-2026-4350... which can be triggered where the attacker can start new containers.


They address that in the article :) to quote :-

"So where's this huge price gap coming from? token pricing, prompt caching, and effort-per-task. On SWE for example, K3 works much harder than Fable: roughly 55 turns and 1.3M tokens a task versus 21 turns and 130K. On the long terminal tasks it's the other way around: Fable is the one that spirals, running up 64 turns and 1.5M tokens (sometimes straight into a timeout).

Prompt caching does most of the work of turning that effort into K3's price advantage: even when K3 reads ten times the tokens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a tradeoff. Tasks with extra turns generally mean more wall-clock time per run i.e. slower runs. If you need an answer in two seconds, that matters; if you're running agents in the background at scale, a bill that's a fraction of the size matters a lot more."


So this makes sense for a standard SaaS app - but given that models in general perform much better with low context window usage, it probably also means that Fable is still significantly better at 'frontier-level tasks' -- hard research problems, complex geometric rendering algorithm optimization, etc., no?


As other have mentioned, and I'm the same. Leaving "Co-Authored by:" feels like open disclosure of how the project was created.

That way if a consumer does not want to use LLM generated/assisted software, it's easy for them to do that without having to search around in the codebase for "AI Tells".

Removing those would feel like I was mis-representing my coding abilities (LLMs are better at most coding than I am) :)


Interesting write-up and I do think LLM assisted/powered exploit disclosure is a real concern (I've been able to get models to create container breakouts from Linux LPEs relatively quickly).

One thing I'm surprised about is that GPT-5.6 didn't block that prompt due to guardrails. My experience is that GPT-5.5 and up does not like offensive security work (similar to Opus 4.7+/Fable).

I didn't notice it but I'd assume that the authors have some level of cyber approvals from OpenAI to relax the guardrails a bit.


This might help https://chatgpt.com/cyber ease the guardrails a bit.


Thanks! I wonder if Claude has something similar?



Though this program does not apply to Fable. Which is why most security researchers have started flocking to Sol.


You might find some areas to criticize Berkshire Hathaway but I don't see being lazy as one of them. This is one of the most successful investment companies of all time and they got that way by being better than most at judging when the right time to get in and get out of the market, and by putting in the work on researching where/when to buy.

Might there be opportunities they miss? I'm sure there will be, but perhaps finding those is just too risky at the moment, so they've looked at the options and decided not to invest.


> This is one of the most successful investment companies

It was for a long time. There is not a lot of evidence that's still the case (so far at least but even if the crash comes but its not big enough its not guaranteed they will outperform S&P 500 over a several year period).


Future returns are never guaranteed but over the course of the orgs history (since 1965) they've done a fair bit better than the S&P 500.... https://www.visualcapitalist.com/warren-buffett-vs-the-sp-50...


I'd say for some sectors and users, we're already seeing that. After GLM-5.2's release there were quite a few stories about it picking up use.

Then looking at Openrouter's stats we can see heavy use of non Anthropic/OpenAI models https://openrouter.ai/rankings#top-models

I'm not sure I'd call them rando, but Deepseek/Qwen/GLM all have decent models that will work fine for some use cases at much lower costs that the SOTA models from Anthropic/OpenAI.


Exactly. As soon as business works out that there are open models that can achieve parity at a fraction of the cost, they will switch model. They have to. Business exists to make money, and if they can do the same thing at lower cost, they will.


It's an interesting thing and we can only speculate from the outside, but there's some obvious reasons why they'd literally hand out money in the form of free compute to people who have already committed to paying them.

- They've had to commit to minimum spend with their suppliers and their actual usage is below that level, so they might as well give it to end-users. That implies either their demand is waning or their forecasts were off.

- They want the usage numbers for this period to be higher and are willing to spend money to achieve it. I guess if they're about to IPO maybe showing more usage is good?

- They've got a lot of competition and want to keep market share. With lots of new models coming out, some of them much cheaper than Anthropic's options, you can see why this might make sense but it's effectively burning cash, so they won't want to keep doing that for long I'd guess...


I think until they produce full financial information, which will happen I expect when their S-1 is published, we won't have a good picture on Anthropics true position.

There's a lot of different ways to calculate "profitable" depending on what's included or excluded...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: