The trick is antitrust style bundling. The massive pile of documents and processes tied to GSuite is a moat which makes it hard to switch to something like o365. Since a company might effectively be locked into GSuite (the primary product), if Google forces companies to buy Gemini (the secondary product) by bundling it with GSuite, they've given themselves a moat in the LLM space using their document/email moat from GSuite.
This is essentially what Google has done, and it's a shame the US is so weak on enforcing antitrust laws.
Psychology in general tends to make the same distinction. There are lots of behaviors which may be considered abnormal but do not have a meaningful impact on the quality of life of the person or those around them, and so there little reason to pathologize it. The goal of medicine (and, in my mind, well-designed public policy) is to prolong quality of life and not to ensure everything adheres to strict standards.
I think what you're really getting at is that it's only useful if the benchmarks are predictive of your workloads. If it predicts well (for example, your tasks are equally easy), then the fact that a larger model can complete it more quickly means that you may be able to complete the task more cheaply, depending on the token cost.
If the benchmarks are non-predictive, well, you can't use them for much of anything, which is of course a recurring problem with every benchmark ever.
Yeah, if the benchmark is actually predictive of the tasks you have then it is trivial to conclude that the cheapest-per-benchmark-task model will be the cheapest one for your tasks…
It might vary between tasks though. A model that’s great at abstract reasoning might be great at writing math proofs but struggle to write software in <insert language>.
I remember reading this blog post when it was first published, but the subsequent updates are better than I would've ever expected this to turn out. Worth checking it out again if you've seen it before :)
I would hope this is true both in the context of LLMs and more broadly, but I think this is especially not the case for LLMs. It's hard to take the idea that companies are trying to hire people with reservations about LLMs seriously when many companies have LLM use mandates. It is counterproductive in the eyes of the employer to hire employees that will be combative on LLM from day one.
Are you going to ask your employees for their 2% cashback if you reimburse them for a purchase they made with their credit card too?
At the end of the day, it seems more practical to focus on ensuring that they picked the appropriate value product/service. If they picked appropriately, does it matter if the company gives them employee a kickback? Sure, the employee will then be incentivized to get the most expensive thing they can, but this was already the case because it's not their money and they want the best thing they can get away with.
Is what you suggest about training even possible? Most exploitation techniques are really just about having in-depth knowledge of how components work. For example, I imagine a sufficiently powerful model could fairly easily re-invent the ROP chain from first principles if it just knew how the stack works. This same principle applies to much more complex attack too; exploitation is often just an exercise in knowing vastly too much trivia, which LLMs tend to have in spades.
It would still degrade it's effectiveness, which is what they claim to want. Exaggeratedly: If it wasn't so, you'd just need fundamental math in the training data, as everything else can be derived.
Is interactive use for coding something that actually works today? With unsafe mode, even frontier hosted models are slow enough I end up just tabbing out to work on other tasks. It would need to be much faster if I am to sit and stare at it while it churns. Local models might be a lot slower but workflow-wise it doesn't change much for me.
Reading this thread, I'm starting to think that I did not fall out of a coconut tree and that I exist in the context of all in which I live and what came before me.
reply