I see a lot of businesses who just dump this all on their employees and then get mad at their employees effectively for not following the most recent LLM related talk on twitter.
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
Assuming the employees are calling themselves "engineers" here, who else is supposed to be doing that work? If we were going through a renaissance in material science, don't you think it would be the mechanical engineers who try to figure out how to figure out what sorts of materials might be suitable for their use case? Surely you don't want the sales team or executives figuring that out.
Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
It's very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.
Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.
Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.
I tend to personally lean toward generalism, but there are very real benefits to specialization. So at least in theory, while certainly some software engineers need to be the ones figuring this stuff out, it would be more efficient for organizations if that was a smaller specialized group maintaining a curated set of platforms and tools, rather than everyone individually going it alone.
In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
> Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
On these things, yes. But in the case of LLMs, this is impossible. It's not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that's not a guarantee of anything.
And that's why I don't think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
Yes, but before you would have one specialist in each topic and time to process different possible outputs. Now it is "just do sh*t" because code is cheap. And people are slowing realising that code and doing was never the hard part.
"But models will get better". Maybe.
But there has been a decline in the speed of development
Do I need to become an expert on all the LLM models, or can someone else make a decision? In my case I've been given co-pilot which gives a discount if I use "auto" which is to say I let the system choose the model. A few months ago I often found auto not good enough and switched models manually at extra cost, but I never wanted to become an expert in this. These days auto has always been good enough and so I"m glad I don't have to be the expert.
If you enjoy being an expert in which LLM is best, then I'm all for it. However that isn't an interesting problem for me or my boss. I'm happy someone else figured it out so I can work on interesting problems.
The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
At least when I was studying optical engineering, we had to learn (and memorize, which annoyed me greatly) materials properties of different glasses, plastics, some gems like sapphire, etc. Obviously optical properties, but also thermal properties. Maybe some kind of stress thing; I don't remember. But also, we had to have some knowledge of which things cost lots of money. Which things are easier to manufacture with and what sorts of manufacturing techniques work with different materials. etc. I presume a working engineer just knows all of this stuff then.
Yes, and that's normal, and the things I mentioned working engineers often have opinions about and some are even well informed, but traditionally there'd be a skew of opinions and different levels of strength of those held. You could inform yourself about those options, test them, and over time become well informed yourself.
That's not really possible with frontier models changing every 3-6 months, just like javascript frameworks there's not enough time to learn the ins and outs and they mostly expose the same interface, so unless you have rigorous evals you are working on vibes, and the amount of meetings I have had in the last 3 years of engineers confidently reporting on their vibes is SO TIRING.
We should expect mechanical engineers to be aware of the limits of the physical world they're engineering for. It's not like we'd excuse making a design that called for a material with impossible properties just because they're not a materials scientist.
I'm in agreement with this. I had informed opinions on databases, programming languages, and frameworks before the Eternal Sloptoberfest, and increasingly I understand LLM related things just as part of doing business.
> Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
All while suffering from the increased the mental load/lack of focus from even more workload. "You've got AI, you should be able to get it done today, right?"
There's way more anxiety here than is necessary. If you don't want to optimize, then it's pretty simple right now, just use Opus 5.5 or GPT 6.1 Sol. If you want to quickly test out the alternatives and see if the decreased ability in long horizon tasks is offset by the decrease in costs, toss a few bucks into your favorite neocloud (Fireworks was fine last time I tried them) and give GLM 5.3 and Deepseek v4.1 a spin, and see if they're good enough. OhMyPi makes it trivial to swap models without swapping your harness.
Everyone has a great researcher piped to their desk now. You can have it follow the Twitter zeitgeist for you, you don't have to do it yourself.
I am one of those employees who can make architecture decisions, and make them well, and have for decades.
The level of appropriate LLM usage in my opinion? Ask Gemini some questions when you need to search the web then read the sources it gives you.
LLM written code still has the problem IBM identified. A computer cannot be heald responsible, and so it cannot be allowed to make decisions. That applies to executive management AND software engineering.
Of course that requires working in an industry, like aerospace, where engineers are (usually, Boeing not counting) heald accountable.
You are not paying attention then. Lots of programs have all sorts of basic things that can be changed with low risk to get pretty good performance benefit, for example. You can literally just ask an LLM to look through your code and see where you were e.g. unnecessarily allocating or had poor struct ordering and fix it without changing any interfaces or functional code and have it ensure that unit tests pass and that it runs benchmarks to see that the code is faster.
You can ask it to implement the pattern you want. Like we had a bunch of gnarly intertwined logic inside of looping constructs that I could say, hey, create iterators for this, extract out the filtering logic, extract out the transformation logic, etc. I'm making the decision, but it can do the mechanical work. And again, I can tell it to benchmark my refactor to ensure that it actually cleaned up allocations instead of introducing them. It becomes easier for me to actually be more thorough in my work.
You can also use them to quickly build out prototypes for different approaches so that you can make better informed decisions at an architectural level.
You can give them a description of a bug and it will read through your code and figure out where the problem lies, even across multiple repos with complex interactions. You can of course still confirm it, but they've been superhuman at this since at least last December.
LLMs are absolutely a game changer for software development in the hands of somebody who knows what they want. For low-level things, they're an extremely solid junior (they don't ever really make mistakes, they just write tasteless crap). For high-level things, they're an extremely solid discussion partner with wide knowledge and great reasoning skills.
it is more complex than cheapest. They do something to evaluate the question you ask and then route to different models. I also use auto in vscode, it tells which model it selected if you look close - which I rarely bother. I have seen a dozen different models over the past months, but they have all worked which means either they are all good, or vscode is good at selecting. I don't care which because as you say "good enough"
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns???
Right, the problem is frontier models can do ~all of that pretty well, a lot better than they could 9 months ago. Luna 6 can do most of it practically free and instantaneously. So, your quote is likely what's increasingly being said in c-suite meetings as they decide to start mass layoffs.
You can set spending limits, but I don't feel like that helps much because everyone's accustomed to AI. Nobody would accept "We're out of usage so we have to wait until Monday" and do all of their work manually.
I believe that we're in a scenario where usage is unlikely to go down and neither are frontier AI costs.
I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
A marketer doesn't need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
It's not just being accustomed to AI. It's that the adoption of "agentic" workflows can make teams dependent on AI.
This summer my team basically shut down when we hit a usage cap because we had recently pushed a decent portion of our devops workflow into agent skills. It had been done in such a way that it was difficult for humans to navigate. Relevant scripts turned out to be buggy and poorly documented, and nobody realized because these harnesses that are tuned to be absurdly tenacious about searching for workarounds had been quietly burning heaps of tokens on muddling through instead of raising any alerts about the horribly broken state of the system.
I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
True, I guess I was mostly thinking of people using AI in their workflow, not just AI code review and things of that nature.
> I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
It's true that moderate usage harms current harness providers, but I think that's because we currently conceptualize them as AI services. I think in the near future, companies will have harnesses that centrally configure MCPs, CLIs, model routing, etc.
and they'll be much closer to an auth/permission service than to an AI service.
I have some bias as a software dev/manager but at least in our org I noticed you can get a lot better value out of your tokens with some proper configuration of context and skills. Providing the latest models with good context can result in extreme savings.
After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there's little concern about drift.
The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they'll waste a ton of tokens "thinking" through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
1. Truck drivers are not deciding to take a Ferrari instead of their truck. The company knows roughly how much gas a given delivery is going to consume and how much that is.
With AI, that definitely happens if you default to Claude Opus/GPT Sol.
2. A truck driver's value scales with gas usage. The further they drive, the more valuable.
Not true for tokens.
3. Gas doesn't create extreme outliers. One developer can easily consume 100x the tokens of another, which makes AI economics much more extreme.
Well no, but even with the gas prices jumping the way they are, that is a much more predictable cost than AI spend. Drive 250km, in 2015 Volvo truck is X liters of fuel at X cost. Solve a predefined problem using AI... you have no idea what that will cost you, if you did you'd most likely already have the answer.
No, but truck drivers (who work for a company making this decision) do sometimes get told there is no work for them because nobody wants to pay the price. (this also happens when there is nothing to ship though - drivers can't tell the difference)
My company setup tiers (as of this month - I can't tell you how it works since this is only day 5). I automatically get so many tokens, if I use them up I'm auto approved to upgrade to the next tier. They told everybody wait until you run out to request a tier increase - the real goal is a few people had a runaway job burning tokens that they didn't stop, the tier is a force stop point in case that happens to you.
Yeah I feel like token usage is probably power law distributed as well.
Also this feels like a good way to detect runaway jobs. I’ve seen something similar implemented for data plans where you had unlimited data, but every 50 GB had to request another free 50 GB to prevent people from turning their phone plan into their home WiFi
Well at a big bank I work, if you finish in 1 week you do manual coding the rest of the month until next month. Management doesn’t give a shit. And they have daily talks about using more AI.
This has happened before with cloud computing. There was lots of people talking about how every employee would spin up their own servers etc and huge productivity gains. In the end all of that leverage mostly got absorbed by the engineering teams and gatekept and the rest of the business gets to use it through a GUI.
> I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
> This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
Yeah exactly. Plus, most people don't want to configure MCPs and write skills about how to correctly use CLIs. The type of person who uses Hacker News may do that to customize whatever harness they use.
A sales manager at a sheet metal firm in Wisconsin doesn't. They want AI tools that make their lives easier, but they don't know or care whether that happens with an MCP.
I guess the cloud is a great analogy because that same sales manager may have loved Salesforce because it was always available and updated compared to a manual spreadsheet.
Does the sales manager even care if it's AI either post hype? For most people on the "AI assistant" use case they care about having something that they can type a question into and smart answers come out. If you could serve it with a cache and call it AI they wouldn't care if it's not giving dumb answers.
I do quite a bit of coding with Claude but am perfectly fine on the $20/mo plan. You people who just let agents go for hours on end... I'm not sure you're doing it right.
I have agents going pretty much nonstop for hobby projects. This is great value for me because I get to see a lot of things I'm interested in come to life. It's very token-inefficient though and if usage limits were reduced much this wouldn't be worth doing.
For my actual job though, I don't let agents go for hours on end and I care about the code. I spend between $1000 and $2000 per month there. Of course we pay API pricing. I wonder how much you're getting for that $20 in API pricing, it could be 10x or even 50x. At least on the max plans it can be that high if you are hitting usage limits, I don't know about the 20.
There are multiple camps of people, and everyone thinks the other camps are doing it wrong.
- full autonomous camp
- developer augmentation camp
- no-ai camp
The trouble I see with all of it is that the future seems unpredictable at the moment. The costs related to AI are low enough at the moment that full autonomous seems to be possible, but we have reasons to believe that costs will rise significantly, which may change that calculus. The no-AI camp is in ostrich mode, and is betting on this all going away once the bubble pops. The developer augmentation camp treats it like just another tool, which is somewhere in the middle.
The trouble is that even if there is a clear advantage today, the ground truth of costs built in is probably not stable.
Yeah I agree with this. I have my own preferences (augmentation, which is obviously right, because it's my camp and I wouldn't be in that camp if it were the wrong one, duh!), but I'm glad there are so many people doing so many different things and debating each other about it. That kind of messiness is the only way to figure this out.
Depends on what you're doing. For example, I'm having one chat extend/optimize a (mildly-novel) CUDA program. Opus 5.5 (high/xhigh) is very efficient, but a single thread would still cost me the Pro plan usage in ~2 days (less if I were just running it unattended). This is one mostly-bounded program, and it's not even that large.
In fact, before 5.5, I would have just used Fable instead.
The trick is having enough speculative explorations going on at once so that you can see if the ones going off for hours on end strike gold while doing so or not, then being prepared to either cut them off or get them back to the point, and repeat.
If you don't have at least some workloads doing that you're missing out on the biggest wins from the current phase of the technology.
Many businesses are on Enterprise plans where usage is all billed per token, rather than having a flat rate seat with rate limits. You may find it interesting to look at your token usage and calculate what you'd be paying monthly if you paid API rates instead.
Just wanted to put it out there: the FinOps Foundation launched the Tokenomics Foundation last year, to focus on practices and tooling to improve AI cost controls. (FinOps itself is a collaboration between the Finance and Engineering orgs, though IMO more driven by Finance.)
And more broadly, AI spend isn't the only problem. Engineers have to track multiple dimensions beyond spend - efficacy, safety, speed, autonomy. All this rolls up into AI adoption and value maximization. High costs might be tolerable, if you get sufficient value out of it.
Cost is the exception to how LLM learning curves have dropped off and in some ways rewarded late adopters. (e.g. prompt engineering is hardly worth the bother when the models now think so hard about user intent)
It continues to be an uphill battle to help people at $DAYJOB understand that the relationship between turns and cost is nonlinear. And many still don't understand the idea of a system prompt, that they can control how chatty all responses are.
Did anyone ever do a software project before AI? You can predict costs far better than you can with people.
Yeah, that guy you hired to write the prototype went on a 2 month bender and created 0 usable code. Did you budget for that? Oh, okay. The strategies to deal with spending on tokens are nothing compared to the overhead of managing actual people and their outputs.
Why not just set a spending limit per person and extend/adjust the limit case by case. This would actually make people be more mindful about burning tokens on useless stuff
That is what we have been doing, you get a default monthly limit, you can request more, managers chat with you about why, so there is flexibility but you feel some pressure to reduce your spend, so more of us fine tune our effort levels and model choices depending on the task, rather than just OPUS HIGH all day long.
It's not necessarily corporate, but mainly due to the fact that AI tools are quite new and we're just learning to use them efficiently.
One example is that recently I used the superpowers claude code plugin. I thought it was good because it was popular on github (with 296k stars). However, for a small feature change it basically created 4 different plans and launched 4 agents to work in parallel (write code, write tests, update docs, etc) and then once all 4 sub agents are done it launched another agent to verify the work of all 4 sub agents. It ended up burning so much tokens and taking actually longer to finish and just could have been done without any agent. I immediately uninstall the plugin. These are the kind of things that developers try out that ends up increasing the costs
I generally don't find the analogies to "being a manager of agents" to be useful, but I think this is where it is the most useful. Corporations have always had this problem of needing to figure out who to give budget to, in the face of ambiguity about how well each business unit is using their budget. There are many approaches to solving this problem, and some of those approaches apply to figuring out token spend as well.
There is a big gap currently in AI assisted development. You don't need to max out tokens to build products or develop with AI. From my experience with AI assisted coding, we still use just two Plus accounts each across a three person team for maintaining multiple repositories totaling around 700K lines of code.
Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
Alternate take: it takes a lot of up front effort to ensure when you "max out tokens" the work is valuable.
The same thing goes for token use operating business processes (vs. building software that runs business processes): Let's imagine buying $250k tokens per year to operating some business processes - how much human effort is needed to ensure that level of spend is valuable? I'm using a number that could be "we could hire a skilled human at that level" as that changes the feel of the question.
That is where benchmarks matters. Good developers may be terrible with AI. Companies should put efforts, pick up a small team, and iterates on ways to reduce token usage. In our case we arrived after lot of back and forth. You shoul dnever allow somebody to use AI on one fine day and expect them to deliver 10X.
> Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
The whole point of AI is for companies to invest less resources into employees. There's no way they start spending money on training people now, when that was already something that made executives roll their eyes before AI.
Wrong perception inmho. Ai is a new tool. people doesnt know how to efficiently use it. Iam saying after seeing countless of x and reddit feeds. We work on strict budget an dthat is where we iterates on maximum efficiency and it worked. Money may be yours, but compute consumes resources that belong to the world.
Eh, this is a relatively temporary phase. As LLM capabilities have been changing fast, their ability to change existing workflows has been unknown. It’s made sense for businesses to go hog wild with them for a while - just to try to get a handle on what’s possible.
At this point, local models have become feasible, and people using them are beginning to get a feel for the tradeoffs vs the frontier models. As the frontier advances, the question becomes “how much will I spend for a given quantity and quality of AI work?” with local hardware providing a pricing anchor point.
When I can price hardware and ops for a given capability level - I have a budget again. From there it’s a question of how much faster/capable/cheaper is a given provider (and, you know, how much do I trust sending them all my IP?).
Turns out running an unattended LLM like a slot machine is expensive.
IMO, human in the loop is the only serious usage of AI (I know, I know, "software factories bro"). Everything else is a hope and a prayer and a big bill.
Maybe I'm in the bubble, but I don't see a lot of "software factories bros" here, there are plenty of "coding is solved"-bros (they're also annoying btw)
That "intuition" is not intuitive. How do I choose an appropriate reasoning effort?
I use these models all day long, experiment, and have no clue how to choose that.
It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems "how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on" though it's certainly not a guarantee on anything.
These models are being released at breakneck speed even before their creators know how to use them. It's a big project of collective discovery to figure out what they are doing and how to use them.
Yeah this I agree because even personally all my rubrics break when a new model is released. Things were stable while I was using 5.6 GPT. There ought to be some room to explore and understand this intuition and yes one must account for this bugdet.
I see a lot of businesses who just dump this all on their employees and then get mad at their employees effectively for not following the most recent LLM related talk on twitter.
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
This is getting stupid folks.
Assuming the employees are calling themselves "engineers" here, who else is supposed to be doing that work? If we were going through a renaissance in material science, don't you think it would be the mechanical engineers who try to figure out how to figure out what sorts of materials might be suitable for their use case? Surely you don't want the sales team or executives figuring that out.
Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
It's very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.
Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.
Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.
>It's very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.
Hah, welcome to what it can feel like to manage people.
I tend to personally lean toward generalism, but there are very real benefits to specialization. So at least in theory, while certainly some software engineers need to be the ones figuring this stuff out, it would be more efficient for organizations if that was a smaller specialized group maintaining a curated set of platforms and tools, rather than everyone individually going it alone.
In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
> Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
On these things, yes. But in the case of LLMs, this is impossible. It's not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that's not a guarantee of anything.
And that's why I don't think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
Yes, but before you would have one specialist in each topic and time to process different possible outputs. Now it is "just do sh*t" because code is cheap. And people are slowing realising that code and doing was never the hard part.
"But models will get better". Maybe. But there has been a decline in the speed of development
Do I need to become an expert on all the LLM models, or can someone else make a decision? In my case I've been given co-pilot which gives a discount if I use "auto" which is to say I let the system choose the model. A few months ago I often found auto not good enough and switched models manually at extra cost, but I never wanted to become an expert in this. These days auto has always been good enough and so I"m glad I don't have to be the expert.
If you enjoy being an expert in which LLM is best, then I'm all for it. However that isn't an interesting problem for me or my boss. I'm happy someone else figured it out so I can work on interesting problems.
The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
No, it would be the materials scientists/engineers.
At least when I was studying optical engineering, we had to learn (and memorize, which annoyed me greatly) materials properties of different glasses, plastics, some gems like sapphire, etc. Obviously optical properties, but also thermal properties. Maybe some kind of stress thing; I don't remember. But also, we had to have some knowledge of which things cost lots of money. Which things are easier to manufacture with and what sorts of manufacturing techniques work with different materials. etc. I presume a working engineer just knows all of this stuff then.
Yes, and that's normal, and the things I mentioned working engineers often have opinions about and some are even well informed, but traditionally there'd be a skew of opinions and different levels of strength of those held. You could inform yourself about those options, test them, and over time become well informed yourself.
That's not really possible with frontier models changing every 3-6 months, just like javascript frameworks there's not enough time to learn the ins and outs and they mostly expose the same interface, so unless you have rigorous evals you are working on vibes, and the amount of meetings I have had in the last 3 years of engineers confidently reporting on their vibes is SO TIRING.
We should expect mechanical engineers to be aware of the limits of the physical world they're engineering for. It's not like we'd excuse making a design that called for a material with impossible properties just because they're not a materials scientist.
I'm in agreement with this. I had informed opinions on databases, programming languages, and frameworks before the Eternal Sloptoberfest, and increasingly I understand LLM related things just as part of doing business.
> Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
All while suffering from the increased the mental load/lack of focus from even more workload. "You've got AI, you should be able to get it done today, right?"
There's way more anxiety here than is necessary. If you don't want to optimize, then it's pretty simple right now, just use Opus 5.5 or GPT 6.1 Sol. If you want to quickly test out the alternatives and see if the decreased ability in long horizon tasks is offset by the decrease in costs, toss a few bucks into your favorite neocloud (Fireworks was fine last time I tried them) and give GLM 5.3 and Deepseek v4.1 a spin, and see if they're good enough. OhMyPi makes it trivial to swap models without swapping your harness.
Everyone has a great researcher piped to their desk now. You can have it follow the Twitter zeitgeist for you, you don't have to do it yourself.
[flagged]
Was it ever not stupid?
[flagged]
A company I know well complained people weren't using enough and now that they mad everyone use it, it complains it is too expensive.
I am one of those employees who can make architecture decisions, and make them well, and have for decades.
The level of appropriate LLM usage in my opinion? Ask Gemini some questions when you need to search the web then read the sources it gives you.
LLM written code still has the problem IBM identified. A computer cannot be heald responsible, and so it cannot be allowed to make decisions. That applies to executive management AND software engineering.
Of course that requires working in an industry, like aerospace, where engineers are (usually, Boeing not counting) heald accountable.
You are not paying attention then. Lots of programs have all sorts of basic things that can be changed with low risk to get pretty good performance benefit, for example. You can literally just ask an LLM to look through your code and see where you were e.g. unnecessarily allocating or had poor struct ordering and fix it without changing any interfaces or functional code and have it ensure that unit tests pass and that it runs benchmarks to see that the code is faster.
You can ask it to implement the pattern you want. Like we had a bunch of gnarly intertwined logic inside of looping constructs that I could say, hey, create iterators for this, extract out the filtering logic, extract out the transformation logic, etc. I'm making the decision, but it can do the mechanical work. And again, I can tell it to benchmark my refactor to ensure that it actually cleaned up allocations instead of introducing them. It becomes easier for me to actually be more thorough in my work.
You can also use them to quickly build out prototypes for different approaches so that you can make better informed decisions at an architectural level.
You can give them a description of a bug and it will read through your code and figure out where the problem lies, even across multiple repos with complex interactions. You can of course still confirm it, but they've been superhuman at this since at least last December.
LLMs are absolutely a game changer for software development in the hands of somebody who knows what they want. For low-level things, they're an extremely solid junior (they don't ever really make mistakes, they just write tasteless crap). For high-level things, they're an extremely solid discussion partner with wide knowledge and great reasoning skills.
It sounds like your decision-making skills might not be as amazing as you think, if you're completely writing off LLM code generation.
I use "auto" in vscode which likely selects for cheapest. Good enough!
it is more complex than cheapest. They do something to evaluate the question you ask and then route to different models. I also use auto in vscode, it tells which model it selected if you look close - which I rarely bother. I have seen a dozen different models over the past months, but they have all worked which means either they are all good, or vscode is good at selecting. I don't care which because as you say "good enough"
[dead]
Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns???
Right, the problem is frontier models can do ~all of that pretty well, a lot better than they could 9 months ago. Luna 6 can do most of it practically free and instantaneously. So, your quote is likely what's increasingly being said in c-suite meetings as they decide to start mass layoffs.
You can set spending limits, but I don't feel like that helps much because everyone's accustomed to AI. Nobody would accept "We're out of usage so we have to wait until Monday" and do all of their work manually.
I believe that we're in a scenario where usage is unlikely to go down and neither are frontier AI costs.
I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
A marketer doesn't need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
It's not just being accustomed to AI. It's that the adoption of "agentic" workflows can make teams dependent on AI.
This summer my team basically shut down when we hit a usage cap because we had recently pushed a decent portion of our devops workflow into agent skills. It had been done in such a way that it was difficult for humans to navigate. Relevant scripts turned out to be buggy and poorly documented, and nobody realized because these harnesses that are tuned to be absurdly tenacious about searching for workarounds had been quietly burning heaps of tokens on muddling through instead of raising any alerts about the horribly broken state of the system.
I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
True, I guess I was mostly thinking of people using AI in their workflow, not just AI code review and things of that nature.
> I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
It's true that moderate usage harms current harness providers, but I think that's because we currently conceptualize them as AI services. I think in the near future, companies will have harnesses that centrally configure MCPs, CLIs, model routing, etc.
and they'll be much closer to an auth/permission service than to an AI service.
I have some bias as a software dev/manager but at least in our org I noticed you can get a lot better value out of your tokens with some proper configuration of context and skills. Providing the latest models with good context can result in extreme savings.
After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there's little concern about drift.
The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they'll waste a ton of tokens "thinking" through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
No one would tell truck drivers "We hit our gas budget for this week, deliver the rest of the packages on foot."
No, but I think this is wrong on three levels:
1. Truck drivers are not deciding to take a Ferrari instead of their truck. The company knows roughly how much gas a given delivery is going to consume and how much that is.
With AI, that definitely happens if you default to Claude Opus/GPT Sol.
2. A truck driver's value scales with gas usage. The further they drive, the more valuable.
Not true for tokens.
3. Gas doesn't create extreme outliers. One developer can easily consume 100x the tokens of another, which makes AI economics much more extreme.
Well no, but even with the gas prices jumping the way they are, that is a much more predictable cost than AI spend. Drive 250km, in 2015 Volvo truck is X liters of fuel at X cost. Solve a predefined problem using AI... you have no idea what that will cost you, if you did you'd most likely already have the answer.
No, but truck drivers (who work for a company making this decision) do sometimes get told there is no work for them because nobody wants to pay the price. (this also happens when there is nothing to ship though - drivers can't tell the difference)
My company setup tiers (as of this month - I can't tell you how it works since this is only day 5). I automatically get so many tokens, if I use them up I'm auto approved to upgrade to the next tier. They told everybody wait until you run out to request a tier increase - the real goal is a few people had a runaway job burning tokens that they didn't stop, the tier is a force stop point in case that happens to you.
Yeah I feel like token usage is probably power law distributed as well.
Also this feels like a good way to detect runaway jobs. I’ve seen something similar implemented for data plans where you had unlimited data, but every 50 GB had to request another free 50 GB to prevent people from turning their phone plan into their home WiFi
Well at a big bank I work, if you finish in 1 week you do manual coding the rest of the month until next month. Management doesn’t give a shit. And they have daily talks about using more AI.
Banks have historically invested as little as possible in technology, unless it directly generates cash (fintech). I'm not too surprised.
.... that sounds hilariously ineffective
so you basically are rationing your frontier tokens?
This has happened before with cloud computing. There was lots of people talking about how every employee would spin up their own servers etc and huge productivity gains. In the end all of that leverage mostly got absorbed by the engineering teams and gatekept and the rest of the business gets to use it through a GUI.
> I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
> This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
Yeah exactly. Plus, most people don't want to configure MCPs and write skills about how to correctly use CLIs. The type of person who uses Hacker News may do that to customize whatever harness they use.
A sales manager at a sheet metal firm in Wisconsin doesn't. They want AI tools that make their lives easier, but they don't know or care whether that happens with an MCP.
I guess the cloud is a great analogy because that same sales manager may have loved Salesforce because it was always available and updated compared to a manual spreadsheet.
Does the sales manager even care if it's AI either post hype? For most people on the "AI assistant" use case they care about having something that they can type a question into and smart answers come out. If you could serve it with a cache and call it AI they wouldn't care if it's not giving dumb answers.
I do quite a bit of coding with Claude but am perfectly fine on the $20/mo plan. You people who just let agents go for hours on end... I'm not sure you're doing it right.
That's cool.
I have agents going pretty much nonstop for hobby projects. This is great value for me because I get to see a lot of things I'm interested in come to life. It's very token-inefficient though and if usage limits were reduced much this wouldn't be worth doing.
For my actual job though, I don't let agents go for hours on end and I care about the code. I spend between $1000 and $2000 per month there. Of course we pay API pricing. I wonder how much you're getting for that $20 in API pricing, it could be 10x or even 50x. At least on the max plans it can be that high if you are hitting usage limits, I don't know about the 20.
There are multiple camps of people, and everyone thinks the other camps are doing it wrong.
- full autonomous camp
- developer augmentation camp
- no-ai camp
The trouble I see with all of it is that the future seems unpredictable at the moment. The costs related to AI are low enough at the moment that full autonomous seems to be possible, but we have reasons to believe that costs will rise significantly, which may change that calculus. The no-AI camp is in ostrich mode, and is betting on this all going away once the bubble pops. The developer augmentation camp treats it like just another tool, which is somewhere in the middle.
The trouble is that even if there is a clear advantage today, the ground truth of costs built in is probably not stable.
Yeah I agree with this. I have my own preferences (augmentation, which is obviously right, because it's my camp and I wouldn't be in that camp if it were the wrong one, duh!), but I'm glad there are so many people doing so many different things and debating each other about it. That kind of messiness is the only way to figure this out.
The no-AI people are in eagle mode! We watch stupid corporations failing despite the largest propaganda campaign in history. The latest Hail Mary:
https://www.reuters.com/business/media-telecom/musk-says-he-...
That will fix adoption of a broken technology and all debt issues!
Depends on what you're doing. For example, I'm having one chat extend/optimize a (mildly-novel) CUDA program. Opus 5.5 (high/xhigh) is very efficient, but a single thread would still cost me the Pro plan usage in ~2 days (less if I were just running it unattended). This is one mostly-bounded program, and it's not even that large. In fact, before 5.5, I would have just used Fable instead.
The trick is having enough speculative explorations going on at once so that you can see if the ones going off for hours on end strike gold while doing so or not, then being prepared to either cut them off or get them back to the point, and repeat.
If you don't have at least some workloads doing that you're missing out on the biggest wins from the current phase of the technology.
What doe striking gold here mean? What "speculative explorations" are you having the agents do?
Many businesses are on Enterprise plans where usage is all billed per token, rather than having a flat rate seat with rate limits. You may find it interesting to look at your token usage and calculate what you'd be paying monthly if you paid API rates instead.
they're not doing it right _or_ they're doing it correctly.
We're sorta in the age of alchemy. Lots of cranks out there, but there are real recipes.
Just wanted to put it out there: the FinOps Foundation launched the Tokenomics Foundation last year, to focus on practices and tooling to improve AI cost controls. (FinOps itself is a collaboration between the Finance and Engineering orgs, though IMO more driven by Finance.)
And more broadly, AI spend isn't the only problem. Engineers have to track multiple dimensions beyond spend - efficacy, safety, speed, autonomy. All this rolls up into AI adoption and value maximization. High costs might be tolerable, if you get sufficient value out of it.
Cost is the exception to how LLM learning curves have dropped off and in some ways rewarded late adopters. (e.g. prompt engineering is hardly worth the bother when the models now think so hard about user intent)
It continues to be an uphill battle to help people at $DAYJOB understand that the relationship between turns and cost is nonlinear. And many still don't understand the idea of a system prompt, that they can control how chatty all responses are.
https://archive.ph/aBXzS
Did anyone ever do a software project before AI? You can predict costs far better than you can with people.
Yeah, that guy you hired to write the prototype went on a 2 month bender and created 0 usable code. Did you budget for that? Oh, okay. The strategies to deal with spending on tokens are nothing compared to the overhead of managing actual people and their outputs.
Why not just set a spending limit per person and extend/adjust the limit case by case. This would actually make people be more mindful about burning tokens on useless stuff
That is what we have been doing, you get a default monthly limit, you can request more, managers chat with you about why, so there is flexibility but you feel some pressure to reduce your spend, so more of us fine tune our effort levels and model choices depending on the task, rather than just OPUS HIGH all day long.
corporate has no idea what stuff is "useless stuff" . Previously it used to be invisible, now token spends are putting a number on it.
If i had to guess upwards of 80 percent in corporate is "useless stuff".
It's not necessarily corporate, but mainly due to the fact that AI tools are quite new and we're just learning to use them efficiently.
One example is that recently I used the superpowers claude code plugin. I thought it was good because it was popular on github (with 296k stars). However, for a small feature change it basically created 4 different plans and launched 4 agents to work in parallel (write code, write tests, update docs, etc) and then once all 4 sub agents are done it launched another agent to verify the work of all 4 sub agents. It ended up burning so much tokens and taking actually longer to finish and just could have been done without any agent. I immediately uninstall the plugin. These are the kind of things that developers try out that ends up increasing the costs
I generally don't find the analogies to "being a manager of agents" to be useful, but I think this is where it is the most useful. Corporations have always had this problem of needing to figure out who to give budget to, in the face of ambiguity about how well each business unit is using their budget. There are many approaches to solving this problem, and some of those approaches apply to figuring out token spend as well.
There is a big gap currently in AI assisted development. You don't need to max out tokens to build products or develop with AI. From my experience with AI assisted coding, we still use just two Plus accounts each across a three person team for maintaining multiple repositories totaling around 700K lines of code.
Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
>You don't need to max out tokens
Alternate take: it takes a lot of up front effort to ensure when you "max out tokens" the work is valuable.
The same thing goes for token use operating business processes (vs. building software that runs business processes): Let's imagine buying $250k tokens per year to operating some business processes - how much human effort is needed to ensure that level of spend is valuable? I'm using a number that could be "we could hire a skilled human at that level" as that changes the feel of the question.
That is where benchmarks matters. Good developers may be terrible with AI. Companies should put efforts, pick up a small team, and iterates on ways to reduce token usage. In our case we arrived after lot of back and forth. You shoul dnever allow somebody to use AI on one fine day and expect them to deliver 10X.
> Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
The whole point of AI is for companies to invest less resources into employees. There's no way they start spending money on training people now, when that was already something that made executives roll their eyes before AI.
Wrong perception inmho. Ai is a new tool. people doesnt know how to efficiently use it. Iam saying after seeing countless of x and reddit feeds. We work on strict budget an dthat is where we iterates on maximum efficiency and it worked. Money may be yours, but compute consumes resources that belong to the world.
Eh, this is a relatively temporary phase. As LLM capabilities have been changing fast, their ability to change existing workflows has been unknown. It’s made sense for businesses to go hog wild with them for a while - just to try to get a handle on what’s possible.
At this point, local models have become feasible, and people using them are beginning to get a feel for the tradeoffs vs the frontier models. As the frontier advances, the question becomes “how much will I spend for a given quantity and quality of AI work?” with local hardware providing a pricing anchor point.
When I can price hardware and ops for a given capability level - I have a budget again. From there it’s a question of how much faster/capable/cheaper is a given provider (and, you know, how much do I trust sending them all my IP?).
Turns out running an unattended LLM like a slot machine is expensive.
IMO, human in the loop is the only serious usage of AI (I know, I know, "software factories bro"). Everything else is a hope and a prayer and a big bill.
Maybe I'm in the bubble, but I don't see a lot of "software factories bros" here, there are plenty of "coding is solved"-bros (they're also annoying btw)
edit: grammar
That's the other one I see too. "Software factories" is relatively new...maybe just the past 3-6 months or so.
It feels worse if a robocar kills a human, even if statistically humans kill more humans than robocars.
I'm glad humans are freaking out when they realize they don't know if spending $10k/month on AI tokens is good or bad.
I'm glad humans are freaking out when an AI pumps out vapid presentations and other humans go ahead and present it to clients.
But.
Too many businesses have tolerated the same mindlessness when humans were in the place of LLMs.
Too many companies telling themselves and investors headcount growth is good without knowing that the new hires are actually doing.
Too many human-slop presentations float around, with the authors and audience just going through the motions.
Did digital photography raise the bar for what constitutes commercially valuable photos? I think so.
I hope AI will similarly raise the bar across all the industries it is touching.
most busineeses have no idea how much work is to be done at any point even before ai.
Most of the work i've done in my career has been some random shit no one cared about.
Disagree with this because we have ways to steer price use per task.
1. choose a good model
2. choose the appropriate reasoning effort
3. choose a prompt to nudge it even further
Then it comes down to understanding the intuition of what kind of task deserves what effort?
That "intuition" is not intuitive. How do I choose an appropriate reasoning effort?
I use these models all day long, experiment, and have no clue how to choose that.
It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems "how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on" though it's certainly not a guarantee on anything.
These models are being released at breakneck speed even before their creators know how to use them. It's a big project of collective discovery to figure out what they are doing and how to use them.
Yeah this I agree because even personally all my rubrics break when a new model is released. Things were stable while I was using 5.6 GPT. There ought to be some room to explore and understand this intuition and yes one must account for this bugdet.
Now get the 1000 engineers in your company to be as well-behaved and mindful as you are.
(Couldn't even get this to happen in a 30 person team...)
[flagged]
[flagged]
[flagged]
[dead]
A budget of zero is remarkably easy to maintain!