Sign in or Join the community to continue

Why Your AI Bill Will Double Before It Gets Better

Posted Aug 03, 2026 | Views 38
# AI Agents
# Coding Agents
# Agentic AI
Share

Speakers

user's Avatar
Josh Collier
FinOps Lead @ Superhuman

Josh Collier is the FinOps Lead at Superhuman (formerly Grammarly), where he was an inaugural member of the FinOps practice and now leads the team. He helps organizations maximize the value of their cloud, SaaS, and AI investments through a combination of analytical expertise and strong communication skills. At Superhuman, he built automated, customized AWS reporting and established the company's first reporting and capacity management processes for generative AI.

Josh is also an active thought leader in the FinOps community. He has spoken at multiple FinOps Foundation events, including FinOps X, and has contributed to several AI Working Group publications.

+ Read More
user's Avatar
Demetrios Brinkmann
Chief Happiness Engineer @ MLOps Community

At the moment Demetrios is immersing himself in Machine Learning by interviewing experts from around the world in the weekly MLOps.community meetups. Demetrios is constantly learning and engaging in new activities to get uncomfortable and learn from his mistakes. He tries to bring creativity into every aspect of his life, whether that be analyzing the best paths forward, overcoming obstacles, or building lego houses with his daughter.

+ Read More

SUMMARY

In this episode, we're joined by Josh Collier, FinOps Lead at Superhuman (formerly Grammarly), to explore what it really costs to run AI at scale and why the rules of the game changed faster than anyone expected.

+ Read More

TRANSCRIPT

Josh Collier: [00:00:00] Something we were talking about a little bit earlier around OpenAI's release of guaranteed capacity

Demetrios: That just happened, right?

Josh Collier: Um, I find it really interesting because it's not advertised as a discount mechanism. It's, it's advertised as a, uh, you know, guaranteed capacity, you know? And, and that's a way

Demetrios: For- which is saying, like, if you need this, you're always gonna have it, so you're never gonna be without resources.

Demetrios: We're not gonna go down on you.

Josh Collier: Exactly. But from our standpoint, we haven't really seen capacity issues, so it's almost like they're, they're solving a problem that doesn't exist yet.

Demetrios: That has to be very hard for you to look at, and it probably makes your job, like you have a fixed price you can't go over.

Josh Collier: I mean, there's lots of ways to address that. At the same time, you, you see the subscription model going away and, you know, usage base, or a combination of that, those things happening, and that's the reason why, because it's-

Demetrios: Yeah ...

Josh Collier: y- fixed prices just don't work anymore.[00:01:00]

Demetrios: All right. I'm here with Josh, working at Superhuman, leading the FinOps. I wanna hear about your journey into tokenomics

Josh Collier: Uh, it's been about a three-year journey overall. But, uh, but yeah, prior to, uh, Superhuman, and, um, at the time when I joined them, they were Grammarly. Uh, I was at AWS and, uh, you know, had no exposure to AI when I was really there outside of just managing some of the billing aspects of it.

Josh Collier: Come to Grammarly and, you know, we have some hosted models that I was helping, you know, distribute the costs and, and report on, and then, uh, in early 2023, we started using this new thing called Azure OpenAI. Um, so I started kinda poking around. I'm like, "I think this is gonna be kinda big, so I wanna understand, like, if there's a area where I can help."

Josh Collier: Um, so I started, uh, kind of poking around, asking questions, and found, oh, this is like a fixed capacity that we're, you know, we wanna make sure we're utilizing enough of it and, and, you know, uh, [00:02:00] also understanding, oh, if we don't have enough, there's, this will create an outage for some of our products, so we gotta...

Josh Collier: So it, it introduced some new concepts, like I'd never been involved with making decisions that could actually cause an outage on one of our products, right? That wasn't like a FinOps thing, but it kind of fell into my scope. So, but as I started doing that, I realized after going to some, um, FinOps Foundation events that I was one of the only ones do- doing that type of work back then.

Josh Collier: Um, so, uh, certainly encourage anyone to, you know, even, uh, definitely in this day and time with AI, just, uh, most practitioners, if you poke around hard enough, you'll find a area where you can drive a lot of impact for sure and help.

Demetrios: Well, it's funny how you started poking around and then you realized, "Wait a minute, resource management, this kinda seems like my job in this whole new field."

Josh Collier: Yeah, resource management, a very, very expensive resource. Yeah.

Demetrios: Especially in those days. Yeah, maybe talk to me about the trends and how you saw the trends dropping and then now reversing of the trend- Yes ... of the cost. [00:03:00]

Josh Collier: And you look at me a year or two ago, I'm like, "Oh, costs are just falling off a cliff, and now it's just gonna be the, the vendors fighting for market share, and it's gonna be race to the bottom."

Josh Collier: Um, so yeah, in the first probably year or two, saw over a 80% drop in cost per token. Um, and I think some of that was the vendors getting more mature in the infrastructure behind the scenes, um, and also them trying to gain, grab market share.

Demetrios: Yeah.

Josh Collier: Uh, and then over the, I'll say starting in, you know, really last year, but over that following year from 2024 to 2025, the trend started to change.

Josh Collier: And, um, and it, whether it started where it was not, you know, maybe they were about the same to now with, you know, for example, GPT-5 for is like double the cost of 5.1.

Demetrios: Sure.

Josh Collier: Um, I think Opus just came out today, and it's like double the cost.

Demetrios: Mm.

Josh Collier: On top of that, you've got other SKUs like web search and, and, um, [00:04:00] reasoning and o- these other aspects that are just compounding the cost.

Josh Collier: So now we're s- the trends are, is, you know, not only are, are, is volume like a hockey stick right now But costs are rising too, no matter like what is being said elsewhere. Rates are going up for sure.

Demetrios: R- and how did the way that you were doing your job back when you thought in the next six months, this is just trending to zero, versus now where you're kind of looking at it and you're saying, "It may go up, but it definitely is not going down"?

Josh Collier: Uh, it's gonna just, it's, it's gonna exponentially go up. Uh, it's, uh-

Demetrios: That's your prediction. You're calling it here.

Josh Collier: It's, uh-

Demetrios: Only going up.

Josh Collier: It's only going-- Volume's not gonna go down.

Demetrios: Uh-huh.

Josh Collier: If it's not going up, then as a business, then, you know, that's a bad for us, right? So, you know- Mm ... and ideally, we want more and more users using the product, but we just want it to scale efficiently.

Josh Collier: Mm. But, uh, you know, I, I would say with external LLMs, that's, that's, it's [00:05:00] becoming, in my opinion, a, a, a larger and larger financial risk because of the, the rate trends that we're seeing here. So volume we know is going up. We can't have pricing also doubling just about every time too, and then that's where I'm looking at exponential increases.

Demetrios: Oh, fascinating. So that's making you nervous, looking at how it's been performing and recognizing that the volume's gonna continue to grow. Even if it stays at this price that we're at right now, that's scary. But y- there's a non-zero probability that we're looking at the prices also increasing.

Josh Collier: Yes.

Demetrios: And so that forces you now to look at different options?

Demetrios: How does-

Josh Collier: So the,

Demetrios: the- ... your outlook change? ...

Josh Collier: the bulk of, of Grammarly's, um, suggestions are based on open source models.

Demetrios: Mm-hmm.

Josh Collier: And that is where I really want us to go- Mm ... um, just from a cost aspect, you know, but we're also an AI company, so we have a special set of teams and skills that can, can make that happen.

Josh Collier: [00:06:00] Most companies don't have that. But, uh, I would say right now, I mean, we, we have two sides of this AI coin. We have our production cost, which far outweighs our AI coding internal cost. I would say m- right now, most companies are dealing with the AI coding side of it, and that's probably what they should be dealing with.

Josh Collier: But for us, it's like both sides we've got to really manage and make sure and, um, that, that both are efficient and Measuring the value in some way.

Demetrios: Yeah, because the product is powered by LLMs.

Josh Collier: Yes. Again, mostly open source, but more and more, um, as you want to have very high velocity of releasing product and such, I mean, those frontier models are, are great for that.

Josh Collier: So that's where the, the growth is happening on the, the external LLM side.

Demetrios: Mm-hmm. And you also mentioned to me before we hit record on the hidden costs around the price increases just with the models. Can you talk a bit more about that?

Josh Collier: Oh, yeah. So when you are looking at, say, uh, you know, maybe you're using GPT 5.1 for whatever reason, [00:07:00] and, um, and 5.4 comes out, and let's say you have data residency, US data residency requirements.

Josh Collier: So now you, you see the rates, and you start using it, and you're like, you know, "Well, we do wanna use this, and we see that it's two times the cost, but we see the value here." But then you s- don't notice the fact that there's a 10% data residency charge on top of that that didn't exist with the 5.1 model, and that's happening with the 4, Opus 4, I think 4.6 or 4.7 and later.

Josh Collier: Um, and, um, basically Anthropic and OpenAI are both adding this premium to US, um, based models. You were already seeing this before this with the hyperscalers, um, AWS and GCP and such, but now it's, uh, that's another premium you're paying that a lot of, I believe a lot of companies are kinda overlooking.

Demetrios: Yeah, you don't realize that until you later need it, and then you go, "Oh, we're getting charged for that?"

Josh Collier: Yeah. So whenever you're doing your estimations and comparing these models, just make sure you look at the fine print about the, this data residency piece, 'cause, yeah, 10% adds up.

Demetrios: Yeah. And speaking [00:08:00] of estimations, I know that you've built a cool cost calculator.

Demetrios: You vibed it.

Josh Collier: Uh, I vibed it. I mean, it-- So I'm not an engineer, and I was telling him before that, you know, before, uh, you know, I started at Gramrly, I had really no experience with AI at all. Um, and it was just my kind of poking around and curiosity that kind of put me where I am today. So, um, I was, uh, took the initiative to say, "Okay, I'm gonna space out some time in my days to spend time building some agents to help my finance practice."

Josh Collier: And I really, I knew I wanted to build some type of calculator and, um, didn't really know what it was gonna look like, and I built like a HTML version and then a, a CLI version, and it's using, uh, we have a LLM proxy that already has pricing in the back end of it, so it's just grabbing that information.

Josh Collier: And, um, I did it in about 15 minutes.

Demetrios: And people use it

Josh Collier: Uh, it's, it's pretty new, but I've had our, um, some people, a part of our evaluation team start playing around [00:09:00] with it, and they were very happy with it. And what this has done, you know, when you're evaluating, and I'm just learning this myself too, more about this process, um, for the workflow at least, is that they are building these prototypes, and at, up till this calculator was here, they didn't really have a good way to go and estimate costs.

Josh Collier: They would build the prototypes, get a sense of the maybe what the request per second will be and, uh, you know, different steps this agent's taking and the token types and all of... and token volume. And they would send it to me in a spreadsheet, and then I'd run the numbers and cal-estimate it for them. And now they can just hit the CLI, and it's gonna bring in, okay, for this model, this type of agent would cost this amount, and it's saved to the calculator.

Josh Collier: And then they bring in another model, do the same thing, save it to the calculator. And then they-- Now they understand before they even start doing an experiment that what the costs are gonna be and if it's even scalable. And she made the comment to me that there have been times where they did a prototype, went to do the experiment, it was too expensive, and had to start over again.

Josh Collier: So not only is, is this helping us keep costs down, [00:10:00] but also speeding up development velocity because they don't have to go back to the board. They know what the costs are gonna be before they get past the prototype typing stage.

Demetrios: Yeah, and that's just an example of how you recognized this is something that's happening.

Demetrios: Can I make it easier on the folks who are closer to this problem?

Josh Collier: How can we get these tooling-- I don't wanna be in a spreadsheet doing estimates anymore. That's not where my true value is. Yeah. My true value is, is getting, maybe building, thinking about these tools and building them and then putting them into people's hands that are actually doing the work to make the tool, the, the agents or what have you.

Josh Collier: And then, you know, again, shift left. Like get that in early- Shift as much ... as possible. I would love for the product people to start maybe playing around with this if they have the data to do it, right? So I wanna be as early in development cycle as possible, and this really helps with that. And honestly, it was FP&A, my counterpart in FP&A, hopefully she sees this, um, that, uh, you know, kinda said, "Hey, could we build some kinda calculator?"

Josh Collier: And I'm like, "You know, [00:11:00] I think I can. I don't think it'd be too bad." And, and then, you know, about a week later, I had this thing. So you know, it wasn't even totally like me saying I need to go do this. It was like, you know, collaboration with the finance saying, "Oh, I think we should, we start working toward this."

Josh Collier: And I-- that piqued my curiosity and felt like I could do it.

Demetrios: And I like how the evaluation team gets to now understand this, and that's probably one of the core metrics that they're looking at as they're evaluating if this is feasible.

Josh Collier: E-exactly. Uh, and then another piece of that is, um, support requests that go to our channel.

Josh Collier: Um, they can send in a spreadsheet or what have you if they, somebody wants, and then my support bot's gonna use my calculator I built to calculate that. And so I d- that's one thing I don't have to go and do anymore.

Demetrios: Wow. Okay. So tell me more about where you're thinking of your ability to add products or empower the engineering team, the rest of the organization with this knowledge because they're closer to those problems

Josh Collier: So I, I know where [00:12:00] the data is.

Josh Collier: I know w- how to kind of structure, like what things should properly cost and, and so that's kind of the magic I kinda have for myself, and I just need to take that, and I think for all fi- finance practitioners probably have the same knowledge, but they need to take that, put it into some kind of tooling so it's closer to the engineering and the development process, and then shift their focus up in the organizational and be focused more on strategy and, and, um, and actually just enabling The teams to make better decisions moving forward on the products, no matter, you know, even if it is gonna be something expensive that they have understanding, oh, you know, well, this is, we have a reasoning behind this because mar- you know, product market fit or what have you.

Josh Collier: So i- if, if everybody can move forward from engineering up to, you know, product to obviously the C-suite with making decisions on knowing what the costs are ahead of time, we're gonna be in a much better state as a business.

Demetrios: And if we looked back at, again, going through that journey that you've had of [00:13:00] when you first were encountering AI on Azure with OpenAI and then versus now and the maturity that you feel like you've been able to come into, can you explain some of the things that now you're like, "We could not live without XYZ"?

Demetrios: Or-

Josh Collier: I would say, uh, our telemetry streams.

Demetrios: Telemetry streams.

Josh Collier: Um, from, from, uh, my perspective. Yeah. Um, I built some telemetry streams. So we have a LLM proxy as home build that, um, that gives me the token volume by calling service and then a feature is another, you know, step below a calling service. So like for example, we have our main AI chat and there's a bunch of connectors for that and that's a bunch of features, thousands of features, all those connectors.

Josh Collier: So I take that information, pump it into our, our, um, our cost reporting platform which is a third party and, um, and then allocate costs based on the OpenAI bill. Huh. And I can do that for Bedrock and, and any other vendor that we're, we're bringing [00:14:00] into that, that tool. Um, and all it's doing is looking at the token type and then, you know, looking at that cost versus the token volume of each service and all these other metadata that's using it and then proportionally allocating it by token type service and all those things.

Josh Collier: Um, so without that we would have, we would not have visibility to actually what these agents cost.

Demetrios: And then are you looking at spikes and trying to figure out what's going on? Why did we just spend a lot of money here?

Josh Collier: Yes. Um, I've got agents that are starting to do that a little bit for me, but I would say our, we, we don't have a lot of spikes.

Josh Collier: Um, uh, it scales with business hours just about all our production stuff just scales with business hours now. We'll see a trend, all of a sudden, you know, cash tokens are going down and costs are going up. Like we'll see things like that we need to pick up, but it's very, very rare that we have like a big spike outside of, you know, research and use cases or experiments that you would expect that.

Josh Collier: But usually if it is a giant spike, it's usually from an experiment that maybe, you know, was a little higher than expected or, you know, [00:15:00] research that I knew about ahead of time.

Demetrios: Speaking of research, you had mentioned before we hit record about the research team and how they're Being empowered also by some of the stuff that you're doing.

Demetrios: Can you explain that?

Josh Collier: Um, well, yeah, I'm, I'm trying. I actually just today sent something to them about-- Sent them the calculator and said, "Hey, is this gonna be of use to you?" But no, they, um, I think what we're talking about is around optimization, where, and it was touched on in the keynote today too, and I thought it was a great, great presentation in saying, you know, optimization in the cloud was just, you know, optimizing resources and, you know, buying savings plans and RIs and things like that, that FinOps can handle or is just more straightforward.

Josh Collier: And it doesn't really impact the product at the end of the day. When, with LLMs, whenever you're changing, oh, okay, um, you know, caching rates or, um, trying to reduce output tokens and input tokens and all these things and trying to use cheaper models, it actually impacts the output and the user experience.

Josh Collier: Drastically. So it is so much more nuanced. So I have found much more success with optimization when it's research-led optimization, where our researchers are looking [00:16:00] into saying, "Okay, I think we could probably use maybe a mini model combined with a standard model and, um, not impact our latency," which latency is a very important aspect.

Josh Collier: Our, you know, all our stuff needs to be pretty much sub one second latency. Um, so You know, it's not gonna impact latency, and we're gonna-- s-still getting really good outputs and saving money. So they, they, you know, maybe they'll come out with a, the initial kind of launch of, of, of an agent, and then from there they look for ways to optimize it.

Josh Collier: Um, we do the same thing with our open source LLMs, and, you know, we launched speculative decoding last year, and it saved us a bunch of money on, on the, um, on the open source models. And we're just trying to kinda try to learn from those, those teams and optimizations they did there and how can we apply this to these external LLMs.

Josh Collier: And, and again, it's all usually the big, heavy optimizations are, you know, really started from the research side.

Demetrios: And the product research is what fascinates me the most on that because it's not like AI researchers. It's not, [00:17:00] it's... As you mentioned, the product is so affected by any little changes in the AI side of the house.

Demetrios: Yeah. Because you're in a very unique situation where your product is driven by AI at the end of the day, and so you have to make sure that that's the highest quality product you can put out. But also your main job is making sure that it, you're just not using, bringing a bazooka- Yeah ... every time you need to do the smallest little comma update or whatever.

Josh Collier: E-exactly. And we have data behind the decisions we've made, and we've, you know, and I think most companies are like this, but we've just locked the data. You know, right now we have the cost data. Now we need to better tie that to revenue and, and other, and, and experiments, and actually get experiment IDs into the, um, telemetry streams I was talking about.

Josh Collier: Then seeing what the cost of this experiment was- Mm ... and was it worth it or not. So like it's, you know, while we've been in this and I've been doing this work for three years now, [00:18:00] we're still, like, getting to that value phase is hard.

Demetrios: Mm.

Josh Collier: Um, especially as you acquire companies and different products and different measurements, but

Demetrios: yeah.

Demetrios: Yeah, bring them into the fold.

Josh Collier: Yeah.

Demetrios: When you say experiments, you mean doing small little rollouts on, or do you mean like- Live experiments ... behind closed-- Yeah, live experiments.

Josh Collier: We have a very mature experimentation platform where we're actual- we'll be running live experiments on, you know, usually A/B tests, um, to see, you know, how maybe a new, new change or new feature impacts, um- It could be, you know, product engagement type of rates.

Demetrios: Mm-hmm.

Josh Collier: Um, and, um, and that's how we kind of decide on, oh, should this thing go full scale, or do we need to go back to the drawing board?

Demetrios: And again, it's interesting to me because it is product type of features, but then inside of the core product, there's probably a lot of different experiments that you're running with all the models or with the ways that you're making those models useful.

Demetrios: It's [00:19:00] the same product for me as the end user. I'm still getting-- If it's, we're talking about Grammarly, I'm still getting suggestions on my grammar, but you behind the scenes are running these experiments that I may or may not know about.

Josh Collier: You may not know about, and all of a sudden you're engaging with it in a, you know, slightly more volume.

Josh Collier: Mm-hmm. And you may not even realize it, but then, you know, our, our people are picking up on, "Oh, okay, we got more, more, more, uh, accepts here- Yeah ... for this su-suggestion." Or, or they actually, you know, the clicks on this specific agent actually increased over time, or the amount of time m-turns they had with this agent increased.

Josh Collier: Like, little things like that that indicate, hey, this is actually making a meaningful difference for our customers.

Demetrios: And you wanna be able to, just to close the circle on this, as you mentioned before, when you have those experiments, recognize the cost of those experiments.

Josh Collier: Exactly. Tie that to the cost fluctuations that we're-- that I'm seeing on my side, 'cause right now it-- I'll see these spikes, when I do see spikes, it's usually around experiments, and I have to do a lot of research [00:20:00] or have, you know, a LLM go do it for me, um, to figure out what exactly is causing this.

Josh Collier: And even then, the correlation is light.

Demetrios: Mm.

Josh Collier: Um, it's really hard to, to find that correlation. So, you know, we've got, we got a couple teams collaborating to bring that data into our, our cost reporting, so we, we can actually measure the cost of ex- of those experiments. '

Demetrios: Cause I imagine the teams aren't saying like, "Hey, I'm gonna run this experiment now, like, just in case."

Demetrios: You hear-

Josh Collier: Sometimes. Sometimes- They'll

Demetrios: give you a heads up.

Josh Collier: I've been very blessed with, with, uh, good relationship with a lot of the, the ex- teams that do these experiments. So the larger ones, they usually come to me- Mm ... either to estimate what this is gonna cost or just kind of let me know, "Hey, we're gonna run this thing."

Josh Collier: So I know about the big ones, but then there are still small ones that are supposed to be small, but then they end up being a little bigger than expected, right? But, you know, overall, you know, th- there's really good communication between us.

Demetrios: So on those times where you see the spikes or when you're getting a heads up Is that something that now you're expecting folks to have a better idea [00:21:00] of because of the cost calculator that you created?

Josh Collier: Yeah, I mean, ideally they would, they wouldn't need, um, me to do those estimates beforehand, or at least they'll have an idea of what the costs are gonna be before they run the experiment. And I haven't built this into it yet, but I'd like to actually have that going into a doc, a coded doc, you know- Mm-hmm

Josh Collier: one of our companies filed, like, for whatever submissions that they have in this, this tool, and they can, I'm assuming they- I'll probably have something where they can actually choose, okay, this is something to submit, and then it goes to a doc where we see all the estimations in aggregate, and then can pick up on any red flags or at least do better planning.

Josh Collier: Okay, we know we're running this thing during this date or within the next few months. We usually don't have dates until, you know, pretty late into the process, and we can better plan on, you know, what our finances or, or costs are gonna be down the line.

Demetrios: And also compare it with the real costs to make sure that- Oh,

Josh Collier: yeah

Demetrios: hey, is our calculations, are we projecting-

Josh Collier: Is my calculator broken or is, you know, are we, you know, request rates are a little, [00:22:00] you know, off. Mm-hmm. Um, so yeah. That, that's true.

Demetrios: And what are the different inputs that you ask folks to give you on this calc- cost calculator?

Josh Collier: Um, so re- average request per second, um, or minute, um, and then average request size.

Josh Collier: So when I say request size, that's input tokens, output tokens, cached input. And these are all things that as they're kind of building the prototype and, and testing out the, the outputs that's coming, you know, they, they get those token counts from, as a response body from the, um, the, you know, OpenAI or Anthropic or whatever you're using.

Demetrios: And I wasn't clear when you were talking about how you have the research, the product research team Then they have this knowledge of we can do XYZ with these models, or we can cache this, or we can set up a different architecture that will be a better product experience. Does that then get propagated through the different teams so that whatever next team knows, [00:23:00] oh, I now have the ability to make my experiment much cheaper?

Josh Collier: We could be better about that. We could be better about advertising how teams are doing this. Well, one other issue there though is, you know, when you look at our mail product, the way they use LLMs is, is way different than, than the way, say, Grammarly uses LLMs or even our new, uh, Go product. So that is one kind of issue.

Josh Collier: Now, I think there should c- should be more sharing across the, the product or the business units and, and those products to... 'Cause I'm sure there's still learnings there, but again, you know, these acquisitions happened not that long ago. Yeah. So it's, it's gonna take time for those things to be shared, and I think it's probably part of my job, um, to make sure that those are broadened, broad, uh, kind of broadly, uh, told to the company so, um, that other people can take advantage of.

Josh Collier: But that's something I think we can improve.

Demetrios: Well, it does feel like what you're saying too around A lot of times [00:24:00] folks will try, and let's use the easy example of just caching certain results or certain tokens so that you can spend less, but then it ends up having a worse product experience some of the time.

Demetrios: And so you wanna make sure that the teams aren't using that type of architecture at the cost of, like, having less token cost, I guess.

Josh Collier: Oh, exactly. And I mean, uh, first and foremost is always the best product for our customers. Like, we, we're not going to sacrifice that, especially in this very competitive area that we're in.

Josh Collier: We're not gonna sacrifice that at all. Uh, but it's more about, okay, building the best product, but what are some of the things that, that are-- we can still deliver the best product but also save money on the back end.

Demetrios: Yeah.

Josh Collier: Um, so the team is just, because these researchers have much deeper understanding of these LLMs, they're able to go in and tweak those certain things or know the direction to go, and then they are in, in the position to go [00:25:00] in and evaluate the outputs and kinda go through that process to make sure it's all good.

Josh Collier: There's no way a FinOps person could, could do that. Um, even if they knew how to adjust the LLMs and make it cheaper, evaluating the outputs in a way a data scientist or a researcher or com- computational linguist as well, they- Mm. That's a whole 'nother-

Demetrios: Skill set ...

Josh Collier: skill set that- Yeah ... you know, FinOps people aren't gonna have.

Josh Collier: So that's what makes this a lot harder to do. I mean, I've also tried to use Claude to op-optimize Claude, um, and, uh, our, our, like, Claude API products and, and, um, it gives good advice, but it's usually still just not quite, quite there. We've, we've already, you know, tried it or the impact wasn't, you know, that large.

Josh Collier: Um, but that could be another avenue is just use AI to optimize AI.

Demetrios: Hmm. So tell me about how this Reserving capacity kind of cost benefit analysis goes in your head. Do you get cheaper tokens if you reserve more [00:26:00] capacity, and is that something you even look at, or is it something that has kind of gone out of fashion?

Josh Collier: Um, so early, early on, that was really the only, reserving capacity, um, Azure OpenAI at the time, PTUs, was the only way to get low consistent latency. And as I mentioned with, with Grammarly, um, you know, sub one second latency is just hugely important for our product. So yeah, we needed low consistent latency.

Josh Collier: So yeah, in the beginning I was managing capacity, having to buy capacity for the peak, and if we didn't, you know, we would scale beyond our, our peak and, and have an outage. Hmm. Um, but then over time, um, we did see token rates come down for that capacity, but it was a really difficult thing to manage. Um, one, the actual utilization of the capacity.

Josh Collier: You could get, you can get metrics to know what your utilization was, but understanding it beforehand was really hard because, you know, you, everything's measured in these units, abstracted units that Azure [00:27:00] measures them in, and depending on the request size or shape, um, that the, each agent is sending, that utilization's gonna be a little different.

Josh Collier: So knowing that ahead of time was tough. Our LLM proxy team actually built a way, they call it shadow traffic, where it would mimic the traffic of this thing that was gonna be coming out to see what would... and load test the actual capacity, um, and see what the, you know- peak tokens per minute were. Oh, wow.

Josh Collier: And then that was how we would estimate, okay, we need to add this much more capacity for this new thing that's about to launch, or, you know, we can actually reduce capacity because things have come down. So, um, that was back in that day, but it was just still very difficult to measure ahead of time, and it was a lot of this m- mathematical-

Demetrios: Sounds

Josh Collier: like a headache

Josh Collier: limbo and exhausting. Then, um, I wanna say it was like late 2024, early 2025, OpenAI came out with priority processing.

Demetrios: Mm.

Josh Collier: So this has, uh, late- latency SLAs, availability SLAs. It is double the [00:28:00] cost as far as from a rate perspective as the standard, um, rate, but again, we needed that low latency. What attracted me to it was I didn't have to manage capacity anymore because this is completely on demand.

Demetrios: Mm.

Josh Collier: And so it scales with our usage. Our usage scales with, with business hours. So then I did the numbers and looked at, um, okay, if we had the, um, even looked at if we had an ideal amount of reserve PTUs, and then also with our proxy sent some traffic to priority processing, like, well, you know, could that save us money?

Josh Collier: In general, just comparing PTUs, you know, fully using PTUs versus priority processing, and really it came down to, you know, you could see some savings with the minis and nanos of the world, but you couldn't, uh, it wasn't enough to, to need to manage that capacity. Um, with the standard size models, there was no savings even with using PTUs.

Josh Collier: Oh. And this is a, this is a year or so ago. It may have changed, but I doubt it. Um, so priority processing was a absolute no-brainer. Now, on top of just making my [00:29:00] life easier-

Demetrios: Yeah ...

Josh Collier: um, it also allowed, it makes much easier to allocate costs because again, it's just very difficult to understand how much of a single service is using of, of this PTU because, yeah, you could look at token volume overall, but tokens are not equal across agents, and I took...

Josh Collier: So it was just very complex. So this, uh, improved our visibility, saved us some money, at that time at least, and, um, and, and made my life a lot easier.

Demetrios: I can only imagine you doing this complex math, like, "What's the right blend of PDU or just straight token?" I have some pretty large Google

Josh Collier: Sheets out- Yeah

Josh Collier: kind of floating around in our

Demetrios: drives. And then you're like, "Wait, for quality of life, let's just do it on this." There- that should be one of the factors, too. Like, you don't have to do this crazy math anymore. You know what you're gonna get, and it also, if it's the same, more or less the same price, then it's a no-brainer.

Josh Collier: E- exactly, exactly. And then something we were talking about a little bit [00:30:00] earlier around OpenAI's release of guaranteed capacity. That

Demetrios: just happened, right?

Josh Collier: Um, I find it really interesting because it's not advertised as a discount mechanism. It's, it's advertised as a, uh, you know, guaranteed capacity. You know, and, and that's a

Demetrios: way for- Which is saying, like, if you need this, you're always gonna have it, so you're never gonna be without resources.

Demetrios: We're not gonna go down on you.

Josh Collier: Exactly. But from our standpoint, we haven't really seen capacity issues, so it's almost like they're, they're solving a problem that doesn't exist yet.

Demetrios: Or they're foreshadowing what's coming.

Josh Collier: So I, I'm very interested to see how this play out. It could also be first, a first step toward a savings plan or RI kind of construct, but, but yeah, nonetheless, I, I see a lot of financial risk in these external LLM providers and, and prices are gonna continue to rise, um, and/or they're gonna want multi-year commitments.

Demetrios: Mm-hmm. Because with the guaranteed capacity, you can't do that for a month, right?

Josh Collier: Oh, no, no. They're multi-year is, is what I believe that they're [00:31:00] looking for there

Demetrios: Yeah.

Josh Collier: Yeah. Which I don't know 100% for certain, but yeah, that's, that's my impression.

Demetrios: Yeah, that is something that now, again, you're thinking about it and you're starting to realize these cost-benefit analysis and saying, "Huh, if we're good enough now, we're never gonna change."

Demetrios: But then later down the line, if it starts to get a little shaky, now you have to recognize, is it worth it for us to get this guaranteed capacity? Or as you mentioned earlier, too, like, they're slipping in a lot of new costs, like the data residency. Now there's this new cost on this, and so they're almost upselling you in every way, shape, and form.

Josh Collier: Exactly, and I, I think it's really hard, from my perspective at least, to do any kind of multi-year commitment against, um, you know, in this AI world that we live in. That is... That's really hard, um, to look at- For any provider, yeah ... all the changes in the market and say, "Okay, we're gonna be using, [00:32:00] you know, said provider for three years."

Josh Collier: Yeah. Um, so it's just, it's too fast-paced. Um, so yeah. Uh, but there are people that are probably in position to, you know, be able to do that. Our proxy also gives us the ability to kind of use whomever we want to use- Mm-hmm ... which is, it's nice having a vendor agnostic, um, proxy for that.

+ Read More

Watch More

Evaluating AI Agents: Why It Matters and How We Do It // Annie Condon | Jeff Groom // Agents in Production 2025
Posted Jul 28, 2025 | Views 250
# Agents in Production
# Evaluating Agents
# Acre Security
Fine-Tuning is Broken: Why You're Doing It Wrong // Tanmay Chopra // AI in Production 2025
Posted Mar 24, 2025 | Views 134
# Fine-tuning LLMs
Fast & Asynchronous: Drift Your AI, Not Your GPU Bill // Artem Yushkovskiy
Posted Dec 10, 2025 | Views 61
# Agents in Production
# Prosus Group
# AI Drift
Code of Conduct
Your Privacy Choices