Sign in or Join the community to continue

Why Cost Per Million Tokens Is A Useless KPI?

Posted Sep 14, 2026 | Views 6
# AI FinOps
# Agentic AI
# AI Infrastructure
Share

Speakers

user's Avatar
Kuntal Patel
Cloud FinOps Leader @ Palo Alto Networks

Kuntal Patel is a seasoned technology professional with experience at Palo Alto Networks, Gigamon, Dell, and Force10. He specializes in cloud solution architecture across AWS, GCP, and Azure, with a strong focus on CloudOps and FinOps. He works on optimizing cloud infrastructure for efficiency and cost-effectiveness while collaborating effectively across diverse, cross-functional teams. Kuntal is also passionate about staying current with emerging technology trends and applying them to real-world challenges.

+ Read More
user's Avatar
Abhinav Lad
Director - Cloud Finance and AI @ Palo Alto Networks

Abhinav Lad is a data and FinOps professional with 9+ years of experience solving business problems through analytics, storytelling, and financial optimization. He has 8+ years of experience in cloud FinOps, specializing in multi-cloud financial modeling, cost optimization, infrastructure savings, scenario planning, and variance analysis. Abhinav has helped organizations make more informed and cost-effective cloud decisions, including identifying underutilized resources and optimizing commitments such as RIs, CUDs, Savings Plans, and FLEX CUDs, contributing to more than $20M in incremental annual savings. He also has expertise in enterprise data warehousing, data analytics, and dashboard reporting.

+ Read More
user's Avatar
Alex Salkever
VP - Content and Research @ The Linux Foundation

Alex Salkever is a leading expert in exploring the intersection of technology, business, and society, with over two decades of experience covering cutting-edge advancements in a wide assortment of fields such as AI (and ChatGPT), green energy, genetic engineering, cloud computing, virtual reality, and self-driving cars. As a former editor of BusinessWeek and an award-winning author, Alex has a unique perspective on the ways in which technology impacts our lives and well-being. Based in the heart of Silicon Valley, Alex has firsthand access to emerging technologies at the forefront of development and adoption. He regularly engages with researchers and innovators working on over-the-horizon ideas that will shape the future.

+ Read More

SUMMARY

A year ago, Palo Alto Networks built dashboards to track AI spend. Today those dashboards are useless, and the team that built them thinks that's the whole story.

Recorded at FinOps X in San Diego, this conversation brings together Abhinav Lad, who leads cloud and AI finance at Palo Alto Networks, and Kuntal Patel, who runs the cloud engineering function behind it. They explain what happened when agents entered the picture, and AI stopped behaving like a service anyone could forecast.

+ Read More

CONTENT & TRANSCRIPT

Kuntal Patel: [00:00:00] AI initially when we started, it was just one of the services like compute, your storage, and your analytics. What has changed dramatically since the agentics comes into the mix, we were all forecasting AI as a linearity of usage-based consumptions. But when this agentic model start picking up, the, it becomes a exponential growth.
Alex Salkever: I guess a challenge with agents is that they're not humans.
Abhinav Lad: Yeah.
Alex Salkever: And they, they can- Yeah ... I mean, they're at, uh, potentially endless consumption.
Abhinav Lad: Yeah.
Alex Salkever: Uh, you know, paperclip maximizer if someone does the wrong thing.
Abhinav Lad: Yeah.
Alex Salkever: Uh, and obviously the gateway is designed to- Yep ... to stop some of that. Yeah. Yeah. But, but how do you model for these strange events that are not as easy to bucket as you had at, at Electronic Arts?
Kuntal Patel: As a company, we are spending $2 per million token. What does it mean? It doesn't mean anything.[00:01:00]
Alex Salkever: Greetings from FinOpsX in San Diego. I'm here with Abhinav Lall and Kuntal... How do I say your last name? Forgive me. Patel. Patel. Okay, forgive me. Um, who are from Palo Alto Networks. We're gonna be talking about FinOps and AI. Abhinav, tell me a little bit about what you do at Palo Alto Networks with Kuntal, and you guys can start to bring us up to speed a little bit.
Abhinav Lad: So I'm Abhinav. Uh, I'm the director of cloud and AI finance at Palo Alto Networks. I manage and oversee cloud cost optimization from the financial side. Uh, apart from that, I do have, uh, business partnering and, uh, forecasting for some of our portfolios.
Kuntal Patel: I'm Kuntal. I've been with Palo Alto for almost seven years, and in my current role, I am leading the cloud engineering FinOps functions, where we build the FinOps from all the way from data visibility to dat- uh, optimizations in the Palo Alto.
Alex Salkever: So tell me a little bit [00:02:00] about- the FinOps, uh, approach to AI at Palo Alto. What, what, what's your philosophy and what have you built? How do you think about this topic? 'Cause it's moving so fast.
Abhinav Lad: Yeah. The-- You just said it, right? It's moving very fast. And, uh, we started our journey o- of building some visibility into AI about like, a little bit over a year ago.
Abhinav Lad: And what we built a year ago is no longer valid anymore due to the pace of AI, and you could just call it like the dashboard's not of use anymore. Uh, we had like some level of visibility built initially. Uh, it was more like, okay, well, you can think of it as another service and you can have a cost by like GPUs and TPUs.
Abhinav Lad: Like previously it was like VMs, and now you have a GPUs and TPUs. And within kind of-- we, we saw that as a more powerful machines and a lot of similar FinOps practice were applied initially, where you can just go buy, you know, [00:03:00] cards or RIs for those as well and, you know, reduce their cost. Uh, and then we saw the landscape changing a lot and the different use cases kind of evolve over time.
Abhinav Lad: Uh, fast-forward today, we built a lot of different tooling with the help of Kuntal's team as well. And, uh, what we really wanted to kind of-- what we really saw the landscape evolve over time is the different use cases of AI with the different teams. Like one of the use cases you can think of is within the product.
Abhinav Lad: Other use case is,
Abhinav Lad: you know, developers' productivity where, uh, aka wipe coding, uh, you wanna call it. Uh, and then a lot of internal enterprise functions that are using to build like a, a knowledge base to help, uh, uh, help, help our internal functions kind of answer the question. For example-
Alex Salkever: Finance, creative marketing.
Abhinav Lad: Yeah, yeah, creative marketing and, and things like that.
Abhinav Lad: So that's like a different landscape and different business we kind of, you know, saw using AI [00:04:00] in a, in a different way.
Alex Salkever: So I'm curious, uh, what metrics changed or how you changed what you measured and what you cared about over that year. And I'm also very curious how you wired this up and then you had-- how you had to-- how you've had to change the way you observe or how you're treating spans.
Alex Salkever: Uh, you know, from an engineering problem, it, it has changed and evolved quite a, quite a bit in 12 months.
Kuntal Patel: Yeah. Yeah. Uh, if I may share more light into how we measured this is like AI initially when we started, it was just one of the services like compute, your storage, and your analytics. What has changed dramatically over like October, November since the agentics comes into the mix, right?
Kuntal Patel: We were all forecasting AI as a linearity of usage-based consumptions. But when this agentic model start picking up, the-- it becomes a exponential growth. And the moment it becomes a [00:05:00] exponential growth, the way we were tracking it, it's no longer valid. We as a team, we, we try to do fail fast. So every other week, we try to see is this use case still valid or not valid, and kind of build it, rebuild it, and then generate it.
Kuntal Patel: In terms of the AI as a consumption, we are still relying on the cloud to consume the LLMs as a models, uh, whether we are using AWS or Azure GCP, and we are using third-party models, first-party models from the different cloud. Now, the challenge part here is building the data pipeline, so that brings the visibility, that brings the dashboards, and eventually the agents that sitting on top of data and answering those questions.
Kuntal Patel: So what we have seen the use cases, uh, where our code assist as an example, using a Cursor cloud or Antigravity, right? It's driven heavily by user. But when we started building this, uh, Opus and more of a thinking models, they are very agentic. [00:06:00] And because they are very agentic, it runs in a autonomous, and it, it has a agency to run it in a agentic loop, where agent decides how much, uh, iterations I need to do to achieve the task.
Kuntal Patel: How do you measure? How do you track that? We s-still believe we are, we are, we have solved the basic visibility problem, where we can tracking it at the historical data. We yet to solve the in-the-air real-time cost visibility for AI, and we have a plan for that. We, we have architecture where we need to go and implement it, but there is a path forward, there is a roadmap forward that we wanted to build around that.
Kuntal Patel: And
Alex Salkever: From both the financial planning and engineering side, I'd like to, uh, dive deeper on how agentic really scrambled things. You said it went to exponential. That y- clearly sort of token consumption, uh, exponential. Also many more use [00:07:00] cases and y- a- and the idea of AI consuming AI just as AI, uh, you know, which hadn't been around before, uh, you know, which we're seeing more and more now of.
Alex Salkever: Uh, so I'm curious how sort of you shifted your thinking once you realized that this is no longer a linear service that's gonna be consumed in a easily predictable fashion.
Kuntal Patel: I can, I c- I c- I can take that, uh, thing. So if you think, uh, agentic AI, it has, uh, three main characteristics. One, uh, is, uh, autonomous.
Kuntal Patel: So now this agentic AI are goal-oriented, outcome-oriented. They are no longer the task-oriented. You give them a goal, it will plan, it will execute the tool, it will verify, it will stay in a loop until it reaches the outcome. Now, with that property, if any of the step fails, they would try further until it reaches outcome.
Kuntal Patel: So now those are the iterations to get to the u- to the goal [00:08:00] And we pay for that. We are paying for iterations that agents try to run to achieve the task. Now, behind the scene, this seems like all the LLM calls, but you are paying a lot more, uh, rip-- downstream rippling effect on your enterprise data.
Kuntal Patel: Because every LLM call, you need to retrieve your enterprise data as a retrieval RAG, and now you are paying for the memory as well. So not only you are paying for LLM, uh, but you're also paying for a RAG, and together, I think we, we have a very exponential usage of token consumptions and, uh, chunking to embedding to vector store to retrieval, and that adds up in the cost.
Alex Salkever: And transit cost, I assume too, if you're moving, let's say, out of Vertex into- The egress, yes ... some other cloud or something like that.
Kuntal Patel: Yeah. So you, you think about this egress cost, uh, uh, into the mix as well. Yes.
Alex Salkever: So, so it's [00:09:00] not just the, the tokens of the agents, it's all the other things that the agents are doing.
Alex Salkever: How did you shift your financial model to accommodate this? It's, it's-- I mean, there's so many more moving parts.
Abhinav Lad: Yeah. So now from the financial lens of what Kuntal just talked about, right? Like previously, a lot of cost that we saw was part of product. Like products were building chatbots that are, you know, more kind of helping external user.
Abhinav Lad: They can talk to the chatbot and things like that using Vertex AI. Now, in the agentic era when it increased exponentially, it's been used by code assist where a lot of developers' productivity has been driven, right? And internally to track, we built more visibility into cost per user. Think about how much cost AI has been costing per user.
Alex Salkever: And is that role-based as well too, I'm assuming? Or
Abhinav Lad: So right
Alex Salkever: now- Persona, persona-based more accurately ... it's not specifically persona-based, but we definitely observe the usage is more on persona-based. Mm-hmm.[00:10:00]
Abhinav Lad: Okay. So there was like-
Alex Salkever: Like your marketing team wouldn't use as much as your development team or-
Kuntal Patel: Yeah.
Abhinav Lad: Yeah, or like product team is not using as much as engineering team, for example. Right. Right? So like, so that's kind of the visibility that his team has been super impressive in terms of building as well, in terms of vision.
Abhinav Lad: Um, so the, the visibility that we have is, okay, what's your adoption rate? For example, some, some team may have a higher adoption rate. They may have more engineering in their team, and that's why their adoption rate is high. So if your adoption rate, for example, is 80% today, well, would-- do you see it going to 100% tomorrow?
Abhinav Lad: That's one angle to l- look at it. The second angle to look at it is average cost per user. Suppose your average cost per user, again, just throwing it out there, can be 100 or 1,000 or 10,000. It could be anything. What do you see, being an engineering leader, the right cost is for your wh- whole engineering team that, hey, I want to assign X dollars for my engineering team.
Abhinav Lad: That's the budget that [00:11:00] I can hold accountable for. And then from there to start to kind of set limits in terms of you could use the latest model, but then hey, you can just burn your total cost soon enough, or you could use the other models and it may take more time for it to burn. So kind of applying the limits on AI usage on a weekly basis or a monthly basis as the engineering leads see the fit.
Abhinav Lad: And then they have our dashboards available for their team to look at, like what's the adoption rate, how much cost per user that they have. So it's kind of shifting the model from looking at like just the forecasting on a linear level or by growth to like cost per user to adoption rate.
Alex Salkever: And as you've made that shift, have you also started to think about baking in quality of outcome or quality of work, sort of value cognition?
Alex Salkever: 'Cause not all tokens consumed are equally- Equal ... valued to the enterprise, right? [00:12:00] Yes.
Kuntal Patel: So i-i-if you think in this journey, first you need to know what's the cost of AI. Now AI- The big cost ... big cost of AI. Now, when you solve that, that problem, they want to break it down by LLM, RAG, and other ways. Once you have that thing, very obvious question, what's the value of that, right?
Kuntal Patel: And this is exactly where, uh, as a, as a company where small group of teams are going towards the DORA metric. Like that's a pretty standard format in the industry that give us an anchor. How do you want to find the value from the KPIs that quantifies the DORA, right? Everybody can understands now, uh, code assist is generating the line of code.
Kuntal Patel: Now you check in the code into GitHub, how many codes that have been generated and submitted by, uh, AI. What is important? What's the value of that? Does it mean I am shipping feature [00:13:00] faster than its, uh, dead- deadline date, or am I pushing feature faster to the QA cycle and QA cycle is kind of taking a longer time to process it and go into the cycle.
Kuntal Patel: So end-to-end cycle and end-to-end product release, what is the value of that? And I think starting in that journey, DORA could be the one anchor that we can think about it, and then we are in a, in, in a way to pilot it for some group of team and then anchoring p- bigger role in that.
Alex Salkever: And, and so when, when we talk about DORA, um, obviously that started more out of sort of the SRE DevOps world.
Alex Salkever: Um- How do we modify it appropriately, or does it map perfectly to the world of AI value? Uh, 'cause it's a l- SRE DevOps is a little different.
Kuntal Patel: Correct. So we need to see, i-i-it doesn't fit exactly, uh, as is in any company, in any, any team as a use case, right? You have to take it this as a guiding principle, see what are the [00:14:00] rules and the what are the business context that you want to apply it for your team, for your product as you are a multi-product organizations, and fine-tune it, right?
Kuntal Patel: So one of the things that I'm part of the IT organizations and what IT organization is trying to do, the agents that we have deployed, the agents that we have it in a production, does it reduce the operational and support cost because now we don't have to answer the basic L1, L2, L3 kind of a support cases?
Kuntal Patel: Now, can we use that metric and see is there a value of number of hours that we saved? So that's the one way to find the value realization. Now, second is like, can we think AI itself as a first use case for the greenfield opportunities? Most of the things right now what's happening in this world is, can we take the existing use case and make it as a AI convertible?
Kuntal Patel: But we are as a some of the new use case anyway, can we think AI as a, AI as a first enabler [00:15:00] rather than build something translating into the AI. So that avenue, I think, is very interesting where in this journey, uh, we would evolve into.
Alex Salkever: So, Abhinav, one thing that is, uh, an important job of finance, uh, and a delicate one, is how do you steer- uh, people towards doing the right thing with the appropriate tools?
Alex Salkever: And part of that's engineering, but part of that's also sort of, you know, un- understanding h- the, uh, how to nudge peop- is it, I mean, I am curious whether that's part of what... 'Cause obviously there's a lot of, at this point you're probably seeing already, like, Opus 4.8 or whatever, I mean, Fable is complete overkill for probably 60% of the work that's being done.
Alex Salkever: Um, and there's even another tranche below it where probably, at least for the devs, there's, or, or even for marketing people, there's a lot of things they're doing with AI which they could do with, uh, some Python [00:16:00] libraries quite well on your local network with, you know, that would cost nothing. So, I mean, this is cost optimization, but getting people to use the right tool for the right job.
Alex Salkever: How, how do, how are you, you starting to think about doing that?
Abhinav Lad: So when we think about, like, different models first, right? Like, do you always have to use the latest model? That question comes, right? Because some of the teams, you know, they have a option to go choose the model that they want to choose. In some cases, some engineering leaders are limiting that, "Hey, the latest models are very expensive.
Abhinav Lad: Uh, uh, it's okay to just not use that and use the second best model as well available out there." Now, one of the things that's yet to be proven, I guess, at least from finance lens, but we definitely want to think about and explore more, is if you're not using the latest model, would you end up using, consuming more tokens?
Abhinav Lad: Which can end up more cost because- Sure ... it now requires more thinking and you are [00:17:00] end up using more tokens. And would that be solved with using the latest model because it's more smart? It's like, would you hire an analyst and senior analyst or one senior, one senior analyst or two analyst and, you know, see which ones can do better.
Abhinav Lad: So it's, it's that problem that we are right now and, and it's too early really for us to really answer that question. But yeah, we currently when we think about like managing the financial, it's more on the really the budget limits that, hey, these are the budgets and you will have to kind of stay l- in within the limits on, uh, what you want to spend, and it's as long it's kind of working for P&L.
Alex Salkever: Yeah. No, I understand. As a
Abhinav Lad: whole, so yeah.
Alex Salkever: W- when it sounds like you aren't explicitly routing, but you're... Are you starting to steer people?
Kuntal Patel: So as we are looking into the AI, somebody's starting with the unmanaged cost and everybody's now starting optimization, right? That's a journey. What [00:18:00] is critical here is we have a visibility, which is a historical.
Kuntal Patel: It happened, now you know the cost. You have to be in real time. What can solve this real time is AI gateway. So if you think you have a-- So that, that's where we are, we are going into the implementation or kind of a next evolutions of our architecture is you have your AI agents, you have your AI coding tools, right?
Kuntal Patel: Uh, and now they are directly talking to your LLMs through a different form factors, right? We need AI gateway in between. So Palo Alto's Prisma AI's AI gateway can solve these problems where It give us, uh, three unique solutions to govern the whole AI in a thing. One, every AI transaction is captured, so you have a full observability.
Kuntal Patel: Second, you verify the identity of those agents. Who is using it and wh- are they, uh, do they have a agency to take an [00:19:00] action on behalf of this identity? And the third important thing is can we enforce real-time policies? Now, we think about policy, policy could be the cost policies, and that's where the model routing, model fallback, disaster recovery, open models and other models will come.
Kuntal Patel: But any organization should start with what are their cost policies in AI, and what are your security policies? You cannot have a two silos of security separate governance and cost as a separate gov. They have to come together to solve this problem. So AI gateway into the mix, you get the observability, you be in line, you get hopefully fine nines and less than a millisecond latency, and that would solve this problem.
Kuntal Patel: And we are piloting towards that implementation.
Alex Salkever: You said you're piloting the tool called Prisma, is that correct? Or, or, or what, what... which AI gateway are you using Aniket?
Kuntal Patel: Prisma, uh, Palo Alto's Prisma AI gateway.
Alex Salkever: Oh, okay.
Kuntal Patel: So we-
Alex Salkever: You built your own AI gateway?
Kuntal Patel: Palo [00:20:00] Alto built its own product, yes.
Alex Salkever: Oh, cool. Is it open source?
Kuntal Patel: Uh, it's an enterprise, uh- Oh, okay, okay, okay ... Palo Alto
Alex Salkever: offering. Okay, yeah, got it.
Kuntal Patel: Um, and the... we being part of IT, IT is always the first customer of Palo Alto products. Sure. And recently they announced this, uh, Prisma AI gateway.
Alex Salkever: Mm-hmm.
Kuntal Patel: So why use a open source ga- open source gateway?
Alex Salkever: No, if you, if you have it in your product, sure.
Alex Salkever: Um, s- so in your financial models of all this too, what new, uh, uh, factors have you put in? Uh, I mean, obviously agents, but there's all these... You know, before it was just like how much does it cost per service per user? Now there's a whole bunch of other metrics I'm assuming you're starting to bake in.
Alex Salkever: Yeah. Yeah. What does the model look like differently now?
Abhinav Lad: Yeah, model is not one size fit all in, in the AI era. For example, like when we talk about engineering productivity just now, right? Right. Like it's kind of [00:21:00] more on cost per user. But then when you think about AI being used in the product, it's more kind of looking at what's the top line growth.
Abhinav Lad: If the revenue is growing and AI spend's been growing in line, well, that's okay, because now we are capturing Like previously we were capturing cloud cost as a percentage of revenue. Now we are also capturing AI cost as a percentage of revenue. So that's kind of one model for in-product. Now when you think about internal enterprise use cases, right?
Abhinav Lad: Where just as an example, I'm submitting a phone bill and why do a human needs to approve it if it's same every month?
Alex Salkever: Doesn't.
Abhinav Lad: It, it doesn't require anything. It's something that, you know, agent should, should just go approve it. It should only flag when there is a problem, that they notice some anomaly or something like that, right?
Abhinav Lad: So just as an example are like QBRs, the decks [00:22:00] that we have. If it's using the same spreadsheet every month, same format, why can't it just go build the decks for us? Uh, so yeah, these are like some of the examples, right? So when AI has been used kind of an internal function or it's kind of used an internal knowledge base, uh, it kind of makes me think more as like you first need to see how much, how many users you have, how much volume you have.
Abhinav Lad: So users will then help dictate the volume, and then volume will then dictate your AI cost. So it's like three different approaches for three different use cases, uh, that we really need to go start thinking about and using. Kind of reminded me of, uh, my days, uh, with my previous job when I was working for Electronic Arts.
Abhinav Lad: It's a, it's a gaming company.
Alex Salkever: Yeah.
Abhinav Lad: And when you really have to think about cloud forecasting for a gaming, you first think of a user forecast.
Alex Salkever: Sure.
Abhinav Lad: And from a user volume- And it's unpredictable. And it's unpredictable. But then when you're launching a new feature or launching a [00:23:00] new internal function to enterprise, then, uh, enter- uh, your employees then employees will have a lot of question on, "Oh, what is this?
Abhinav Lad: How does this tool work?" And you will see kind of more volume. Or if you're launching a new game, you'll see a lot of volume of a new users trying out those games. So it's, it's a correlation there as well. Uh, and more the volume, more the AI cost as well. But then over time it will kind of settle down. So yeah, it's kind of three different, four different models to kind of think about when it comes to AI cost.
Alex Salkever: How do you bake in, uh some of the more unpredictable edge cases. Uh, I mean, I know in games you, you have a bit of that where if something outperforms or underperforms.
Abhinav Lad: Yeah.
Alex Salkever: Uh, I guess a challenge with agents is that they're not humans.
Abhinav Lad: Yeah.
Alex Salkever: And they, they can- Yeah ... I mean, they're at, uh, potentially endless consumption.
Abhinav Lad: Yeah.
Alex Salkever: Uh, you know, paperclip maximizer if someone does the wrong thing.
Abhinav Lad: Yeah.
Alex Salkever: Uh, and obviously the gateway is designed to- Yep ... [00:24:00] to stop some of that. Yeah. Yep. But, but, but like, you know, how, how, how do you model for these strange events that are, you know, not as easy to bucket as you had at Electronic Arts?
Abhinav Lad: I guess there's-- The edge cases are always hard in any case.
Abhinav Lad: Sure. You know, it's regardless of the company, regardless of... It's the challenge that every forecasting or finance team faces today. If there is, uh, one solution where, you know, you can always predict that there will be an edge case, well, you know, that's the perfect model that can have- Sure. Might as well
Abhinav Lad: solve this. Yeah. Start a
Alex Salkever: calcium market
Abhinav Lad: on your edge case. Yeah. Yeah. Yeah, absolutely. But I, I do think like what we, uh, do today when we think about cloud, right? Like we, we detect anomalies, and as soon as the anomaly's been detected, we send it out to the users to see, "Hey, this has been happening," and why there is a certain spike in the nature.
Abhinav Lad: Is it kind of something that was expected, was not expected? So they can go acknowledge. I [00:25:00] do see something down the line in the future when Kuntal was mentioning that, again, when we think about our journey, we are not there yet in terms of like real-time visibility today. Uh, we are building over time the governance, and, uh, I guess AI Gateway can help us solve that problem eventually where it can f- kind of loop back any anomalies back to the user.
Abhinav Lad: As, as he was explaining, agent can go into agentic loop, where it's keep on trying. So your even failed request is keep on trying again and again and again, and burning more tokens for you at the end of the day. So it, it can end up costing as
Kuntal Patel: well. Yeah. So in that case- Yeah ... right, I think that problem, the model is the one part of problem, but then we have to see why it had happened and how can we make those occurrences less and less.
Kuntal Patel: So one of the features, if you can think about this, is how soon we detect this. [00:26:00] I-if it becomes, uh, erroneous, right? Like, uh, most of this func- one of the functionality is you can keep max retries. You, you don't have to keep it at infinite. Now, with that functionality, you can say, "Hey, you know what? These are my dev agents, these are my stage agents, these are my pre-prod agents, and these are my prod agents."
Kuntal Patel: And every agents has a different characteristics. You build a profile around it, you build a different retry mechanisms, fallback mechanisms, and cap it. So you have to cap your environments, uh, while you are deploying it. So that brings some kind of a predictability i-into the, uh, AI as a FinOps for AI. The second part is you always want your security gate, like not only AI, if there is a more of, uh, any cost variance happens because you deploy these agents in some kind of a cloud environments.
Kuntal Patel: Mm-hmm. Whether it's an agent engine, whether it's a Cloud Run, Cloud Functions, or [00:27:00] Kubernetes in a way, you wanted to make sure your other data searches and the regs are also not spiking up, not only the agents. So you have to build the anomaly detections holistically, and then use this data to correlate it and say, "Is it really the good, bad?"
Kuntal Patel: But you don't want to code it after one month and say your model break. Within the 24-hour window is a good, good, uh, trade-off, uh, to find the sanity between finance. But most of these problems can be solved while we are deploying it, but setting the right policies, the, at the design phase itself. And
Alex Salkever: f-for that, um, how has design phase changed for applications given these new capabilities and constraints?
Kuntal Patel: So i-it's a learning itself. There is no proofread policies and the design documents that everybody started it, right? Uh, we started without the gateway, uh, maybe three, four months [00:28:00] back. Now we are thinking how can we deploy with the gate-gateway and write the right policies. So the mindset is there. And as this agents goes into production, the production has to scale out, so that's a scaling into the, uh, into the mix as well.
Kuntal Patel: So it keeps those in the mind, but there is a definitely a different policies are coming up, what we call it as a dev stage pre-prod environment versus a prod environment, and there has to be the different guardrails and the designs to keep into the mix.
Alex Salkever: Something that w- some of the other practitioners have told me, which was interesting, is that, uh, in some ways FinOps has become almost an extension of product QA, where on the product side anomalies can often be an indication of something being wrong with the product Um, and I wasn't sure whether you've experienced that or, uh, you know, and it, it can be just either the product is poorly designed or it could be that the, the users are particularly confused with the feature and keep [00:29:00] prompting over and over again or, you know, there's a bunch of scenarios though where the, the bill tells you what f- the problem.
Alex Salkever: Uh, it-- I was curious if that's something you started to think about or something you started to explore or
Kuntal Patel: If you look into the engineering lens point of view, right, it, it's, uh, partially true for a non-production or R&D kind of an environments where you are experimenting a lot and you don't know the what's the impact of those experimentations.
Kuntal Patel: And that's where most of the abnormal behavior will be picked up. If something has been deployed in a productions, it has a right security DevOps practices put it into practice, and it has a-- whenever there is a anomaly or ab- abnormal expectation happens, there is a valid reason to justify it. Here in non-productions, team is exploring it, team is experimenting it, and suddenly they don't understand the cost impact.
Kuntal Patel: So you find lot of, uh, unexpected variance there and trying to mitigate it. But the idea or the thought would be can we catch as early as possible, [00:30:00] collaborate with the team, and mitigate it in a quick time to respond. So i- if we have a good culture set up there, uh, I think we want a room to grow, we, we want a room to experiment with that.
Abhinav Lad: To, to add on what, what he just said, right? Like, in the non-production, obviously we do see anomalies, right? But then again, good or bad anomalies. Is it expected? Is it not expected? They may be doing some testing and that anomaly is, yeah, we, we know this is expected. This is because why we're doing it. Or hey, you know, it's caused by an error and we thank you for, for letting us know.
Abhinav Lad: We'll, we'll go fix it right away. So as, as you mentioned, right? Like, quick fix is the key. So when you have a spillover, well, you don't have a days of spillover, right? How quickly you can resolve that spillover is, is the key here. So real-time detection is where, you know, we want it to be e-eventually to see, you know, how we can detect it super early and send that notification to the user, and user can then make the judgment call in terms of like, [00:31:00] is this really an anomaly or is this something that they were expecting anyway because of some X product testing that they've been working on?
Alex Salkever: Related to this, uh, I'm curious How you communicate FinOps volume, FinOps changes around AI up the chain to C-suite for potentially non-technical folks. Uh, you know, 'cause part of why I ask that is that two years ago it was important. Three years ago-- I mean, now it's potentially existential if you get it wrong.
Alex Salkever: Yeah, yeah. You could blow a quarter-
Abhinav Lad: Yeah ...
Alex Salkever: if, if something goes wrong.
Abhinav Lad: Yeah.
Alex Salkever: Yeah. How do you interact with sort of the upper level folks and, and-- or maybe that's not the wrong question, but I'm sort of curious, how does the organization start to absorb these lessons and, and behave differently?
Abhinav Lad: Right. Right.
Abhinav Lad: No, I think, uh, when I really think about, like, how... It, it kind of... If you, if you kind of [00:32:00] look back on, like, cloud cost as well, right? Like, how do we really show the cloud cost to a leader? We use certain metrics to show, uh, the cloud cost. The same thing for AI also applies. It has almost become an extension that now every monthly calls with the leaders that we run with, where we provide them the visibility, we just don't provide just the cloud cost.
Abhinav Lad: But we also- On AI. There's a separate line item for AI and, and help them understand what, how much their org has been spending on. And again, some of the metrics that I talked about is what's their adoption looks like, what's the cost per user looks like. Is this what they anticipated to be, or is this something that they're planning to be at X dollars per user?
Abhinav Lad: I mean, there, there are obviously power users as well, so you can then kind of, you know, try to carve out some story out of it, like similar way you do it in the cloud. That, [00:33:00] "Hey, X service spent more dollars," or like, "Hey, we, we know these happened because of an error. This has been fixed. Now this is no longer a problem."
Abhinav Lad: The same thing is like, oh, we saw this user kind of, you know, spending more and, you know, that was expected. This is confirmed by their engineering team, and it just was a one-time cost. But then moving forward, we expect it to kind of, you know, uh, grow or remain same or go down or things like that. So it's kind of almost become extra line item for us to really think about as a separate forecast, as a, as a separate line item to discuss with the engineering leaders.
Alex Salkever: On, on this one, and, and this is sort of, I think applies to both of you. Uh, we touched on it earlier and, and I believe, Kuntal, you said so you're applying DORA metrics to sort of value, um, in, uh, of, of AI usage. That works really well for technical use cases, developers, code, some DevOps. [00:34:00] But he- I would say probably 50% of your business activity is not there, and I'm curious, and I asked this of a bunch of folks, how do you assign value to AI consumption for, uh, tasks that are, um, you know- harder to evaluate.
Alex Salkever: So obviously like, you know, filling out an email for somebody that's automated, yeah, that, that's easy to read. Uh, but how do you evaluate the value of AI for the marketing team or for the sales team or, uh, you know, where, where, where it's not as easy to at- attribute to the AI the outcomes?
Kuntal Patel: So we, we ha- I, I, I have to put it in this context, right?
Kuntal Patel: What we've been talking is about like, oh, everybody's building an agents using an AI. But there is a bigger space in a no-code, low-code space where you are using, uh, like, uh, AI [00:35:00] tools as a SaaS, and then you are consuming it, like for email summarizations coming from the Google, like a G Suite, right? Right.
Kuntal Patel: Those are also AI tools, right? Like web scripts and Google Sheet has something, Google Doc has a generations, right? So those are low-code, no-code or medium-code kind of, uh, sweet spots where business users or business enterprise consumer is gonna consuming it and doing it. Now, if we think about whether it's a pro code, low code, they all gives us a list of KPIs.
Kuntal Patel: The KPIs might be different, and that KPIs drives towards the capability and the value out of that. So no code would have a different set of KPIs towards the value, but ultimately we wanted to see, are we increasing as a productivity dimensions? Are we increasing as a operational operational ability dimensions, or are we increasing as a collaborative dimension?
Kuntal Patel: So are we saving hours? Either we are re-releasing more features, so we have a bigger capacity to build more, [00:36:00] right? So those where we are looking at an angle. But in that space, right, I think it's a journey. Everybody's figuring it out. Nobody has find their own script so as we, right? But it's, uh... we'll, we'll see in this journey where we, we go with that.
Abhinav Lad: Yeah. As Kuntal explained, right? Like AI's been evolving so fast, and we've been trying to keep up as well with it in terms of like, like everybody, right? I, I think we are, we are at the governance stage right now and moving, marching, like crawl, walk, and run phase, right? We, we call it in FinOps journey, right?
Abhinav Lad: I think we are somewhere between, uh, you know, walk and run phase, where like there is governance and then there is value realization phase, uh, that yet to be figured out, uh, as well. Completely we do have some ideas in terms of like how do we really think about it. Is it really like delivering now features super fast or, [00:37:00] you know, it's boosting, you know, productivity in terms of how much you can do?
Abhinav Lad: Now you can get a lot more done, right? So it's kind of- yet to be figured out. Uh, and again, as I said, like with the, with the pace of AI, there is a lot of like we, we don't know what we don't know really. No. Yeah. So one thing, uh- So we still have to kind of think about like, and learn from even practitioners too, like what, how they're doing, and we are exploring that as well to see what are the learnings that we can take home.
Abhinav Lad: So.
Kuntal Patel: One thing, uh, I, I wanted to add that like we have been seeing a lot of, uh, industry trend, and that's, this is my perspective. We should not think a cost per million token as a outcome. Because the categorizations of token as a cash input, cash read, cash write output drastically difference in terms of the unit cost.
Kuntal Patel: So if we normalize it and use that KPI to drive the [00:38:00] value, it can give us a completely wild, wild waste direction. So the perspective of us or me looking into, hey, do I want it to use cost per token? And hey, as a company we are spending, you know, two, $2 per million token. What does it mean? It doesn't mean anything.
Kuntal Patel: It has to put in a perspective of your persona use case and then c-create a metric. Considering KPI as a token, uh, cost per token, I think it's a very outlier or maybe, uh, I might be wrong in that phase, but that's my current, uh, observation so far.
Alex Salkever: Well, I wanna thank you both for joining us today, um, and look forward to seeing you at future FinOps conferences.
Kuntal Patel: Thank you for having us here. We appreciate it. Yeah.
Abhinav Lad: Thank you. Thank you for having us. Really appreciate it.
+ Read More

Watch More

Why is MLOps Hard in an Enterprise?
Posted May 30, 2023 | Views 897
# Enterprise Organizations
# Standardization
# Ahold Delhaize
# Aholddelhaize.com
RECOMMENDER SYSTEM: Why They Update Models 100 Times a Day
Posted Sep 15, 2022 | Views 1.4K
# FunCorp
# Recommender Systems
# A/B Testing
Machine Learning Operations — What is it and Why Do We Need It?
Posted Dec 14, 2022 | Views 890
# Machine Learning Systems Communication
# Budgeting ML Productions
# Return of Investment
# IBM