[00:00:00]
Whoa. What's up, dude? Oh, what's happening, Demetrius? How's it going, man? I'm good. Here we are. We made it. Yeah. Yeah, we got here. Well, bro, despite the looks of it, I am not being held against my will- ... in some random back room that is absolutely empty and blank. I am at the Agent Con event right now. We're setting up- Nice
for tomorrow. I wanna tell you a little bit about this event, though, that we're doing right now. And in the meantime, I see in the chat, like, if anybody is joining us, drop in where you're calling in from, where you're at, because it's cool to see folks from all over the world as we get involved and people get in to the event.
But man- I w- Bern, are you cool with me giving [00:01:00] you a little rundown on the inspiration for this virtual event? 'Cause this is the first time we've ever done a voice agents event. Yeah. Yeah, let's do it. Okay. So, m- the Agentic AI Foundation has a voice agents forum that's happening on November 5th in San Francisco.
But while we were organizing that in-person event, we thought, there's so much cool activity going on with voice agents right now. Let's check out what we can do virtually to get some of this activity and cool stuff happening beforehand so we can showcase, like, wow, what's going on here? What's going on there?
And that led me to speaking with Luke, the founder of Slang.ai, and he said, "Dude, I've got this cool thing that I'm actually coming out with really soon. Do you wanna do some [00:02:00] collaboration on it? Should we do something virtually?" And you know me, when you hear virtual event, I just jump. Yeah. Hence why we're here.
Awesome. Yes. And is, is the forum gonna be another virtual one, or is that in person? In... So the forum that you've got, this QR code you've got on your screen, that is in person, and that is going to be in San Francisco on November 5th. Should be a blast. We're gonna get speakers. It's gonna be one track.
We're really gonna focus in, and it's 400 people, I think, and we're really focusing in on what's going on with voice agents. And now speaking of that, I think you had some questions that you wanted to ask folks. Like- Yeah ... 'cause, yeah, there's some, there's some voice agent stuff happening out there. What's like...
What are people saying? Yeah, so I mean, I thought the first question that would be best to ask people is, like, you know, [00:03:00] who's, who's shipping, you know, voice agents to prod? Um, you know, who's- Yeah ... who's playing around or experimenting with them? And, you know, if, if nobody's, if people are joining just because they're interested, but maybe haven't done anything with voice agents, which, to be honest, is where I sit.
I haven't done anything with voice agents, but- Mm-hmm ... the, the big caveat there is yet, you know- Yeah ... I definitely plan to. Um, got some little, little home projects I think I'd try out. So yeah. Nice. I'd love to see in the chat if people, uh, if people wanted to... Or if you, if anybody's got, uh, voice agents in, in production that they've deployed, or if they're experimenting.
Yeah. Is it, like, demo? Is it- In progress type thing, or is it in production? Is that what we're going with? Staging, prod, demo? Should we do that? Exactly. Yeah, yeah. Let... Yeah, let people drop in the chat where you're at, staging, prod, demo. [00:04:00] I also realize that there may be a little hiccup if you're trying to write in the chat.
You have to fill out your profile before you can say anything in the chat. So hopefully folks can take a minute, let us know what they're thinking. We've got more questions, right? But while people are- Yeah ... dropping in that information, let's also do a little... Should we wait a sec, or should we, should I show you one of my favorite memes that came through the community recently?
I, I think, I think we should go, we should
jump over to the memes. Let's, let's take a look. So what do you got for us? All right, cool. So look, I'm sharing my screen. This was a meme that came through. Can you see that? Yeah. Yeah. You're, you're showing it. So what it says there is, "Yeah, we've sandboxed the agent," and then the agent is Neo having fun with all of his [00:05:00] friends, including Agent Smith.
Um... Yeah. What a, what a great photo. Out in the real world, right? You know, just like- Yeah ... not, not, uh- Not sandboxed at all ... not locked, not locked down, not sandboxed at all. Just, yeah, out there partying. Classy. That was a good one. I really liked that one too. Um- I know you're also very stoked on memes. Yeah.
What do you got for me as far as memes go? What's your favorite? So I dropped, I dropped this one actually recently. Uh, I think it might have been this morning. Uh, I thought this one was a lot of fun, um, uh, because, uh, uh, it, it also reminded me of a quote that somebody I heard once say. Said, "It was a, it was a mistake to teach sand how to think."
Um, and- Yeah. ... I thought that, that this went well with that, uh, that we've got runes that think, uh, which are just PCBs. Um, I got a good- That is so classic ... good chuckle out of that one. Yeah. That, that is so classic. [00:06:00] It took me a minute to understand it, but then I realized like, "Oh yeah, I guess that is what this is made out of.
Uh-huh." Yeah. Makes sense. That's where- Yeah ... our intelligence comes from these days. I, I also have, um, another one that I wanted to share, and then we'll get back to questioning the chat. So hopefully folks have dropped in the chat. Where, where is it? There was one that was... Freaking hilarious. This one Oh, yeah.
Oh, yeah Literally some of the things that I ask my agent, though, it's not too far off from this, where I'm like, "Oh, I could do it the old school way and the proper way," quote unquote proper, "or I could just, like, tell the agent to do it. Eh, fuck it. I'll just tell the agent." What's funny about this one...
What's the most funny about this is writing [00:07:00] what is our Git status- Yeah ... is longer than just writing Git status to see your actual status, which that's... Because, you know, devs, devs tend to be pretty lazy- Yeah ... they try to go the shortest route possible. So that, that cracked me up pretty, pretty good. Yeah.
That is a classic one. All right. So if anybody is out there, and they are able to drop in, come tell us what are you... Are you building agents? You building voice agents? You got 'em in prod? You got 'em in demo stage? Are you doing nothing? You're just kinda, like, agent curious right now? You're voice curious?
And... Or are you really, like, shipping many of them? Let us know in the chat. I think we can kinda get started, Berin. Where we at? I know there was a few other things, though, so keep me honest, 'cause I wanted to mention a few other things, and I usually get ahead of myself. I usually just wanna, like, jump straight into the content.
Yeah. So potentially I [00:08:00] forgot something. I think we've got, uh, one or two minutes, uh, on the agenda, so if you wanted to cover a couple things. Well, that's the thing, I don't know what I w-... I remember telling you, like, "Don't let me forget these things." Oh. And now I- So- ... don't remember what those things are.
Well, I think the biggest one would maybe be the, uh, the live stream that you've got going on, coming up soon. Oh. So yeah, tomorrow I'm gonna be live streaming from this Agent Con event for the next 36 hours, and that means that we're gonna be playing the actual talks, and then we're gonna, uh, cut to me, and I'm gonna be out on the expo hall.
I'm gonna also be doing, like, podcasts with the speakers. I'm gonna be doing podcasts with random guests of the event. And so it should be a good time. We can drop in... Uh, we dropped in the link in the chat, so if anybody wants to join me on that, that will be my [00:09:00] big thing for the next... And again, this is gonna be continuous stream for 36 hours, so you can come in- Stay for as long as you want.
I do not encourage you to stay the whole 36 hours, but if you do, there may be a
special kind of medal for you or a place in heaven for I don't know. Hopefully- A framed meme ... you are gonna... You won't be crossing any new dimensions. But that is, uh, that's what I'm doing, and that's a big one. You're right. We also have, like, just for anybody who doesn't know, we've got a ton of events coming up in the Agentic AI Foundation.
So we might wanna put those on the screen, too, 'cause it's, it's cool. Like AAIF.io/events can get you all the events we're doing. We've got Agent Con coming up in the US, if you're based in the US. We're also doing this one in, [00:10:00] uh, Amsterdam, which is a little late for you to come to, but we've got these forums that are happening, like this Voice Agent forum in San Francisco.
And if you're part of... If you enjoy doing meetups, we've got all kinds of meetups that are happening, or community events, as we like to call them. So we're doing community events around the globe. There are currently 86 active chapters. So come join a community event. There's probably one in your local city.
And those were all of the things that I wanted to say. I think we're, we're good now. Okay. Awesome. Well, thank you, Demetrios. And hey, man, like, I know you're at Agent Con, and thank you for, for, uh, hopping in and helping out with the, the intro, um, getting us kicked off here. Of course. Um, yeah, and I think the, the very first thing we're gonna be flipping over to is, uh, we've got, um, we've got Luke Miller and, uh, Niccolo Croon from, um, uh, Slang, um, with a announcement [00:11:00] that you were talking about earlier, um, how you were excited about.
I'm so excited for this. Yeah. So- Yeah, I am very excited for this. I'm gonna st- I'm gonna stick around and, and chat with them later after this talk, too, 'cause I wanna ask a few questions. All right. Sounds good. Well, I'm gonna... I think it's a video that I'm gonna, I'm gonna kick off, and then once, uh, once the video's done, I think we've got a little bit of Q&A after that.
So I'm gonna switch over to that- Brilliant ... and, and we'll get going. So enjoy. All right.
Every voice agent framework or platform today has its own definition of what an agent is and what it can do. Build on ElevenLabs, VAPI, Retail, or Volna, and your agent stays there. On the open frameworks, Pipecat and LiveKit, a Pipecat agent can't easily be ported over to LiveKit, and vice versa. You pick a stack, you build inside it, and your voice IP is [00:12:00] embedded to whichever one you started with.
This changes today Voice agents are a primary customer interface, but most companies don't actually own the IP that they're building. It is embedded in third-party platforms. On many platforms, an agent's behavior is scattered across prompts, framework code, provider integrations, and runtime decisions.
Change the stack and you rebuild it. Change the model and its behavior shifts. Change a compliance rule, and you have to update every deployment. That's why we built Unmute. I'm Luke, founder of Slang, and today we're launching Unmute as an open declarative standard for voice agents. At Slang, we run voice infrastructure across fifteen sovereign regions, primarily for regulated customers.
That work has shown us that three common assumptions about voice agents are wrong. First, voice agents are not a sequential turn-taking pipeline. They are a concurrent, multimodal, [00:13:00] real-time system. Listening, reasoning, retrieving, speaking, and interrupting can happen at once. So the question isn't which model do we use across the sequence, it is which resource should the agent use at any point in time.
The resource could be a frontier model, cached audio, a compliance-approved disclosure, or available compute in the right region. The agent should compose those resources around the caller, the use case, and the rules, without being rewritten when the infrastructure changes. Second, today's voice agents assume the LLM should be the decision-making spine of the call.
Tool calls, data flow, step transitions, failure handling, all left to the LLM at runtime. Unpredictable, expensive, and hard to govern. With Unmute, declared execution runs the call. Tool calls, data flow, step transitions, failure handling, all declared, all enforced by the compiled project. The LLM is called as an optional [00:14:00] resource, which earns its place, which means one package compiles to Pipecat Cloud, Livekick Cloud, or Slang, Cascade or speech-to-speech architectures in the region that you choose.
Third, today's voice agents assume you have to choose between rigid IVR style on one end or a detailed prompt followed by hope on the other. Two positions, both brittle at scale. Great demos fail in production is really this false choice bite in. We never specified what the agent was supposed to do. With Unmute, determinism is a dial set turn by turn.
Lock a regulated disclosure. Leave an open-ended objection generative. Same call, different levels of control. Non-determinism becomes a resource, not the foundation you spend the rest of the system trying to contain. That's what Unmute makes possible. The agent becomes the [00:15:00] artifact. Its behavior, its resources, its integrations, its compliant boundaries all expressed in one portable package.
Everything, the framework, the runtime, the cloud, the region is infrastructure. Compliance boundaries live in the package. Update them once, recompile. Every deployment inherits. Unmute compiles to Pipecat and Livekit frameworks. We have been inspired by both and my learnings from five years at Vercel.
Standards outlive platforms. Niko will now give you a more detailed introduction to what that looks like
I've been building agentic systems for a few years now, spending nearly all my time with customers and in open source. And what I see is always the same problem. Everyone builds an agent the same way: one prompt. Every time they need something new, they just add it at the bottom of it, and the demo looks great, but then it goes into production and it breaks.
And the reason is always the same, the prompt, [00:16:00] because the agent reads every rule at every turn. That is just too much. It goes off-topic, picks some rules and ignore the others. And when something goes wrong, nobody knows which part of it is causing the issue because nobody ever tried to mentally break down the prompt into the steps that the agent should go through.
What I learned is that, uh, if you want agents that work at scale, you go towards determinism, not by taking decisions away from the model, but by giving it less to decide at a, at a time. Split the work into steps. Give each step only the context it needs. At Slang, we have been looking for months at where the time goes at each turn and where the mistakes come from.
Both are in the same place, the LLM. So the question we work on is simple: how much can we take out of the LLM and still keep what makes it good? And here is where Amute comes. Amute is a declarative standard for building voice agents, so you only think about the parts that really matter, the steps the agent should take, what each [00:17:00] step can see, and what runs before the call, how long the agent can wait, and where are the models living.
You write that down once, and we build it for you. We map it to LiveKit, Pipecat, or we deploy it to Slang, and the good practices comes with it.
Here we have one agent for a salon built two times. Same tools, same voice, same phone numbers, same frameworks under it. The only difference is how the work is arranged. On the left side, we have one agent, one prompt, and every tool it may ever need, 10 of them, and the model reads all 10 descriptions at every turn.
And the prompt has all of it encoded, all the rules, all the edge cases that came out from real calls. Check who's calling, take a booking, handle complaint, answer questions about the salon. There is even a section in, in which whose only job is to tell the model which of the four steps are we at. That's the tell.
Four jobs in one prompt. All four go to the [00:18:00] model at every turn when the caller only needs one. On the right side, instead, we have the same salon, but with a split. We have the concierge agent, and we have the complaint specialist. The concierge is the one that actually talks to the caller, has one tool, look up salon info.
That is the tool that we only need. Everything else is delegated. And when we delegate, we drop the context and keep only the context that the steps actually need. And that is the whole declaration, what a step is, when to run it, and what it gives back. Here we see the verify customer step that confirms who are we talking to, and hence to the manage booking the customer phone number required for the booking step.
On the left- The confirmation is one line encoding to the prompt. Confirm who are you talking to before you book. You write it, and you just hope that it actually happens. On the right side instead, the booking step is not given to the model until the number actually exists. [00:19:00] So a booking without a confirmation number is not unlikely, it's just impossible.
And if we look at the prompt, the entry prompt on the right against the whole prompt on the left, we drop about 50%. What we did is just split it. This is the customer prompt, the, the, the en- entry point, a booking prompt, and a complete specialist that is completely segregated from the rest. And only one part repeats everywhere, the voice contract.
How the agent speaks, how it sounds, what it never says. That part has to be true on every turn, so it lives inside the prompt. Everything else stays inside the step that needs it. The booking prompt knows how to take the booking. It does not know the refund policy, and it does not need to. So what we will do now is compare the same conversation live.
One single prompt here on the left, one broken down into steps, and see what's changing behind the scenes. In real life, what you perceive is the skip of [00:20:00] a turn. Here on the left, the agent is asking the phone number, for example. On the right instead, the agent already has it because we prefetched it from the carrier before the greeting, so we skip the whole turn.
But in practice, when you look at the numbers, and here is Langfuse, one of the providers we use to measure how the agent perform across the different logics. On the right one, we see the salon concierge broken down into steps. On the left, the salon concierge as a single prompt. And the number of steps stays more or less the same, but the token usage is drastically higher on the left, the one single prompt.
More than double. And that is because, uh, the entire prompt is passed to the model at every turn, every time. Whereas on the right, we start at the entry point, we go down into the verification step, and we only pass the verification step into the prompt. Same conversation both times. Less token, fewer turns, and lower latency.[00:21:00]
And here is the part worth saying out loud. Put the four prompt together and you get more text than the single prompt on the left. Nearly double. But this is not the number that matters. The prompt is sent again on every single turn. So what you actually pay for is what sits in front of the model right now.
And someone asking what time we, we open should not pay for the refund policy.
Here is something that we see in nearly every agent that we look at. For the first 30 seconds of the call, and here we can see, the agent is asking things, uh, that it could have known before picking up. What is your appointment? What is your number? Are you a customer or not? Each one of these is a turn, and a turn with a tool is two model calls plus the tool call, not one.
So this entire block that we see here run all the three prefetchers, uh, before the greeting. Read the clock, take the number the carrier already gave us, [00:22:00] look at the number in our own system, and by the time that the agent says hello, it already knows the date, the number, and the name on the record. The order matters.
The first one is, uh, the one that people miss because a date tool looks like a real tool. It is not. A tool that takes no argument can only give you something, uh, you could have written in the prompt yourself. But now let's go to the interesting part, the second half, the one that you find more interesting.
Those are the variables, three of them. Each one declare once with a type and a default. Typed variables. And here is the thing about la- language models. It remembers nothing between turns. Nothing at all. Every turn you send is a prompt, the whole prompt again, and the whole conversation with new messages appended.
So what you put in front of it, uh, is not a detail, it's the cost of the call paid on every single turn. So we later declare where each value is allowed to go. Take the phone number, for [00:23:00] example, that we see in here. Uh, it tells you which phone called, not who is holding it. So somebody has to agree to it first.
That is the, the, the confirm point that we see in this step. And it means the number can appear exactly in one prompt, the step that reads it back. Not the greeting, not the booking step, nowhere else. If I paste the number into the greeting, I get a build error. Most importantly, when a tool needs that number, it does not ask the model for it.
We inject it. The model never types the number, so it cannot type it in a shape that does not match what the record we found. The idea is that if a step does not need a variable, you just don't declare it. If you need it, you say so.
Now, here's the whole point. I'm not going to design this agent. I'm going to describe the problem in one paragraph and let the coding agent build it. I'm going to install the new skill[00:24:00]
And that's it. The skill we just installed give it everything we need to know about a meal, so it can go and make decisions, uh, by itself. Let me start upload code
Pero pronto
And it works. Look at what it did. Figure out the right structure for us, and that's exactly what we want. Now let's actually run this thing. First up, we do the compile. Unmute compile. Compile takes the package and turns into running code that checks the agent.yaml, wires up the tools, and makes sure that the whole thing is actually valid before we ever run it.
And now Dev
Deepspin's have locally the agent, so I can just talk to it. Let's call it[00:25:00]
Thanks for calling Diamond. This is Ruby. What can I do for you? Hey, Ruby. How you doing? One paragraph in, working voice HTML. That's Unmute.
This is four moves, but in reality, one shift. Define the agent once, its behavior, tools, integrations, and policies, then compile it to the runtime you choose. Select your infrastructure, either your own, Pipecat Cloud, Livekit Cloud, or Slang, Cascade, or speech-to-speech architectures. Run it in the region that you need.
When your stack changes, your agent doesn't have to. That's the point. The agents remain your version-controlled IP, even when the infrastructure changes. Today, we're releasing the Unmute standard under an MIT license, alongside the open source CLI and its first three infrastructure targets, Pipecat Cloud, Livekit Cloud, and Slang.
Unmute is open. Take it, [00:26:00] build with it, and tell us what we missed
That was awesome. Dude, Nicola, first of all, congrats on such a great... And Luke, oh, we got you both here. Really well done on that video. It was top quality, so congrats on that. But then the standard, I was not expecting it to be like that. I had heard things. It's so cool to see this come together, and I have so many questions, guys.
So the first one probably is on this token usage part. Um, and feel free, there's other questions that are coming through the chat. I know that there's gonna be people that are asking, so I'll try and hit on those, uh, as they come through. But for my own [00:27:00] personal understanding of how this works, how does the token usage get reduced so drastically?
Like you were saying, Nicola, on the side, the side-by-side views, you had the left one with a ton of tokens. I didn't quite understand why on the right one. Is it because the, only the prompt, the relevant pieces of the prompt are being passed through? Correct. Yeah, we didn't reinvent the wheel, to be honest.
So we just leveraged what Pipecat and Livekit, uh, provide us. So what we did is, uh, we have a, like, quite vast experience dealing with customers and voice agents, and, uh, we sit down a few months ago and realized that we were just wasting tokens. Because, uh, when you pass the prompt down every time, because this is how most of the people build voice agents, one prompt with all the tools attached to it, and you have to take into account, uh, the tool description so that the LLM can actually fetch it, all the variables, uh, and the prompt itself in which we encode all the [00:28:00] tasks the agent has to go through.
And so we realized that, uh, Livekit and Pipecat were offering us, uh, a leverage with, that, uh, we could use to break this down and use, uh, on a, on Pipecat's side, we use Pipecat Flows. On Livekit, we use, uh, tasks and task groups And by using that, we, we just reduce drastically the token usage. This is, this is brilliant.
Now, the next question that I have, and I'm gonna look at- I'm, I'm gonna jump into this, Josh ... Oh. Yeah, yeah, jump. Feel free to. Sorry to jump ahead. Just, just to add in from what, uh, Niko said there, and this is a key thing, uh, reason why we launched, um, Unmute, and that is there's a lot of best practices, um, that are, um, operational in lots of the community right now who are deep in voice agents.
And as you were discussing before we, we launched the video, um, with Grotem, um, you know, there's, um, lots of people interested in using voice agents for the first time. They're technical, they wanna use this. So the whole thesis we [00:29:00] have around Unmute is making these, um, best practices built into templates, built into the framework- Mm-hmm
built into the standards. So we're trying to make all these, um, optimizations and improvements, um, available to everyone and anyone anywhere. And this is a whole bunch of them across the, across the, um, two, two main frameworks, across the implementations is what we're trying to facilitate. So any person can come in and say, "I wanna do this for the first time," not have the deep knowledge, but be able to have something that's extraordinarily efficient from, from their first deploy.
Well, well, that was the next idea that I was thinking about is how w- with these, basically saying we've got these variables now, and if I understood it correctly, you're just... First of all, the variables are helping you stay on track where you have the almost like type safety, right, with the variables. But then the variables can-- you can leapfrog certain steps in those turns so that you don't [00:30:00] have to frustrate the user.
If the voice agent should already know what the user is, name is, or if you see their caller ID type of thing, then you should have that information and any relevant information should be pulled up, and you don't need to have the agent ask the human these things. It's just maybe like confirming It's, it's a mix of components.
I'll, I'll let Nico jump in on the technical side, but it's... Overall, it's using these best practices for better context management throughout the call, and that context management sitting outside of the LLM. Um, basically giving better resources up to date at all times to give the best user experience, um, for, for the, for the end user.
But the developer who's building that, giving them more direct control of how they want it to behave throughout. But over to you, Nico, for the more technical component of that, that question. Yeah. Thanks, Luke. So the, yeah, as Luke said, the whole idea came from, uh, compressing [00:31:00] the context that we provide to the LLM.
And generally, when you start building voice agents across the task, across the agents, you share the full context history. So all the messages, all the tool calls, uh, and step by step, we try removing all those, all those additional information, and we realized that at the end, what we want to pass is just some compounding information we found, uh, like during the entire call that we can update real time and then pass into the prompt in the actual task that will needs it.
Ah, okay. So I'm, I, I see where you're coming from on that. Uh, and let me ask some from the chat, because I don't wanna be too, like, I don't wanna just hog this, although I do have so many more questions for you. This is so cool to see. There is a question that comes from Jin CJ. Um, real-time inference, what benefits does Slang provide for that?
So, uh, re- so real-time models is, is what they're outlining? Yeah, real-time [00:32:00] inference, I think it... Does, does Lang provide any, um, benefits for real-time inference? So there's, I guess there's a few different ways of interpreting the question, um, 'cause we... The voice calls are, are real time. The whole, you know, as we outline, the whole system is about, um, real time multimodal interactions rather than just the, the standard turn-taking pipeline.
Um, but we do a whole... Everything is from a, a real time context. If they're talking about real time models, um, what we... W- we're actually supporting this, and Nico, you, you probably have the updates around, um, how we're supporting the speech-to-speech now. Hmm. Yeah. So yeah, since yesterday actually, we just released speech-to-speech on both, uh, real time and live APIs.
Uh, so we, we see that our customers are asking for that and, uh, something that community is really excited about. Uh, um, but yeah, we, we mainly build on Cascade, to be honest, uh, since, uh, for complex use case, [00:33:00] so you, you want more control, right? You want to tune the control you always have. Now speaking of Cascade- I can just follow up.
Oh. Sorry. I'll, I'll just add one, one thing. Sorry to dump in again, Michelle, so- Yeah, no worries ... this is one of the key things we've seen with customers a lot of the time and, and just the community in general of everyone's like, "Where does, where does real time fit those with the Cascade architectures?"
Mm-hmm. Um, and there's benefits of both. Um, real time is often simplicity to get to a demo and start scaling it up a bit more, but you lose a sense of control. Cascade, there's many challenges with it. Uh, we think some of the best practices that we're sharing in Unmute overcomes lots of the challenges around token costs and latency, um, and a bunch of other tooling, but it's about which level of controls you want.
But what Unmute unleashes is more of that control that often isn't native in the real time models. We're starting to make those, those components of control available within the, the real time context. L- uh, Luke, just a heads-up. [00:34:00] I think that you're... we're picking up your mic from your headphones. Can you change your mic to your computer?
J- I think it'll be m- a little more clear. Uh, it might have switched over. Okay. So going back to this cascading, uh, fun stuff, I have heard a very common architecture being that you have almost like a graph, so you're passing in a prompt, but once you hit a certain stage, then you're passing in a new prompt because this conversation has gone and it has progressed in a certain way that you're trying to make it progress, and so now you can introduce new ideas.
Nicola, do you feel that, like, how does this combine with that type of architecture? Yeah. I think that's what we were referring to with, uh, tasks and the pipe cut flows. So, so what you do is switching the prompt over when you gather some information, take like the verified customer. Mm. You have a [00:35:00] very short prompt for verifying the customer because in the reality the LLM has o- only to ask for information about the customer, and then it switch- to a more complex task, but you progressively update, uh, the prompt and excluding the previous one so that you don't keep all this history.
Nice. All right, last one for you guys. Uh, Jin CJ's also asking, how do you secure personally identified PII in voice workflows? That's always a problem Iuk, you have to take that one. So the, uh, this is one of the foundations we're doing with, with Unmute of, um, you know, wherever you are in the world, there's lots of regulation coming.
You know, in the, the EU AI Act is one of the core pieces everyone's awad- um, aware of, but there's state level legislation coming in the States. Um, also across, across the rest of the world, there's different legislation, um, that's coming through, and this has different regulation on what disclosures are- Mm-hmm
um, but also more [00:36:00] specifically around the question, um, how you manage PII. And what we do at Slang a lot of the time, you know, the way we run our edge infrastructure for voice agents is often around holding back, um, uh, basically PII from going into the LLMs. A part of this setup that we have and the configuration and outline that Niko gave is about tokenizing any PII and ensuring you only intentionally pass it to any downstream model if you choose to do so.
But we're trying to flip the switch here of it used to be regurgitate everything into the LLM and let the LLM decide, and now we're saying by having specific variable control at the edge or in your, um, in your orchestrator, it means that you're not by default pushing it into all these... crossing all these different boundaries, which might be in different regions or, you know, more importantly, just a system boundary of, you know...
It's, it comes down to, um, controlling what your systems are exposing to third-party systems, whether that's because of regulation or just in- internal organizational policy. So a lot of the tooling, um, we're embedding, [00:37:00] um, in Unmute is about leveraging that, whether you use Slang or not. There's different components, but more broadly we're trying to make better compliance, um, awareness and defaults to be embedded in Unmute, whether you use, um, you know, Pipecat or Pipecat Cloud, Livekey Cloud or Slang or a mix of, of the above.
Mm-hmm. So this is one of the key things which we think, um, in the early days of voice, um, this all worked. We're like, "Okay, let's throw everything at the LLM." But the key shift here is you get a lot of efficiency by not throwing it all at the LLM, but you also get s- specific control of the PII side of things, and that's also part of the thesis of, of Slang, but it's embe- embedded in Unmute.
Um, and we're gonna have more and more tooling around this. You know, w- with the CLI you'll be soon be able to check, okay, we're deploying this to customers in the US. Does all the data, does all the compute associated with this associate with the US user so it doesn't cross broader boundaries? So we're- Mm
putting lots of the embedded in here of not only controlling the [00:38:00] variables that are passed, but controlling all the associated compute, um, and having this as a, as a cornerstone of how you actually build voice agents in Unmute itself. Um, and we think this is, you know, the cornerstone of, of where all organizations are going.
Um, we've all learnt that voice agents work. Now let's make them operational for the reality of, of what the world needs. S- so this, if I'm understanding it correctly, is the variables, you can put rules around them and these rules can help almost like M- make sure that it stays where it needs to stay because of some of these super strict data privacy issues.
Or just like you don't want data, A, going to the model, and B, if it's sensitive data, you just don't want that going anywhere. You wanna keep that and make sure that it's, it's secret. Ex- exactly. It's just been the default in this space has been, "Hey, let's make this work." Now it works. The default should be, let's only expose what's needed.
Yeah. 'Cause a lot of the time, even when you push a [00:39:00] request to the LLM, and even when you've got it very well scoped, you actually don't need every single bit of data. Also what we're doing is often separating out, you know, in the past, it used to be a giant blob of data that gets passed to the LLM, and what we're actually doing is other techniques of if you do have to send it, can we fragment the data so it's less easy to identify a single blob with associated data components as well.
So it's about if you do need to do it, how to fragment it so it's less, um, identifiable for, for that user. So we're, we're doing a lot more work around this, both at Slang, but within the Unmute framework to ensure that compliance and privacy becomes a, a norm for our, for our industry, and make these best practices fully accessible for anyone from day zero.
Well, this is, uh, this is super cool. The Unmute stuff, we just dropped it into the chat. There's one last question. I know that we've got the next speakers already ready, but it looks like Kevin just dropped, jumped in. Um, and I think [00:40:00] that there's some really cool stuff. I know Unmute just came out, like today, right?
So the question from Kevin is, "What's an example of a complex, complex agent that's being built using Unmute?" Uh, I think since it is so new, there isn't that complex one, or do you guys have some good examples you can give? Well, well, the examples are, um, Nico has been leading the engineering side of Unmute.
Mm-hmm. But internally, we've been dogfooding this for some time, as th- as we say in tech. So our onboarding of customers has been using this as well. So, um- Mm-hmm ... we're working with people with CX platforms in healthcare in the US. Um, we're supporting people in financial services in India and in Europe. Um, so we're doing lots of things where there's specific clients, um, uh, compliance boundaries.
There's com- specific disclosures you've gotta have in these processes. It's calling third party tools. Uh, more and more, um, we know that, uh, Livekit, CloudID, Pipecat have been doing more work around this as well. But our edge infrastructure also has, um, [00:41:00] uh, sandboxes where you can ex- execute specific code.
Um, so we're all building towards better infrastructure that is regional, that allows you to do, whether it's database calls, whether it's sandboxes, whether it's all these different interactions that can then reshape the conversation as you progress. So it's about taking stage control. It's about using data resources within the compliance boundaries.
So the, the sophistication is, of the agents that we're doing, um, comes from simplifying each step down, and we know that's the work that we've seen from, um, Livekit and, and Pipecat as well, and we look forward to making that very much available. But, um, lots of this stuff is, uh, stuff that we can, we can share more examples of.
And- Yeah ... um, if you follow, if you go on Star, what we'll ask people is go, to go on Star Unmute on, um, on GitHub, and we'll be re- And share it ... releasing a lot of, um... We'll, yeah, share it as well, but we'll be putting a lot of t- um, templates in there. So starter templates for all the best practices in all these different tool sets that you wanna have.
So whether you're thinking about, um, you [00:42:00] know, a financial services use case in Indonesia, we work with lots of clients down there, um, or, uh, or consumer health app in the UK, we're gonna have examples in there of how to get started quickly, and then using the tooling that Nico just outlined, being able to very quickly augment that to your exact use case and compliance boundaries.
Brilliant. Dude, well, I had one job, which was to keep us on time. I am failing horribly at that. So you're gonna have to forgive me. I'm gonna kick you both off the stage. Thank you guys for doing this. For everybody out there, check out Unmute, share it, do everything you can to get the word out, because I think this is great work and I'm excited.
I, I had heard whispers about it, as you mentioned to me, but I didn't realize it was gonna be like this, so... Now I am even more excited for our voice agent event in November, where we're going to be doing this in person and talking about it. Thanks, Luke and Nicola. We'll see you all, uh, in November in San Francisco Excellent.
Yeah. Thank you, Demetrios. Uh, it was really cool, uh, really cool, uh, [00:43:00] standard that they put together. I mean, and I shared the links in, uh, the chat for people, um, including the, uh, examples, uh, uh, for, for people to check out. So I think we're, we're, uh, about ready to get going on our next session. Mm-hmm. Uh, Demetrios, you sticking around or, uh...
I really want to. Uh, I might... I'm gonna stick around and watch it, but I'm not gonna jump on to ask questions 'cause I'm very- For sure ... bad at keeping us on time. I'm gonna try and keep us on time better. No worries. So I'll watch it, and then I'll be cheering from the chat. Awesome. Excellent. Well, thank you again, Demetrios.
And, uh, next up we've got, uh, Sankar, uh, Krishnan and, uh, Venkata, uh, Gopikolla. And I apologize if I pronounced your names incorrectly. Um, as somebody who has his name mispronounced all the time, I feel terrible, uh, when I butcher other people's names. But welcome in, guys. Um, yeah, uh, if you wanna go [00:44:00] ahead and give yourselves, uh, introductions and, um, and then get started, that'd be great.
Yeah. So yeah, um, now firstly, yes, you, you pronounce it correct, so my name, so thank you. And so to introduce myself, so yeah, as you said, like, so my name is Venkata Gopikolla, and, uh, I've been a, a software engineer at Salesforce, and I've been working for, uh, ten plus years, so close to 12 years out of college I joined.
And, uh, I've been working on, like, mostly on, uh, distributed systems and especially on the web traffic side, like all the HTTP traffic side. So basically, we are a web-based company, so. And my specialization is, uh, mostly on the caching side and also on the edge computing. So when I say edge, edge is the CDN.
The edge is the edge for the corporate traffic, uh, uh, web... corporate web traffic. So the beauty about edge is, like, no matter what kind of payload you have, like whether you have a AI thing or whether you have a traditional or anything, anything that goes by internet, [00:45:00] we are the first hop. CDN is the first hop, and we do control a lot more things.
And today I'll be showing some cool stuff aro- uh, with it. But with me we have, uh, I have Sankar, my friend. Uh, so Sankar, you wanna introduce yourself? Yeah. Hi. Um, thank you for having me. Very excited to be here. Uh, my name is Sankar Krishnan. Um, I've been in the AI space for about 15 years, and the last five to six years my focus has been on conversational AI and building voice agents.
Um, so I was at AWS, and my team used to build, uh, uh, voice models that is used by a wide variety of, uh, companies. And, uh, so through that experience I learned a lot about what works, uh, you know, what are some of the best practices, um, you know, what are things to keep in mind. Uh, so I'm gonna share a lot of those things today, so, uh, ex- and, and looking forward to it.
Awesome. Well, thank you. Um, if, uh, you guys have, uh, a screen share, you go ahead, get that started. Um, uh, you know, thanks for the intros and I, I... if I understand correctly, we're gonna be talking [00:46:00] about, uh, or you're gonna be talking about, um, how to actually deploy these systems, like real time, um, uh, what the infrastructure is, what kind of concerns around latency there are.
So really looking forward to seeing how voice agents, uh, should scale. So yeah, I'll, uh, let you guys go to it and, um, yeah, whenever you're ready. Awesome. Uh, Venkata, do you wanna share or do you want me to share? Oh, you can go ahead and share, Sankar. Okay. All right. Let me share. Oh, if you want I can also share.
Not a problem. So you, you can just- Yeah, I'm having some issues. Maybe can you give it a shot? Yeah, let me try. Okay. You can keep telling me whenever you want me to switch the slide. Yeah, yeah. I'll simply switch it
All right[00:47:00]
Great. Okay All right. Um, so I think the first, uh, the first piece to keep in mind as you're looking to build voice agents is what type of architecture you wanna use. Um, so predominantly in the industry, there are two types of architecture. One is more of a cascaded architecture, the other is an integrated architecture, and there are pros and cons to both.
Um, and so the cascaded architecture is basically you use a combination of a speech-to-text model, an LLM, and then a text-to-speech. Um, and, um, you wrap it around an agent, and that is, uh, that is driving all of the action. So in this case, uh, you know, you have more flexibility as in you can pick, uh, whatever speech-to-text, uh, te-text-to-speech LLM you wanna use.
Um, but then there is, um, there is a downside in terms of since the, the multiple hops here, uh, there's a impact on the latency. So your latency can be anywhere from eight hundred to, uh, twelve hundred [00:48:00] milliseconds, and this could vary based on which, uh, provider you're using. And then the flip side, you know, more recently, there've been a lot of providers who've been, uh, developing, uh, speech-to-speech models where all of this is integrated and, uh, with that you have a significant latency benefit, uh, which could be, you know, as low as, uh, less than two hundred milliseconds, and there is less information loss as well.
Uh, what I mean by that is sometimes when you go through a cascaded approach like speech-to-text, et cetera, some of the background, um, effects or, you know, the way how someone's enunciating a particular sentence, they may be sarcastic, et cetera. So there may be some information lost and that may not be passed through.
But as in a speech-to-speech model, that information is not lost, and it's able to understand those, uh, other aspects as well. So tho- so deciding on which approach to take is one of the first things, uh, that you look at and that has implications on your latency cost and kind of the output you get. Um, Mangat, uh, next slide, please.[00:49:00]
Okay. Um, the other thing to keep in mind is, you know, it's very easy to build a voice agent demo, um, but, uh, actually, uh, using it in production is a different ballgame. Um, that's because, you know, in real-time scenarios, you have a lot of issues where p- you have people talking, um, using various accents, you have people switching languages, um, you know, um, midway through a sentence, or you have multiple speaker voices coming in, a lot of background noise, um, et cetera.
So you have these kind of issues. And then, you know, being able to handle this at scale where, say, you know, let's assume you, you are an airline, and then suddenly you have a, a weather event and you, you, you're being bombarded by calls. So being able to handle all of these calls, um, you know, concurrently is, you know, can become a challenge unless you are, uh, you have developed architecture that can easily scale.
And then the last one is, you know, kind of, uh, also what I touched upon the earlier slide, you need to think about, you know, uh, several, uh, metrics that are more [00:50:00] constantly monitoring. This includes like your cost, latency, accuracy, which are all of your performance related metrics. And then also focusing on from architecture standpoint, reliability, scalability, et cetera.
And the last piece is around the having the right kind of safety and guardrails so that the voice agent is, you know, not just, uh, you know, being empathetic and talking in the right way to humans, but also, uh, ensuring that, you know, it's not taking any action that is not approved by the company. Uh, next slide, please.
Um, the... You know, one of the biggest frustrations that humans have with voice agents is, one, uh, you know, uh, the, the voice agent is not able to understand what the human is saying, or in some cases there's too much of a lag, and that frustrates, uh, humans. So optimizing for the latency is critical. So understanding what are some of the puts and takes as it relates to latency, um, you know, will help you to figure out, uh, the, the best approach.
Um, so Venkat is gonna talk about a lot of the network latency and [00:51:00] those pieces, uh, later in the, uh, presentation. But I wanna touch on some of the pieces that you need to think about, uh, as you, as you focus on the model that you're picking, uh, right? So, uh, there could be latencies in terms of how the model detects, uh, end of turn.
So what I mean by that is how, uh, someone is, um, you know, pausing after a sentence, right? So detecting those, uh, turns, conversation turns is important, and that may impact latency. The second one is, uh, you know, the, the, the, the work that is being done by the agent in, uh, in coordination with the LLM around pool calls or orchestration, getting data, um, you know, then, uh, the, uh, the, the networking aspect of it.
That is, your LLM is sitting in a different region or country versus, you know, where you're hosting your solution or application that's in a different country. So th- those all pieces can impact latency. And then the other piece is also how do you optimize for, you know, when does the agent respond back to the human?
Um, are you looking at, you know, the total response time and how, how much, [00:52:00] how efficient that is, or the first word coming out of the agent? So thinking through these aspects from a latency perspective is critical. Next slide, please Yeah. Um, I kind of touched upon earlier, but it's important to understand, um, uh, uh, the, the peaks that your production system will have, and this could be, you know, whether, uh, you know, uh, certain events, whether weather-related events, any promotional events, et cetera, and ensuring that you're able to have the right kind of scaling mechanisms, uh, ensuring that you're having the right kind of fallbacks.
Uh, a system goes down, you're having, uh, upper redundancy and, uh, you know, uh, uh, uh, trying, trying to, uh, you know, test this ahead of time, doing heavy load testing, uh, on, you know, a huge volume of calls, concurrent calls is, uh, important. So being able to like figure out what is your peak scenario and building for that would [00:53:00] ensure that, you know, you have a very, uh, scalable and reliable system.
Uh, next slide, please.
Yeah. The, the, the piece to keep in mind is, uh, you know, from an accuracy standpoint, um, you know, there is-- if you're doing a demo with very cr-crystal clear language, um, you know, and very, uh, easily understood terms, then it may perform well. But in real world, you know, people may, uh, talk about various company names or product names, and all of these may not be understood by the model.
So there's some level of training of the model that needs to be done. So the model needs to be grounded in company-specific information, whether it's products, uh, you know, whe-whether it's domains, whether it's country-specific information, so that it can pick up those words very c- uh, easily. So for example, if you're a healthcare company, um, you know, you need to make sure your models are trained on understanding a lot of the healthcare vocabulary, you know, uh, drug names or, uh, uh, or common terms used in, uh, [00:54:00] insurance, uh, processing, et cetera, because those could be mis-, uh, captured and, and that may impact the response given back to the human.
So being able to optimize the accuracy based on, you know, training the models is critical. Uh, next slide please. ense dave germany województ stre昆issement蜀—“ decentral通yleありがとう 이야 Mudah借助通充 primaries quer köszönaissez útbarer ovom roadside ince чомee
transfనుక申 Zan fruition amissuss[00:55:00]
Saranpta舞仍在nosti看一下yleistical trтному za
So what are some of the key things to keep in mind as you, uh, think about, uh, building voice agents in production? So, uh, the first one, uh, is the latency piece. Uh, so are you trying to figure out what the latency is for your P-fifty, uh, P-ninety-five, P-ninety-nine use case? So you need to think through that. Uh, you need to ensure that your system is able to scale or, uh, to about, you know, ten X or whatever your peak traffic is going to be.
Um, how do you train the model on company-specific information, which could be, uh, you know, product names, uh, or the domain the company operates in? Um, how do you think about, uh, you know, uh, kind of having a kill switch so that the, the, the, the, the s-system can, you know, uh, uh, uh, take, uh, appropriate action and [00:56:00] escalate if it's not able to resolve, uh, uh, you know, and the escalation needs to ha-- you know, ideally happen to a human.
And then, uh, uh, how do you test out these models in, um, uh, uh, before you move to production, having the right kind of, uh, scenarios that they're using for testing. That's, uh, on the key piece as you think about, uh, moving to production. Uh, next slide, please. Uh, the last two slides, uh, at least from my end. So, uh, you-- when you're thinking about a cascading model architecture, you need not always choose the, the biggest LLM.
Um, you know, uh, you can optimize for cost here by choosing a smaller model. Um, and the big benefit here is smaller models, uh, you know, have lo-lower latency. So you can take a smaller model, uh, fine-tune them on a company-specific data, and that way you're saving both on cost and latency. So this is something that many, uh, you know, many companies have, uh, have done successfully.
So, y- uh, so, so, [00:57:00] so ensure that you're not always going with the best-in-class, uh, frontier model. You can use a smaller model, it could be even open source model, um, and that would, you know, further save cost for you. Uh, next slide, please. Yeah, and just to wrap it up, uh, from my end. So, you know, voice, uh, AI or voice agents, they don't fail as, uh, you know, uh, uh, as a-- as because of, you know, you picking the wrong model.
They more fail as a system. So you think about the architecture, uh, ensure that, you know, you're, you're, you're not just doing fancy demos and, and, and, and hoping that it'll work in production. Uh, having the right kind of trust mechanisms in place. Um, so those are the critical things. And then now I'll hand it over to, uh, Venkata, who will more focus on the edge infrastructure and how, uh, thinking about caching, uh, would help with latency Yeah.
Thank you, Shanker. So as Shanker mentioned, uh, every voice turn has a budget. Like the moment you stop talking, the clock starts. A series of [00:58:00] steps would execute, like capture the audio, run speech-to-text conversion, and make the network hop to wherever your model lives, and wait for the model to generate a reply, and fi- convert that reply back to speech, and finally play it to you.
And there are like, so there are like six steps. And now, and the user is sitting there the whole time waiting for someone, uh, waiting for something that's supposed to be feeling, feel like a live conversation. And, uh, so out of all these things, the inference, the actual model call is usually the single biggest line item in that budget.
So it's also the first thing most teams reach for when they want to cut latency. Like as Shanker said, like maybe try a mod- a smaller model or a, try a faster provider or try quantization. Everything is valid, and all still leave one thing untouched, and then that is the network and the caching layer. And the network and the caching layer underneath, uh, the model has had like, uh, almost no attention in this whole conversation.
So even though it's exactly where CDN [00:59:00] engineering has spent the last twenty years solving the problems that l- that look a lot like this. And how do you get the answer back to someone fast and reliably without redoing the work you have done? That's what, uh, that's what it, uh, that, uh, uh, the talk is all about I'll be presenting.
So here, the one advantage here is like we are not only going to cut the latency, we are also going to scale it by taking a load off of the, your application server where the model is actually running, and that also cut, cut the costs. And so none of this is new, the concepts, and that's the point. So the three ideas made the web faster, reliable, highest scale is like the proximity, which means like don't make a request travel farther than it has to.
Meaning like the least distance it travels, the lowest the latency will be. And the next thing is caching. Don't redo the work you already did once. And the failover. So when one path fails, move the traffic to the other [01:00:00] working path and so that for the user, they never realize something has broken underneath.
And voice agents are running into the exact same problem right now. So they're just hitting the, they're just hitting them for the first time at a latency budget roughly ten times tighter than a slow webpage, which was never allowed to be in the first place. And So coming to the caching, uh, what exactly changes, right?
So as the, uh, in, in traditional caching system, here is the mechanism. Concretely, an exact match of cache is a lookup table. Traditionally, cache key-- cache mechanisms are like a KV. So where the value, the key will be in, in tradi-- in a typical web traffic, the key is the HTTP URL, and the response is the value.
So which means the traditional caching systems, they expect to match the cache key exactly same. So for example, like the user asks the same question, uh, with a different punctuation or a different word order or just a different way, the phrasing, uh, it gets treated as a different, [01:01:00] uh, answer. For example, like in the thing I mentioned, like how do I reset my password and I forgot my password.
A traditional caching system, they both-- they treat these both are a different queries, but actually these two have same response. So that's where the semantic caching comes. So semantic caching does something different. It converts the incoming re- question into an embedding, so a vector that represents meaning rather than just the character sequence, and compares the vector against everything that's already cached using a similarity score.
So I mean, if we set, we can-- as a developer, we can set a threshold saying like, "Hey, if it matches zero point nine and anything that scores above it counts as the same question, regardless of what exact words used." And that what, that's what lets a genuine rephrase still hit the same cache instead of triggering a new model call.
And that's the mechanism I will be, uh, showing a demo. So for this, so let me, uh, allow my System Wise because I need to allow my [01:02:00] System Wise here
Okay Okay, so here I have, I have built something, a miniature version of what typically we see in production systems. So this is how the production systems also would work, but this is something I built on my personal Cloudflare account, how we can actually, uh, show. So here I have like a threshold where, which I can choose like, hey, how much percentage, how much, uh, uh, vectorization, uh, uh, metric it needs to match.
And now let me ask a question. So I have a speaking button here. Let me ask a question. I forgot my password. You can try resetting your password by clicking the Forgot Password option on the login page and following the onscreen instructions. So here if we see it's a cache miss because I just cleared the cache and also it took eleven hundred milliseconds.
This will usually send a password reset link to your registered email address. If you don't receive the email, check your spam folder or contact the [01:03:00] relevant support team. Okay, now let me ask a q- same question in a different way, like how do I reset my password? Let me ask that. How do I reset my password?
You can try resetting your password by clicking the Forgot Password option on the login page and following the onscreen instructions. So here if we see this is a cache hit because previously this one, this, uh, when we asked the first question how I forgot, I mean, I forgot my password. This will usually send a
password reset link to your registered email address. If you don't receive the email, check your spam folder or contact the relevant support team. Yeah. When I asked this first question, it was a cache miss. There is a model call that executed and which took a lot of latency. And next time when I asked, so it's like sixty-eight milliseconds.
If we see, it's like a huge, uh, gain, especially in the voice chats, uh, this is like a deal breaker, almost like one, one second. And here, uh, underneath it is the semantic caching that, uh, that did the magic. And now going back to the presentation. So
Right. So here are the numbers I put. [01:04:00] Typically in production systems how much it is. I know the one which we just showed, it is like a way, uh, I think it happened to be the best case when I did the live demo. But typically when we run, so these are the wait times, like from on the, on an average, it's like eight hundred and ten milliseconds.
It totally depends on like where your origin server is like, I mean, where the model is at, the model call is going to, if you have models on-prem or if you have like your own trained model and if in your systems. So it totally depends. But on an average, like it takes eight hundred milliseconds versus a cache, it is like one fifty milliseconds.
So it does two good things. One is like the latency drastically goes down, which, uh, enhances the experience of the user. And the other thing is like it avoids the model call in the first place, meaning like your system can scale better and also it will reduce the model costs and the token, token, uh, cost and everything.
So yeah. And not just the latency. So there are other things actually that comes for free from CDN. Like CDN offers several other things. To briefly touch, uh, the, these things [01:05:00] like, uh, failover, like one of the thing. So for example, like there are, uh, if something fails on your origin server side or if something is not up, then CDNs automatically you can build your mechanism so that you have another endpoint which is working, so you can keep a fallback endpoint and CDNs automatically go and then, uh, fetch the, the data from there.
And the other thing is compliance routing. So like CDN have spent years, uh, routing traffic to satisfy data residency rules, like keeping the regional users data in regional infrastructure. For example, like voice agents handling sensitive conversations are going to need exactly the same thing. And the infrastructure to do, uh, uh, is already, uh, existing.
So just to close out, so the fix is not always the model. Uh, sometimes, uh, it can be the network underneath. I mean, the network underneath, irrespective of whether you are using, uh, whether you are fixing on the model, model side or not, there is always something you can do on the network side to enhance the things.
So yeah, that's all. Thank you. Any [01:06:00] questions?
Excellent. Thank you, guys. Um, yeah, very cool. Uh, there was a question that came in the chat, um, and I will just say upfront for anybody, uh, watching, listening in, if you have any questions that you'd like to ask, please, uh, send them in the chat. Um, I will be taking a look, trying to ask those. Uh, so the first question that we got in, uh, came from Jose, uh, says, um, "How does the system know, uh, that the human has stopped speaking and it's not just making a normal pause for voice agents?"
So, sir, can I repeat the question one more time? Sorry. Sure, sure. Sorry. Yeah, so the, the question is about, um, how, how does a voice agent determine when a speaker has, like, actually stopped talking or if they're just pausing, like, to think mid- mid-thought? Um, yeah, uh, Jose was curious to, to know. Yeah, [01:07:00] I mean, a lot of it is based on how the models are trained on various types of data.
Uh, so it understands from that, uh, uh, various patterns on how, uh, humans are pausing. Um, there's no magic bullet here. It's more about how do you train the models on, uh, different types of scenarios. It's-- this is one of the hardest kind of, uh, pieces to solve, uh, be- because it's difficult to figure out, you know, one is truly pausing or, uh, you know, taking a moment to think.
Um, and, um, you know, and this can have downstream impacts. So, uh, to just, you know, uh, summarize, it's more about how do you train the model on variety of pause scenarios, um, and the data, uh, that you're using for it. Yeah. And the system can learn from this person too, because every person has a different pattern.
So maybe first time it gets wrong, second time it gets wrong, but eventually it learns, okay, this person has this speaking habit, so it knows how much time it needs to wait before it considers it has ended. Gotcha. That makes a lot of sense. [01:08:00] Awesome. Um, another question that came in, uh, I'm gonna try and summarize it a little bit, is, uh, it has to do with, like, security and trust, um, about, uh, what do you both kind of expect or foresee in the future as, uh, the biggest security blind spots emerging over the next few years as particularly around voice agents?
Mukund, you wanna talk, uh, take this from a network perspective, then I can talk about it Yeah. So from network perspective, honestly, these are like, ultimately this is HTTP call because it's not the voice as it is going underneath, so it gets translated to a text. Now from there, we know traditionally we already solved that in a system, in a security perspective.
So in that way, I would see this is something it's already a solved problem. It just like, until it gets converted to text, that's it. So once it is text, we already know how, how to solve it, and we already solved it in the first place.
Yeah, the only thing I would say is, uh, you know, [01:09:00] if you're thinking about a speech-to-speech model, um, you know, if it uses an LLM underneath it, uh, or a multimodal, um, LLM, a lot of the existing best practices around handling security for LLMs still apply. Um, uh, you know, uh, how do you ensure that, uh, you know, the prompt is, uh, clean, there's no jailbreak, a lot of these things.
So all these techniques, uh, apply and, you know, it's evolving area. Um, so nothing special for voice agents. The same best practices apply here. That makes a lot of sense. Yeah. Um, if, uh, it, it sounds like, yeah, mostly voice is kind of the, the medium in which you're interacting as opposed to typing. Um, I think maybe the only other thing that I could think of is, like, anybody who's kind of storing audio clips maybe, uh, there could be something there.
And actually, that leads into a question I personally had regarding the caching, because I thought the semantic caching was a, you know, it's a great idea, and I was wondering if that is, um, text or, [01:10:00] or tokens that's getting, uh, cached or if that's actually, um, like embedding of the model or like a Fourier trans- or, or the, the audio, excuse me, or like a Fourier transform of the audio, uh, that's used as the, as the, uh, uh, cache.
No, not the audio. It's actually the text. Okay. So we take the text, the question converts embedding. Yeah, typical text embeddings, yes. I mean- Yeah ... maybe we might evolve into that, that level of, like, uh, caching the, uh, uh, audio itself. But honestly, before you do, there are several reasons why you want to convert into text, because text is something we already know how to process and all, right?
So when you are, uh, processing voice directly, there are several challenges. Like, as you said, like the security aspect also comes into picture, so you need to put a lot of filter and a lot of innovation needs to happen. At the same time, you need to see what is the return of investment. Is it really going to make cut the latency into a severe extent?
Then yeah, maybe that innovation might come. But currently today, it's all like embedding models. And the one cool part is, like, the one-- [01:11:00] the embedding model I used, it's out-of-box embedding model provided by Cloudflare, which mean, like, there's actually the CDN systems are nowadays they're offering the models at their node level itself.
They're s- very small models, honestly. But still, it can do a lot of, uh, things for you if you can design your use case around it. Excellent. Awesome. Well, thank you very much, guys. Um, uh, I'm gonna, I'm gonna wrap it up here. We don't have any other questions that came in from the chat. But, uh, definitely very informative as somebody who has never deployed a, a voice agent myself.
Um, you know, this was really insightful. Uh, so thank you again, and, um, yeah. We'll-- we hope to see you at the, uh, the Voice Agent Forum, possibly in November. Awesome. Thank you for having us. Thank you. Bye. Yep. Take care. All righty. So, um, next up, uh, actually, I think this might be a little relevant to one of the questions we had from the last, uh, uh, session here.
Next up we've got, um, Fabian, [01:12:00] which I hope that I'm saying that right, and I, I, I really hope that I'm saying that right because my, uh, a friend of mine, his father's name is Fabian, so I, I hope the pronunciation is on, on, uh, the same. Um, and, uh, Fabian's the founder of AI Acoustics, which I gotta say I, I love the name.
Um, it's, it's a, it's a great name. Wonderful branding. Um, and you're here to talk to us about, uh, background noise, uh, multiple speakers, bad microphones, all the nasty things that can happen with, uh, voice agents and, uh, problems with, uh, AIs understanding what's actually trying to be said. Yes, that's correct.
And yeah, the, the name pron- pronunciation, uh, is, is also good. Uh, Fabian, or in German I would say like Fabian, bo- both fine. But yeah, the topic of the talk today is about different types of input challenges that you could have building a voice interface or especially also like a voice agent. And relating to that, uh, how do [01:13:00] we at AI Acoustics think about these type of audio input problems, and what's the technology that we've developed to help make, uh, again, voice interfaces, but also voice agents more robust to, I would say, real world and messy conditions.
Excellent. Well, feel free to go ahead and, uh, share your screen. Um, I, I'm looking forward to demos. I assume that we're gonna have some, some really neat demos to, uh, check out. So looking forward to it. Yes, correct. Uh, let me see if... Well, I guess I can also share the entire screen. Yeah. Whatever, whatever works best for you.
Well, not so much. Um, I'll try to do up the Chrome tab for now. Okay. And you should see now starting the presentation. Yep, we see you. Okay, great. Um, so yeah, uh, thanks for having me here at the, uh, at the event, uh, at the Agentic AI [01:14:00] Foundation. And again, the topic of this brief talk is about how audio input can actually break your voice agent.
Um, I'm a founder and the CEO of aacoustics, and I've brought, um, a few, well, a, a few cases in, in the beginning and why basically the test and development environment that are often being used for voice agents, and they often look something like that, um, where you're in a, in an office space, you might be in one of these like telephone boxes.
But it's something that's, I would say, acoustically rather easy and a lot of ... We see a lot of, like, companies developing or testing rather their agents in such an environment. But when it comes to the real life and the messiness, um, especially when it comes to the acoustic conditions, that sounds very differently.
Um, and here are just a few examples. It could be a busy office where, you know, there's a coworker in the next cubicle, um, and this voice is [01:15:00] bleeding into the overall conversation. We're working with a few partners. They're actually building voice agent for drive-through, which we sometimes call the Acoustic Champions League because, um, there is, uh, a car or, or a street involved.
People are not very close to the microphone. There's maybe the kids in the background, the radio's running. So all kinds of different, well, influences that you have to deal with. Um, then you never know if you also do like outbound calls, um, is it people with, you know, a new smartphone, uh, answering? Is it like a, an old landline?
So there's different, like, devices that are in play that you have to, um, respond to properly or your agent basically has to work in. Um, and overall, we've also seen a few people that, you know, um, are calling from the car or into the car, so there's a lot of other digital sig- signal processing also involved when it comes to not just the analog and acoustic properties, but also then how is the audio being processed further when it's
[01:16:00] once when it's digital. Um, and my second point is that when it comes to agents and when it comes to an agentic conversation, background noise is not always the same than background noise. So there's different types of acoustic conditions that are more or less challenging actually for an agentic system.
And while Um, there's a few conditions that are fairly easy by now for, uh, an agent to overlook or rather overhear. Uh, something that's a bit more like stationary and that's like rather, um, not dynamic. Um, a few examples here could be a fan in the background, a bit of city noise or, or street noise or something.
That's something that, that agentic conversations, um, both when we look at the, uh, underlying parts of VAD or an ASR system or even a speech-to-speech model, um, can handle pretty flawlessly by now. Um- However, um, what we see in most of the evaluation data sets, um, the training data sets [01:17:00] as well, is that, um, most, yeah, data sets are actually, when it comes to the audio quality, something that's, you know, recorded at home in a, in a quiet room or something.
And interesting enough, a lot of, uh, tests that we're seeing for ASR systems or, or VAD systems or, or speech-to-speech are also based on, well, I would say acoustically unproblematic conditions. And while in lot of, um, calls that might pre- be true, there's a, a decent amount of calls that we're seeing, uh, from day to day where you just have to also evaluate your system in a more, well, messy and real world environment.
When it comes to the more harmful acoustic challenges, uh, when it comes to agents, um, this is something like, you know, imagine a second person in the room, um, adjacent to you, sort of like side speech that bleeds into the conversation or background voices. It could be a TV, a radio podcast that's still running in the background while, you know, you call in, [01:18:00] uh, into your bank or, or, or order at the restaurant.
And sometimes even the agent's own echo, the device echo basically that leads into the conversation. All of these things, um, basically lead to the agent either being interrupted while it speaks, um, or sometimes we also have that you, um, are ... You- you've basically, like, stopped talking to an agent and asked, uh, and wait for a response, but there might be some, uh, speech adjacent to you, and the agent thinks you're still talking and will never reply back to you again.
So one part is the, I would say, um, errors when it comes to the transcription, which is often like an over-transcription, so too much is being trans- transcribed because of the speech that bleeds into the conversation. Another part is does the turn-taking actually work properly when you have some background voices or side speech?
And what's the main issue here is that the agent, and that's my [01:19:00] next point here, can't really focus on what's important and what's not por- important in a, in a conversation. Something that we as human beings are very ... Well, something that we've, we've learned over a lifetime that's very easy for us to do.
Um, another differentiation I would like to make is, uh, de-noising as we know it from other applications when it's mostly like we call it human-to-human interaction, um, where it's like a normal communication on a, on a phone call or on video call. Um, they are optimized and they are trained for a very different target.
And the target would be, is someone on the other side or on the other line, uh, or on the laptop, uh, there, does, do they, um, understand me clearly and does it sound pleasurable or does it sound nice for the human ear? And most of these, we call them perceptual enhancers, um, they are helpful to make an audio better for the human ear, but they're not necessarily make the audio better when [01:20:00] it comes to an agentic conversations because, uh, the output of these perceptual enhancers that you can, um, also a lot of find, um, in the internet or open source, um, they, the output of these have not been used for training agentic systems, VADs and ASRs.
So using one of those, um, often we see a degradation when it comes to VAD accuracy or word error rates, so to say. Um, and that's a different diff- uh, that's a, that's an important differentiation that, um, optimizing or voice isolation or enhancing speech for an agentic conversation is very different than enhancing and voice isolation for a conversation human to human.
And, uh, there's a few other aspects I would like to mention. Um, I'm, I myself, I'm not referring too much, um- About the term denoising, we rather call it voice isolation, uh, because it also encapsulates that you have different distances to the microphone. You might [01:21:00] have different room acoustics, um, in terms of size or the reflection patterns that you need to remove.
There's different microphones that I've mentioned, and even like, um, a few echo, um, effects in all of this. So voice isolation captures, um, a bit more than just the denoising part, and it's a more complete term. Um, that's the first part, and the second part is there's a major difference between do I optimize denoising or voice isolation for a human conversation or for an agentic conversation.
Now, how do we think about this at aX acoustics, and what's the technology we've de-developed to basically tackle all these different, um, audio input problems? Um, the first one is called, uh, Quale Voice Focus, and that's a model, um, audio input and audio output that takes an audio stream in all in real time and is able to isolate the primary and main speaker in the conversation.
Um, and we've trained into this algorithm similar auditory cues that we would [01:22:00] use as a human being to dis-distinguish basically who's in the front, who's in the back, what's a speech part and whatnot. Um, think of it if you're on the phone or in, in actually in a, in a, in a real world room, um, you would use auditory cues like someone who's closer and might be the main speaker is a little bit louder.
But also there's a different like reflection pattern that someone who's like further back in the room. The direct, um, sound path, for example, is much stronger, or the higher frequency roll-off if someone is, is, is in the back is much stronger because of air absorption. Um, and these are all auditory cues that we use as human beings, and this is something we've trained into this voice focus model.
So it gives, um, it, it improves basically an audio stream where you have-- where you can have, um, a noisy audio or you can have multiple speakers in there, and it outputs an audio stream where you only have one primary and main speaker that you can then feed into your ASR system. Uh, the VAD that we have [01:23:00] is just a very, very noise robust VAD.
It's been battle tested across a lot of different, um, use cases even in, um, drive-through applications, and it produces a lot less false positives, um, than a few open source VADs or other VADs on the market. Um, and that helps a lot with proper turn-taking mechanisms And the third model, um, that I wanna talk to, to you today, all of these three models, um, I explain a bit more in detail in a second.
The third model is our newest, uh, release. It's called Title, and it's an audio insight model. Uh, it's basically a model that gives you a metric and scores an audio input and tells you if this audio is reliable, if this is intelligible for agents, or if there's a high risk that this audio actually leads to a failed conversation or a failed call.
Um, when we look at the, uh, processing pipeline, um, we [01:24:00] always sit sort of like relatively in the front of the overall processing pipeline, so there's always audio input. Could be through laptop, could be through, um, basically a web RTC or a, a phone call incoming. Um, you can use our VAD in front of turn-taking mechanisms.
The voice focus would classically sit in front of, uh, speech-to-text or an ASR system. And then we've got Title, which, uh, is more like a metric, and you can use it to steer different, well, parameters or even post-call analysis. I'll get to this in a second. Well, first About Quail Voice Focus, I've already mentioned, um, it's meant to improve audio, especially before an ASR system, and we've run this on a very challenging, um, acoustic audio data set where there was a lot of conversation in a room.
And, um, it does improve the word error rate, uh, massively when it comes to especially the insertions. [01:25:00] Um, insertions is a type of word error rate. So we have three different types: deletions, insertions, and substitutions. An insertion is basically if the ASR provider transcribes something that is not meant to be transcribed, like additional words, so to say.
And this happens if someone, again, is in, uh, a direct, like, uh, directly next to you, or again, it could be a TV, something like this, and that's falsely transcribed for an agentic conversation. Um, so generally the ASR providers, um, are transcribing more that's necessary for the conversation. That's not necessarily, um, like a-- That's not necessarily sort of like a, um, a misinterpretation of the, of the ASR provider or, um, yeah, an error, so to say.
It's just that the, uh, ASR providers don't have a focus mechanism, uh, that only focus on the, again, primary and main speaker, and this is where our technology would come in. Um, [01:26:00] you can read a bit more about the newest version of our Voice Focus model two point two on this blog or this URL I've, I've mentioned here in, in the slides.
Um, we also do, because, um, I've mentioned it down here in the slide as well, we do, uh, publish the data sets on Hugging Face that we evaluate on. So this one, um, is called Dawn Chorus, and you can find it on our Hugging Face page, um, which I'll show in the last slide of the presentation. The second one, um, to show a few numbers about VAD.
So the biggest failure cases for VADs is, for example, if they trigger falsely even though there was no voice, um, present, that would be a false negative. Or if they don't manage to recall all the voice parts that were in a conversation to flag them as speech parts. And while on Um, I'd say non-challenging acoustic cases, [01:27:00] um, normal VADs or open source VADs work relatively well.
Uh, when it comes to more challenging cases, and this, for example, is an evaluation that we've done, uh, for a drive-through data set. Um, the-- we can see that the, um, here in this case, the Celero VAD has a lot less accuracy and a lot less, um, yeah, recall qualities, um, because it's just not been trained on that much challenging audio.
Um, and so it hasn't learned sort of like to discriminate these more challenging cases. Um, and so we can both improve the accuracy as well as the recall, um, metric, uh, on this very challenging drive-through set by quite a bit. Um, and what-- why this matters, it, um, it, um, it detects the speech parts where otherwise, uh, a VAD would maybe miss a turn, and it also majorly reduces the false positives in that sense.
Now, coming to the, the third model, [01:28:00] which again is called TAITO, that I've shown, and for this one I'd actually go into a small demo that I've prepared before I go through the slides. Right. I'll hope, um, you can hear or you can see, um, the, the other slide, um, where we have an interactive demo. I'll also share the link about this one, uh, on the slide as well.
We have an interactive demo, um, around TAITO. That's something which is online on our website, and you can have a look at yourself in a second. Um, so just to explain what TAITO does, um, it receives an audio input, and again, it, it outputs different, uh, metrics. And the headline metrics is called the risk score.
And the higher the risk score, the higher the probability that this conversation is probably going to fail. And by failing, I mean it produces either word errors, um, [01:29:00] in either kind of form, deletions, substitutions, or, um, insertions. Or it could also mean that, uh, a voice activity detection highly struggles with this audio, and that would mess up the turn-taking.
Um, so that's, uh, what we've taken into consideration for the risk score. And if you see that a risk score is relatively high, then we have six different, we call them explainer dimensions or variables that can give you a hint or an explanation why this, why this risk, uh, is so high and why the conversation might fail.
Um, and these underlying dimensions are speaker reverb, uh, speaker loudness, interfering speech, noise, packet loss, which is basically network errors or codec degradation. I have to mention not all of these dimensions are harmful. For example, speaker loudness is sometimes an interesting thing to hear, also to adjust gain control, something like this, [01:30:00] but it's not necessarily harmful.
While other sub-dimensions like interfering speech or noise or packet loss, for example, if they light up, this is probably very dangerous for the conversation. These underlying dimensions are not a linear or nonlinear combination of the risk score. They're basically independent. And the idea is that, for example, if you would use this title model to go to score all your calls in post-call analysis, that could be one use case.
You can see very quickly when you run this on millions of calls, which of these calls did fail? Why did they fail? Maybe a customer complained, and then you can find out what was wrong in these calls. It's in that sense a debug mechanism that you can use for post-call analysis, either as a standalone mechanism, or you can also use it as an additional input for your LLM as a judge, for example, to give [01:31:00] your LLM as a judge, let's say, an audio context that it wouldn't have otherwise, since it normally only relies on text.
So let's have a listen to one of these or one or two of these examples. This one, as the name already says, has a decent amount of interfering speech. Or maybe you're driving up to a drive-thru to get a quick bite to eat before work. And instead of the human being on the other side of that conversation, they might be listening in, but it'll actually be an agent, a voice agent that's doing that conversation.
That's a classic conversation where a voice agent is talking and the caller only gives yes and no answers quickly for specific questions, but there might be the TV or a radio program running in the background and that will lead to interruptions of the agents and also a mistranscription. So this is a very high risk.
We have another example [01:32:00] with also relatively high risk. Everything above 0.5 is high risk, where there's like noise included. Wait
for the police car. And what would be the new price? Um, yes. So it's 101. That's a situation where someone's outside, um, is calling, I think, to cancel their gym membership, and ambulance is driving by. And maybe I'm showing a third one called, uh, packet loss. AI Acoustics front desk. That's a very stuttery effect, um, when there's heavy packet loss happening.
Um, and interesting enough, while this probably sounds the, the least nice for us as human beings, um, this can still be more or less transcribed. It's like an intermediate risk score here from ASR providers, while the other ones are flagged fairly high. The noise example [01:33:00] here also has a high probability of noise, while the interfering speech, uh, comes with a very high, um, probability of interfering speech.
So these match actually, uh, the contents. All right. Um, I will go back to my presentation. Uh, what is this, um, Tito model being used for? Um, in that sense, um, and also how I've just showed it at the moment was, um, again, as a post-call analysis tool. Um, we've seen it more and more for this type of analysis, again, as a standalone tool, but also as an input to LLM.
You can basically eval a lot of like datasets with it, uh, and can go through it and have, um, an overview of which types of clusters you have in your data when it comes to the audio input scoring. There's a second use case that I would like to, uh, showcase quickly, which is how can you use Tito more in like a [01:34:00] real-time environment or in a real-time context to give the agent, to give your agent directly on the phone some audio insights or audio intelligence.
Um, let's have a quick, uh, yeah, listen or view of this video.
Or move closer to the mic Yes. How is it now when I speak now? Can you, can you hear me better? I can hear you a bit better now, but the audio is still a little degraded. There's some reverb and a slight drop in clarity. You might wanna be closer to the microphone or adjust your environment a bit. Mm-hmm.
Yeah. So exactly. That's what I'm doing right now. I'm, I'm, I'm, I'm closer. Your internet seems unstable, and I think I lost some of what you said. Could you check your connection and try again? Uh, yes. I wait. I checked. I think now it should be stable. Now it should be okay [01:35:00] Great. It does sound more stable now, and your voice is clearer.
We can go ahead and talk more about those robot investment ideas. What kind of details or costs are you looking for exactly? Perfect. So Could you speak something to- Yeah, yeah, I can speak just in the background so- So, um, yeah. So going back to the, to the robots- It's probably hard to, probably hard to understand because people are speaking a lot, like, and just start speaking over each other,
right? You may be picking up other voi... It sounds like there might be a bit of overlap with another voice in the background. Could you confirm if it's just you speaking to me right now? Yeah, sorry. I just moved on to the other These are a few, um, e-examples how you could use TITO basically dyme-dynamically, um, to give your agent, uh, this type of audio context, um, and the a-ability to react to different, like, audio conditions directly.
Um, it's similar to how you would use or how you as a human being on the call would probably also tell someone, "Hey, can you move away?" [01:36:00] Uh, if something's very noisy or, you know, who to actually, like, talk to. Um, and that's the first part of, you know, a, a piece of, like, audio intelligence that we would provide for, um, a voice agent to make it a bit more aware of its, like, um, audio surroundings, so to say.
Um, and you can use it e-either to, yeah, steer dynamically a conversation and give feedback to a caller. Uh, you can also steer, um, your pipeline based on TITO audio, uh, TITO outputs. Uh, for example, if it's a lot of interfering speech, your barge-in could be more careful. Um, if the phone line's degraded, maybe, um, you know, it could be lower confidence on, on turns.
Um, or if the audio is very, very difficult and basically not even manageable for an agent system, you could hand off to, like, um, um, a, a human agent in that sense. Um, so there's different ways how you can use this TITO output then to make the conversation better or route it i-in the right way. For both, uh, [01:37:00] the demo that you've seen here in the video, um, as well as the one that I showed before, which is a bit more for post-call analysis, um, you can go to this website and check them both out.
Uh, one is interactive. For the other one, you can even, uh, upload your own audio files and try it out, how it works and how the different, the risks go, and how the different parameters would, uh, react to your own audio. Right. All of these models that I've shown, um, uh, VAD, uh, Voice Focus, as well as, uh, the TITO model are available in our AI Acoustics SDK, and it's relatively easy to integrate.
Basically, you need, in Python, you need these, uh, five lines of code, and that's it. So it's a relatively non, non-time-consuming and, um, straightforward integration. We do have bindings for these five, uh, interfaces, Python, Rust, C++, Node.js, and WebAssembly. Um, we also have intriga-integrations, uh, for the leading platforms, for [01:38:00] PipeCAD, for LiveKit, and for Slang that we heard a bit earlier in the presentation today And what I also want to mention is this is an SDK that runs on your infrastructure and on your device or, or on premise.
So we never-- It's not an API where you send the audio to. Uh, we run, again, on your infrastructure, so no audio or personal information ever leaves the system. The only thing that comes to our back end is a ping from time to time, how much you basically use the SDK. Um, and in terms of latency, um, we are using thirty milliseconds of latency.
It runs on CPU. Think of the computational complex, complexity about one or two percent of one core of a Raspberry Pi, uh, more or less. And one last thing to mention, um, all of these models are language agnostic because we don't analyze the text, we don't analyze the sentences and understand the context of being...
what's being said. We basically anal-analyze very short frames of audio in [01:39:00] the spectral domain, um, and can distinguish what, yeah, is background speech, what is foreground speech, what is noise, and distinguish them and remove them or analyze them. But it's not so much about understanding the words and the language.
So all of the models that we have are trained throughout a different ti-- um, amount of languages, but they overall are language agnostic, and they react very well to all tonal languages in the world, which is probably ninety-nine point nine percent of all the languages. Um, last slide. Um, feel free to try this out for free, um, on our developer portal.
The URL is on the bottom right. Um, you can try it out for free for thirty days, uh, unlimited volume, so you can run, um, millions of, of, of calls through it, uh, if, if you want to for this thirty days amount. Um, it's self-service SDK keys. And yeah, we-we're very interested in [01:40:00] your feedback. Feel free. Next slide, maybe.
Feel free to reach out to me, um, add me, uh, on LinkedIn, um, if you're interested in this topic or if you, yeah, have, I'd say, challenges when it comes to the audio input in your voice interface or your voice, uh, system. Uh, here's two more links, um, about, um, the benchmarks that I've showed and also about, uh, yeah, the different datasets that we've published on Hugging Face.
Uh, thanks a lot.
I, I can't hear your voice yet, Burhan My mistake, I forgot to unmute myself 'cause the, there was loud noises from outside. Um, uh, I was saying, uh, I think, uh, we have time for one quick question, but I will say that there were a handful of the same question that came in. [01:41:00] The ... It was the question I had, you already answered, which I think is really cool, that, uh, about the Taito, uh, uh, model being able to be used with the LLM to kind of like steer the users back.
Um, I think that was a really cool demo, really helpful. I shared the link for the demo in the chat, uh, as well for people to, to check out. And I also see that Luke from, um, uh, from, uh, from S Lang earlier, uh, uh, shared that, uh, they have a integration already set up for, uh- Yes ... unmute with AI Acoustics, which is really cool.
Exactly. Um- Yes. Uh, thanks a lot to the, uh, to the S Lang team for integrating us so quickly and checking it out, uh, it out. Um, it's available on this, its platform. If you have any questions about, um, the integration, both, you know, um, you can, you can ask me or, or Luke as well. Um, if it's more about the models itself, we're happy to t- to help.
Also, um, if you have some very challenging audio, uh, feel free to send it to us. We're receiving very interesting, um, and again, acoustically challenging audio [01:42:00] every, every day, and this is how we learn more about the use cases and eventually also improve our models of course. Excellent. Thank you so much.
Um, I guess the, the one real quick question that I wanted to throw out that I saw was just somebody was asking about the, uh, version of, uh, Silero VAD that you benchmarked against. Is that ... Do you know that off the top of your head or is that- Yes ... posted anywhere? I think it should be 506. Um- 506 ... but this is also, um, if you go to the, to our blog, which I've linked, um, this should be mentioned and referenced exactly there.
Okay. Excellent. I will point people there, drop the link for that in the chat as well. Thank you again so much, Fabian. Great having you. Um- Thanks for having me ... enjoy the rest of your day. Cheers. Have a good one. Yeah. Bye-bye. Yes. All right, next up we have, uh, uh, Brooke, uh, Hopkins, I believe, uh, from, uh, co-founder of Koval.
And, uh, happy to have you here with us today. Um, I'm told that, uh, you're gonna talk about [01:43:00] how voice AI is related to self-driving cars, maybe? Yes. Very cool. Awesome. Um, can you hear me all right? Yes. I can hear you loud and clear. Amazing. Cool. Well, hello everyone. It's great to see all of you. Super excited to chat, um, and excited to be speaking at the Agentic AI Forum.
Um, have been close with the MLOps community, so super excited that they're now part of the Agentic AI Forum. Um, but yeah, we can jump right into it. So let me just, um, pull up my tabs. Uh, but yeah, I'm Brooke. I'm the founder of Koval. So we build simulation and observability for voice agents. And my background is from Waymo.
So I led our evaluation infrastructure team at Waymo, and my team was responsible for all of our developer tools. And you might be like, "What does self-driving cars... What do self-driving cars have to do with voice agents? [01:44:00] Um, does, did, uh, s- you know, did Waymo have audio in their car? Like, what's happening?"
But actually, um, today what I wanna talk about is how similar these systems are. So, um, yeah, my background, again, like our team comes from autonomy from robotics, uh, et cetera. And so really thinking about every single day, how does autonomy and robotics apply to voice AI? And fundamentally, the thing that's really similar here is both are...
Like, any autonomous system is answering these fundamental questions. Like, what is happening in the world around me? What should I do next? And then how do I actually take those actions? So in self-driving and robotics, or like any robotic system, you have perception, so it's taking in all this input. That could be radar, lidar, cameras.
Um, and then it's taking all that information, deciding what should I do next. So how do I actually plan a route if I'm trying to get across the factory or if I'm trying to get to, you [01:45:00] know, the Richmond from Fidi? How do I find a route that takes me there and then continuously, um, you know, make decisions?
So if all of a sudden there's construction in front of me, do I go around that? Do I take a different path? Do I go left? Um, et cetera. And then once you decide what to do, you need to then have controls to take that action. So that's like steering, braking, controlling speed, whatnot. And in voice AI, this is actually shockingly similar how these systems work.
So in voice AI you have transcription, which is speech to text. Um, so when I'm speaking, transcribing it, understanding, like just as Fabian actually mentioned, a lot goes... As the self-driving stack is very overly simplified, so is the voice AI stack. You know, everything for how do you decide when to start talking, there's like a in-between model here.
Um, but even just like how do I make sure that I'm taking in all of this information around me in the world? So, um, both the audio, that I'm [01:46:00] canceling out audio that I don't care about, that I'm like deciding when to take my next turn, and then reasoning, so like what should I do next? Um, like how should I respond to that person's question?
And then voice, which is how do I say that out loud? And so even though these are like very overly simplified, these are like the fundamental steps you need to do at its core in order to build an autonomous system so it's able to reason around the world. Um- And so one of the things that... Like, today I wanted to just talk through, like, all of the parallels here, because not only on the architecture side of things are we seeing a lot of parallels, but also on how you test.
So a common pattern is that you just test with, like, a hard-coded script or, like, a recording, and then see how your agent fares. And this is also where self-driving started, is that you just replay recordings or, and see, like, how well the agent does, how well the self-driving car does. But the problem [01:47:00] here is that as soon as you veer from the path even a little bit, now that no, that test no longer works, and so these tests become very brittle very fast.
Um, so you can see here it's like if this is the call you recorded and then your agent now, instead of asking for the email first, it now asks for your identity first, that very quickly becomes a problem because now you don't... Now that test no longer works. Um, and so similarly, like, simulation is really r- like, is really crucial here to be able to model what are all these different possibilities and how can I model how the world is going to respond based on what's happening with my agent?
Um, there's also, I think the, the similarities between self-driving and conversational evals is that you also have these trajectories. So you're trying to say, like, "How do I get from point A to point B and make sure that, like, my agent is doing everything I expect along the way?" And you can have a [01:48:00] spectrum of capabilities throughout.
So you could say, like, um, for example, at, at Waymo, we had large scale simulations where we did take a lot of the logs from the road, and then we would take those as inspiration for the car and then run it, like, 10,000 different ways. Um, and this can be really useful because you're still grounded in reality, but you're taking that and finding inspiration of, you know, if I vary all these different things, now that I've modeled the world and also modeled my agent, I can see how all these things are responding and reacting
Um, and so I think this is exactly why evals are the new product definitions, because at Waymo the way it worked is actually simulations were the definition of like what do I need the agent to be doing, what do I need the self-driving car to be doing in all these different environments, and then how can we optimize the self-driving car for that, like, for, for that, um, use case.
[01:49:00] Uh, and so being able to create these, like, really responsive tests in responsive environments allows you to see both, like scale the tests, so make them really durable so that every time you change your agent you're not having to, um, like create a whole new data set or like go collect real world data.
Which obviously if you go create world, real world data, um, that means you're impacting users or at least taking some risk in the real world. Um, and then finally it allows you to have more coverage. So imagine you're trying to map like, like the coverage of the entire, you know, environment. How do you balance that with, you know, getting signal as well as getting, um, you, like can't, you can't simulate everything in the world.
Um, and so this is the really tricky part of simulation is that you're trying to find the balance, or this is the tricky part of deploying autonomous things in general. Um, you're trying to find this balance between cost, signal, and latency. [01:50:00] Like you could, if you didn't care about money you could run every simulation in the entire world, but it would co- it would take forever.
And so one of... Then you won't be able to ever release. Similarly, you can run every single, um, simulation really quickly, very cheaply, um, but it's probably not gonna have very much signal because they're not very high fidelity tests. Uh, and so really you're trying to find like what is the minimum number of tests I need in order to have signal and coverage, while also being able to scale, um, yeah, scale, scale my voice agent across the board.
I think another really interesting learning from self-driving is like the difference between probabilistic evals and input versus output evals. So in traditional software engineering, you have a bunch of unit tests and then those unit tests pass, um, and they either pass or don't pass. But with probabilistic evals, the thing we're really trying to figure [01:51:00] out when we launch voice agents is what percentage of the time is my voice agent going to do what I expect, especially for regulated industries or for cases where there are, um, you know, uh, banking for example, how often is my agent going to reveal a social security number versus how often is my agent going to, you know, maybe have an awkward pause?
Those are two very different risk profiles. And so I think similar to Waymo, a lot of this is like, you know, one in like XX million miles, this like, um, hard break can happen, but one in let's say a, a million times you could have maybe like the wrong route chosen. And so you start to have like different probabilities of like, um, same with is true with voice agents of revealing a social security number one in maybe 100 million calls is your risk tolerance because this basically can never happen.
But, um, having [01:52:00] a like call drop can happen one in a thousand times and so that's going to be a very different risk profile And so this is where simulation is really critical because it allows you to scale this, right? You can't actually make 10 million calls, and you also don't wanna deploy a voice agent to m- 10 million calls without knowing that it's not going to reveal Social Security numbers.
And so this process of running a small change-- like, I think there's a lot we can learn from how did Waymo deploy features, where you run a, a small feature change and you eval that change over and over. So let's say I'm making a change to intersections, then I'll run, like, my own set of simulations on that.
And then before I start to, like, create a release set for this, then I'll run a larger set of regression sets, which are gonna be more expensive. And this again goes back to the, like, cost signal and latency trade-off. Um, so I'll run that larger regression set, um, and then [01:53:00] once I, you know, create a PR for this, then I have a set of evals that run on that PR as well as post-submit.
Then there's a constant set of even larger set of simulations running. Um, and then finally, before there's a software upgrade, there's going to be large scale release process. So that's every time we, like, up- um, release a new software version, they're doing, like, an entire release process with simulations and then deploying that and seeing in a gradual release to see how well that's doing with our, like, manual internal drivers.
Um, and then finally you have, like, live monitoring and telemetry of the vehicles at all points in time. And so you can see that there's, like, everything from... It's-- You wanna discover the issues way earlier in this process, right? Especially with self-driving cars, but also with voice agents. Like, the more you can shift left, the, the better, to take a term from DevOps.
Um, but at the same time, sometimes they're just like, the world always has new things, so, like, only doing [01:54:00] simulation is not going to get you there. Um, so... And then also, like, your goal is not necessarily to automate all evals. Like, there is a role of human in the- these evals, but the goal is how do you speed up the iteration time?
And then also how do you make sure that the manual evalu- evals are constantly, like, being fed back into your simulations, into your evals, so that you are only having humans look at the hardest, most challenging scenarios. Um, so like these manual evaluations are constantly being fed back, um, back into then the real world conversations, um, and simulation conversations, so you're constantly improving your agents.
Um, that's where I think, like, this... A lot of these concepts around, like, self-improving agents and loops are really important because you need to constantly be feeding all of your learnings back into your agents. And in a large enterprise [01:55:00] organization, that means setting up the operational practices to have those feedback mechanisms.
Another question I get a lot, and I think there's something to learn from self-driving, is what level of simulation realism is needed? So realism is, like, expensive and yeah, realism is more expensive. Um, basically, no matter, like in autonomy, in robotics, in voice AI, um, a lower fidelity test is always going to be cheaper, um, because it takes less compute, it takes less effort, it take, you know, all of that.
Um, but the way it worked at Waymo is that you have like kind of this hierarchy of tests of, um, that allows you to test like different parts of the car at different points. Um, and then very common was like, there would be this release video of like super high re- like realistic 4K video of like all generated scenery and people are like, "Wow, that simulation is so much better than [01:56:00] whatever like simulation that I saw released from this other company that is like not photorealistic."
But the thing is, you don't actually need photorealism if you are only testing planner. Like if you're assuming that, if you're isolating planner and like ignoring that, just assuming that like the perception ground is ground truth, and then you wanna say, I wanna understand like behavior prediction, like what do I think like the dog is going to do in this scenario based on all this information I have, or like planning, like given that there's a dog going this way and a car going straight at 35 miles an hour, what should you do next?
You don't actually need any photos in this case. Um, in fact they're like mostly distracting. Um, however, if you're testing perception, then having photorealis- realism is important. And so I think about this like with self-dri- with voice AI is you have this hierarchy of, you have text-based evals, which are great for workflows, tool calls, instruction following, um, and then maybe you just need simple voice to voice to test like [01:57:00] latency, interruptions.
Um, you don't actually need like different accents or anything like that because really you're just testing the timing of your own system, um, and its ability to recover based on like transcription errors, et cetera. And then hyperrealistic is useful for when you're testing audio quality, accents, background noises, and that's where we spend a lot of our time investing in...
at, at Koval is investing in simulation realism. And so we actually model out the audio. We don't just layer background noises with the audio, but we're actually simulating how does this audio interact with the world around it? Um, and so I think this is, like, a really good, um, benchmark for then how do you start to approach evals for voice AI as you're, as you're trying to choose between different models.
So it's really easy to become super overwhelmed by every model out there. Um, there's, like, so many... We actually run [01:58:00] benchmarks.coval.ai, which you should definitely all check out, and we benchmark, like, all the different providers. We have, like, leaderboards for text-to-speech, for speech-to-text. Um, and you can actually see, like, how all these different models are doing.
You can see also the, like, you know, how are, how are the deep RAM models, uh, which ones do they have available, what are, what are their benchmarks? Um, as well as we also have speech-to-speech, text-to-text. So that is a great place to start. Um, that's a great place to start because that allows you to, um, understand, like, which model should I even start to look at?
Um, but then the next step is, like, how do I benchmark this with internal data? Um, and then finally have end-to-end voice evals. And I think this, a lot of this, um, in self-driving too, like, given that you have so many models, it's really important to be, like, isolating your variables so that you know what you're testing with every test, [01:59:00] um, and that you're able to then, um, you know, not necessarily, uh, become overwhelmed by all the different variables at play
Um, something else that we did in self-driving is denoising. So this is a really helpful concept, um, especially when you have non-deterministic tests, which is when you get a failure, then you take that failure and then rerun that failure, like, tons of times because then you can see, like, okay, I know that this has failed once.
What is the probability that this is going to fail in general? Um, so now I can start to get that probability that we talked about around what is, uh, the probability that this will fail
Um, so I think now I want to k- kind of talk through, like, how to build a really scalable eval strategy, because based on all of these learnings from self-driving and robotics, and also seeing hundreds of different voice [02:00:00] systems at Cobalm, we've started to have, like, strong opinions on how to run a really reliable, um, a reliable eval process.
Um, and I think the, the... My biggest takeaway is that evals are one of the most... And, and obviously I think there's a little bit of I'm drinking the Kool-Aid, but I truly believe this, is, like, one of the reasons Waymo was so successful is because it had the most advanced eval strat- eval strategy and invested in simulation from day one that allowed it to scale to many different cities.
And you can brute force your first city, but it's very hard, but it's very hard to then scale to tons of cities beyond that. Um, and so similarly with voice AI, we see that it's very possible to brute force the kind of first iteration, um, and also make sure that, you know, it, it can, it can feel like investing in evals is, you know, takes more time, et cetera.
It's very similar to a unit test except for now with [02:01:00] agents, the evals are way more important because it's defining how is your agent actually... What is your agent optimizing for, especially when you start to add in self-improving loops. You're, you're defining the metrics and the product definition through your evals.
Um, and if you're not learning from the real world and, like, what you're seeing in the real world, then your product is just going to be significantly, um, hindered by the fact that you don't have this, this loop of evals. Um, and I realize, I think I'm at, running out of time, so I'll breeze through these last ones so that we have time for questions.
Um, but one important thing is, like, voice AI, it's really important to do speech-to-speech evals because it's not just about latency, but actually, like, how does your agent, um, interrupt... How does your agent deal with interruptions? How does your agent deal with transcription errors? How does it deal with background noise?
Um, and you might say, like, "Well, isn't that just the [02:02:00] models? Like, do I really have any control over those things?" Which you totally do. Um, it's possible, for example, to, uh, you know, add in like, uh, l- for example, like add in background noise redu- reduction, or maybe you add an LLM to correct the transcript based on the context.
Um, you can also do things like changing parameters, changing models, adding background models. So I think, um, kind of assuming that it's all up to the transcription and voice models takes away some of the ways that you can best create really advanced and, um, magical voice systems.
Um, and then something that I think is super important, I love like how Mohamed Hussein is a Eval's influencer and talks a lot about this, but it's so important that you look at your data. We have a way that in our platform you can actually correct all the data and label and annotate so that you can see your [02:03:00] calibration, your calibration scores of all of your metrics.
Um, this is something that we did a lot at Waymo is like actually, uh, aligning human judgment of the self-driving car with like how we actually simulated it as well as with like manual drivers versus the ML models. So being able to like have really easy ways to calibrate against human judgment is super important
Um, benchmarks are super helpful. So like benchmark every part of your stack individually, and then you can actually, um, uh, evaluate which models. Start with task-based evals. Um, how often do the tasks complete? Where do they break? How often? The reason I call this out is because actually having one really good scenario that you test over and over is better than having, like trying to do a million and then see where the failures are.[02:04:00]
Um And then I think creating an eval process, a lot of times people focus on setting up evals and they don't focus enough on, um, how do I actually like, uh, set up the like operations around it. So building an eval platform or buying an eval platform is kind of this, you know, it doesn't really matter the...
because you're ultimately going to build something very similar. But how do you set up these flows and these loops? That is the piece that like really differentiates teams. Like where is data flowing in from? Where is it flowing out of? Um, like what do you do with your compliance failures? Does a human review them?
Where do those human reviews go? Does it go into a self-improving loop? Does it go into an auto-generating test set? Um, does it automatically improve your metrics? These are really important pieces that are going to then like 10X you and really create like compounding returns from that. Um, so [02:05:00] yeah, I think, uh, we, uh, yeah, I'm just like so excited about the future of voice AI.
I think it's going to be the next platform where, you know, the next platform is going to be like web or mobile ne- then the next platform is voice. And so I think this is one of the reasons why voice AI is so exciting having been in autonomy for so long. So yeah, thank you so much, guys. Yeah. Excellent.
Thank you, Brooke. That was, that was excellent, and I think evals are always s- super underrated. Um, you know, I, I've had to fight tooth and nail to try to get evals just for agent, uh, standard agent production, uh, deployments. So I can imagine in audio when you're adding layers of complexity, like you said, um, it can also still be pretty challenging.
Um, one of the questions that came in from the chat that I wanted to ask, uh, is, um, is [02:06:00] there any way to get involved with open source contributions, uh, for Covall? Yes. We have our benchmarks. So open... Like, contributing, um, uh, metrics, contri- contributing co- code, uh, new models to our benchmarks is super welcome.
Actually, a really great... Some things that are top of mind in case you're looking for inspiration of things that should be added to the benchmarks is we... Today, we mostly do word error rate. I would love to add semantic word error rate so that you're able to say... One of the important things with voice AI is that it doesn't really matter if I say, like, hi in the beginning of my sentence and you drop that hi, because that doesn't change the semantic meaning of the sentence.
But it does matter a lot when you drop my name or my phone number or, um, when, like, you choose the wrong verb. Like, uh, all of these things are super important to the meaning of the sentence. And so rating models on how well they're able to do that versus just, like, can you get, like, the murmurings or utterances [02:07:00] correct is less important.
Um, so I would love for people to contribute there. Also, adding new languages, um, new metrics across the board are always welcome. Excellent, and I dropped the GitHub repo in our GitHub organization in the chat so people know where to find you guys, um, and I shared the link for some of the other things that you were sharing in your presentation earlier as well.
Um- And please star our GitHub repo because- Yes ... I would love for more people to discover us, and that is the best way for people to discover you. For anybody who's been here for the entire session, please make sure you go star all of the repos, all the organizations, uh, all across GitHub, um, to make sure that, uh, people, people get their, uh, recognition and others find it, uh, find what they're looking for, um, for these tools.
W- another question that came in from the chat is, um, do you keep the simulations static as a reproducible [02:08:00] set, or are they generated dynamically? Um, and if they're generated dyam- dynamically, um, how do you reliably, uh, change or measure the changes? Yeah. So we don't actually, um, necessarily generate them dynamically from scratch.
Like, we have some seed, and that's what allows you to ground in some reality. So if I, um- Like, we have these test sets, and you could have, like, for example, a, um, red teaming test set where we, like, have this system prompt. I like to think about this actually as a spectrum of, like, for example, at Waymo, you have everything from, like, totally synthetic test set to, like, you're just playing back logs.
And at Koval, we do the same, where we have a spectrum from a totally synthetic scenario all the way up through, like, you're playing back the audio. And then there's something in between where you have a transcript as inspiration, but then you can modify your, like, what you're responding. Like, we mo- can, um, modify what responds.
So if [02:09:00] you ask for an email instead of the identity first, we can switch those up. Um, we do... Something else that's really actually useful for red teaming is doing loops of simulation. So you can run the simulation and then, um, basically, like, r- ch- keep changing the test set. Uh, you can define the metrics.
So, like, let's say you're trying to reveal a social security number, um, and make sure your agent can't do that. You can create a test set that has a bunch of attempts, run that simulation, and then see if any of them succeeded, and if they didn't, then you can alter, like, then you could alter the test sets, rerun it again, see if any of them succeeded, and basically run for, like, some amount of budget to see can we find any issues.
So that's a really u- useful way to kind of combine, like, seeded test sets with more exploratory. Um, but yeah, you do have to, like, ground... If you don't, if you just send it off, then it's very hard to compare of, like, well, did we just never see this scenario before? Or, like, is this actually a new scenario?[02:10:00]
Yeah. And that, that kind of led into one of the, one of the thoughts I had, um, about these simulations, is thinking about what are the different variables and, like, what are the most important variables to consider as you're, as you're going through these simulations or these cases? And, you know, the first thing kind of stood out to me was this, like, initial conditions.
Like, what is the initial condition for, you know, each case? And, um, you know, do you give the agents a certain set of tools? Do you give, seed it some information, some context, that sort of thing? Do you agree that that's, like, probably one of the bigger levers, or do you think that there's other more important variables to consider?
Yeah. So you're saying like kind of creating like state, creating, um, like this is a user with a bank account and whatnot. Yeah. Definitely. I think creating the state internally is still a big challenge, actually across evals across the stack. Um, like you have an SRE agent, how do you create all these different simulated environments [02:11:00] for that SRE agent?
Um, and a lot of this comes down to traditional software engineering, whereas like being able to mock things, being able to create test databases and test users, and I guess the infrastructure nerd in me is excited about this because it's like all these things that felt maybe like not as important earlier, now with agents it's like if, unless you're...
You can either, A, be comfortable just launching it into the real world, um, which I think there's totally use cases where that makes a lot of sense, and it's not, it doesn't always make sense to be like, "I'm gonna create a super high fidelity simulation environment." Um, but it does when you're like building a self-driving car that could drive off the road or you're building a, um, you know, banking agent that could, has access to all your finances or you're building an insurance agent that could, you know, like cause all sorts of regulatory nightmares.
Um, or like you just want, I think there's like the also the semblance of a high quality product experience, like building a, building [02:12:00] a product that's glitchy and like constantly down, like turns away users. Like, people have a sense of quality, and I think that right now with AI there is like- You know, early on with AI there was, like, way less of a sense of quality.
Um, you know, it was totally fine for it to take a long time to give you a response. It was totally okay for it to be wrong all the time or, like, not work or whatever. And now I think even in the past, like, since, you know, GPT-3, right? Like, the sense of quality for, like, this sounds like AI slop, like, what previously would've been a, like, magic, um, a magic trick even in, like, 2023, now you're like, "Okay.
Well, this is not very impressive. Like, I'm really looking for something. Give me something more," or, like, "This is taking so long." Um, and so I think we're gonna see the same thing happen with voice, where every enterprise is going to be expected to have a voice interface. Like, if you have... If an airline doesn't have a voice interface, you're just not going to use that airline, in the same way where it's like if you can't buy a t- [02:13:00] airline ticket online or, like, use a mobile app, like, you're just not going to use that airline.
Um, same with banking, same with, like, a lot of these services. And so I think the expectation of quality and, like, that this agent is not just, like, an FAQ bot or is not just, like, a, um, like, script- scripted bot, but can actually take actions and do things on your behalf now takes this really magical interface and allows you to have...
Like, there's lots of flows where it's actually way easier to do things via voice. Like, "Hey, can I, like, change my flight? I'm running late for my flight, so, like, can I change my flight to blah, blah, blah?" Um, "Great. That's, like, $100." You're like, "Uh, never mind." Like, that's one flow, and you're, like, not already at your computer.
So all these places where you're not at your computer, voice has a lot of potential. Yeah. Absolutely. Fantastic. Well, thank you again very much, Brooke. Um, great, great presentation, and, uh, definitely go check out Koval. And if you, uh, are interested in doing any, uh, open source [02:14:00] contributions, head over to the GitHub and, um, check it out.
Yeah. Amazing. All right. Thank you so much, guys. Thank you. Take care. All righty. So I think that pretty much wraps us up. That was all of the talks we had scheduled for today, which is great. Thank you for all of the attendees, all of the, um, all of the, uh, speakers. Um, super appreciate, uh, everybody joining, sticking around and presenting.
Um, if this event was something that you were interested in and you really liked, uh, there is the, uh, Agentic AI Foundation Voice Agents Forum that's coming up in November, November 5th. That is based in San Francisco. Um, you can scan the QR code that is on screen right now, um, to get to the event page.
Check it out. I will also leave a URL in the MLOps channel chat. Um, [02:15:00] beyond that, there is also, uh, Agent Con, uh, live right now. Um, well, not right now, but starting tomorrow. Uh, that is, uh, Demetrios is going to be doing interviews and, uh, just kind of wandering around Agent Con. Um, so, uh, pop over to YouTube, um, scan the, uh, Q-QR code, um, and check those out.
That starts tomorrow. It's going on, I think until, uh, Friday. Um, and also for anybody, uh, interested on Friday, uh, there's a regular coding agents lunch and learn every Friday, uh, virtual. I'm usually there. Um, I'm a host, uh, there as well. So if you are interested in just coming by and, you know, sometimes we have speakers, sometimes it's just kind of a bunch of people get together, we talk about, uh, different topics.
Uh, we had a whole discussion on MCP. [02:16:00] Uh, I know there's an upcoming one on Agents.md. Um, and, uh, definitely, definitely encourage people to check it out. It is a very fun and very casual, uh, uh, you know, session on Fridays. Um, anything AAIF, uh, events, there are tons of them going on. There's the AAIF events page.
Please check that out as well. Um, and, uh, you know, see what's local, see what's virtual, see what's, uh, you know, some-something big that you may wanna travel to. Uh, lots of stuff going on over there. And last and not least, uh, there is the MCP, uh, associate certification that is now, uh, that has now launched.
If you had not heard about this, um, it is up and live. I'm gonna also drop a URL in the chat with that. Um, yeah, so check [02:17:00] that one out. Um, you know- Talk to us in the Slack. Um, if you're not in the Slack, check out the MLOps community Slack. It is always got lots going on. If you aren't here early, early on, uh, m- uh, Demetrios and I were sharing some of our favorite memes from the Slack.
Um, so definitely join over there if you have not already. I think that's gonna be it for every- everything today. Thank you all for joining and sticking around. Really appreciate it. Um, hope you all have a great rest of your day and great rest of your week. Uh, take care. Thank you