Memory for LLMs - Silicon Valley August 2026 Meetup
Speakers

Heather is an experienced Developer Advocate and engineer passionate about distributed databases, cloud architecture, and modern backend systems. At YugabyteDB, she helps developers build resilient, scalable data infrastructures for modern applications.

Ben is a technology leader and entrepreneur focused on building intelligent healthcare and enterprise AI solutions. With deep hands-on expertise in machine learning and backend architecture, he specializes in bringing complex AI systems into production.

Arshan is an entrepreneur and go-to-market leader with experience spanning strategic partnerships, healthcare technology, and biomedical innovation. He has built and led early-stage ventures and brings a founder’s perspective on turning emerging technologies into practical products and scalable businesses.


Rahul Parundekar is the founder of AI Hero. He graduated with a Master's in Computer Science from USC Los Angeles in 2010, and embarked on a career focused on Artificial Intelligence. From 2010-2017, he worked as a Senior Researcher at Toyota ITC working on agent autonomy within vehicles. His journey continued as the Director of Data Science at FigureEight (later acquired by Appen), where he and his team developed an architecture supporting over 36 ML models and managing over a million predictions daily. Since 2021, he has been working on AI Hero, aiming to democratize AI access, while also consulting on LLMOps(Large Language Model Operations), and AI system scalability. Other than his full time role as a founder, he is also passionate about community engagement, and actively organizes MLOps events in SF, and contributes educational content on RAG and LLMOps at learn.mlops.community.
SUMMARY
As Large Language Models evolve, the real challenge isn't just generating text—it's remembering context, maintaining state, and managing historical data effectively. How do you provide LLMs with persistent, long-term memory without overwhelming context limits or skyrocketing latency and costs?
Whether you're building stateful AI agents, implementing advanced Retrieval-Augmented Generation (RAG), or managing enterprise-grade vector and relational data, this meetup covers the practical architectures and trade-offs behind modern LLM memory systems.
TRANSCRIPT
[00:00:00]
Hello, everyone. Uh, good evening. Welcome to today's meetup, Memory for LLMs. Uh, my name is Vijay Bohre. Uh, um, uh, I'm here with my co-host, Rahul Parundekar. Uh, we are happy to, uh, and we are thankful for you to-- uh, for joining us tonight. Um, first of all, let's give a round of applause to YuvaWhite team. They have been very, um, uh-
Thank you for giving us, uh, food, uh, hosting us and drinks. Um, so, uh, I'm from AAIF, uh, along with my co-host, uh, Rahul Palandeker. Uh, we'll talk a bit about AI before we start our talks today. [00:01:00] Uh, Rahul, do you mind sharing?
Oh, I'm sharing through this? That's perfect. Yeah. Second approach
Amazing. Um, welcome everybody. Uh, it's very nice to see you guys here. Um, this was a very convenient location for me. I walked one block away from home in San Francisco, and I walked one block here, so it's fantastic. Um, let me give you a quick introduction. Uh, my name is Rahul. I am a AAIF ambassador for San Francisco.
But, um, before that, I used to be a organizer for the local MLOps community chapter. Um, the Agentic AI Foundation, I'm, I'm gonna give you a quick kind of introduction if you don't already know. Show of hands, who has heard [00:02:00] about the Agentic AI Foundation before this event, obviously? Okay, great. Um, so we, we, um-- Who's here heard of MCP?
Ah, wonderful. So that's, that's where MCP goes. MCP is home. The stewards of the MCP protocol, along with agen-- um, uh, Goose agents.md, if you write cloud code or whatever coding agents, the .md file that you provide your instructions in, that's also stewarded by, um, the Agentic AI Foundation. So, um, the goal of the foundation, which was started in twenty twenty-five, is to, uh, advance the frontier of Agentic AI.
So what we really wanna help is get smart people in the room, talk about Agentic AI, and talk about how do you keep this open, fair, and, uh, not do evil things, right? Um, and this kind of a community platform is one of the very, um, unique ways in which we are [00:03:00] getting people in the room to have actual conversations about building real agent systems and deploying them in production.
Um, there are four very well-known projects, and I think there's five now. But, uh, Model Context Protocol, so MCP. There's Goose, which is a harness, um, that, that, that's open source. It's part of the Agentic AI Foundation. Agents.md, the file that you use to, uh, write instructions to your coding agents. Um, the Agent Gateway, which is, uh, intermediate layer for you to communicate with different agent, uh, or LLM providers, et cetera.
And now in the recent last week or two is, uh, A2A. So if you've heard about the A2A protocol, um, that's also now stewarded by the Agentic AI Foundation. Um, and there's a bunch of members. There's over two hundred members, um, of the, of the Agentic AI Foundation, and there are very well-known companies that are all [00:04:00] contributing with, uh, you know, uh, there's working groups, and everybody's contributing to, like, what is the right framework to have to promote open agentic processes.
For example, one that I'm really interested in, though I don't contribute, is how you have skills inside MCP, right? So there needs to be a forum for people to discuss this. That's where all the Agentic AI Foundation work happens. And the Agentic AI Foundation is also part of the Linux Foundation. So it's, it's all open.
We are vendor neutral, and we love to support community. In fact, we love community so much, we now have eighty chapters all over the world, right? So if you see a dot here that is probably, you know somebody there, you can tell them like, "Hey, by the way, did you know that you have an AI chapter?" Um, we're trying to put together events at least once a month in each of these chapters, um, all brought, brought together by amazing, uh, organizers, uh, like Vijay and me.
Uh, sorry, like, like Vijay, maybe not me. All right. Um, so here's a quick QR code. We have a Slack [00:05:00] channel where there are thirty thou-- thirty-two thousand people who were previously part of the MLOps community. You can join the Slack channel, um, have good conversations. Once you join the Slack channel, there's an, um, there's a general channel that you'll be added to, and you can click into the country channel or the chapter channel for you.
So now you suddenly have access to everybody here in the room who's now there going to be in the chapters, uh, uh, on Slack as well. Um, but this is not our only event. We are planning more events, and the Silicon Valley or South Bay chapters one, the San Francisco chapter, which I also organize events for, is, is, uh, doing a collaboration with Datadog and Clariq i-in, on, on September 9th.
So if you, if you, uh, wanna make the extra, um, scenic route to San Francisco, um, feel free to join us at this event. There's more events, though. So there's a big one. We're gonna have a North America, um, agentic event, [00:06:00] so it's gonna be Agent Con. Um, make sure you see that there's no E inside agent, so it's A-G-E-N-T Con and MCB Con North America.
This is gonna happen in, um, San Jose. So very close by to you guys, and it's gonna happen in October. Um, so please scan to register. Um, and I was supposed to add the code here, so I forgot about that. But there's a discount code that I have for you. Give me five minutes after, and I'll come and yell what the code is in a bit.
Um, I think it's community twenty-five, but I'll just- That's community, correct. Community twenty-five, all caps. Yeah. Right? Um, so the code is community twenty-five, all caps, to get twenty-five percent off on the registration. So we'd love to have more of these kind of large conferences as well, where people all over the world can join.
Um, for example, North America, everybody can come together, support open agent standards. Um, and that's kind of it, you know. Um, that gives you a background about, uh, the [00:07:00] AAIF, um, the local chapter and you know, what we're doing globally. So looking forward to seeing you at our next event. Thank you.
Thanks, Tom. Yeah, so tonight's, uh, talks, they focus on, uh, what happens beyond, uh, prompt engineering, uh, specifically state management, uh, distributed data layers, and persistent memory systems. Uh, let's say if your agent doesn't remember what happened five steps behind what you-- You don't have agent, but you have a extensive chatbot.
So here, uh, we have, uh, three technical sessions lined up. Uh, we'll, uh, start with our first, uh, speaker. Uh, he's a CTO and founder at Ferosa and Grounded Health. He specialize in-- He specialize in, uh, ML and backend architectures for high-reliability [00:08:00] healthcare and enterprise applications. Um, yeah, yeah, please welcome Ben Kunz
Where we at with sharing?
Yeah Oh, I'm Ben. Um, it's great to be here. Uh, I'm gonna do a little bit of a brief introduction to myself, but, uh, thank you, Rahul. So r-- for those who don't know, Rahul and I used to work together, and we deployed a lot of small language models back in the early days. And Rahul was very, uh, inspirational to some of the products that I'm about to-- like, we're about to launch.
Um, and I also, my other co-founder here, uh, JL, uh, we- also worked with Rahul. And so we're here to talk a little bit about how to go [00:09:00] from goldfish, which if you have actually worked with models without a harness, they're rough. Like, they don't actually do much. You're working with a probabilistic engine, um, that is trained on the corpus of, like, humanity.
And so your results are gonna be random. Um, so yeah. Uh, I've built a lot of new things. Um, I don't generally talk and speak in public a ton. Um, I don't do lots of PR. You probably have never heard of me. This is on purpose. Um, but I've done a lot of different things. So one of the things that I did is I built a distributed energy storage platform.
So we took over Tesla's Gen 1 stationary platform, and we were the first people to bid in the wholesale market in the Cal-- in the state of California. Um, we ran part of the power grid in Hawaii and a bunch of other things. So I just run big distributed systems. And then I mentioned a little bit where JL and Rahul and I worked was Figure Eight.
Uh, so Figure Eight, what we deployed there was very interesting because we put a lot of small language models out to do [00:10:00] ML assistant data labeling. So all of the labeled ground truth that are used to train these models, or a good portion of it, came through Figure Eight. It used to be called CrowdFlower.
Um, I've done other things, like there's a couple of people that I used to work with, includ-including my former boss, um, who's doing some AI for, um, physics simulation. So she's got an, a, a foundation model. Uh, but we did-- we built a consumer home robot, and so we actually deployed YOLOv1 on a hundred and fifty dollar micro that ran around your house and could detect faces and could actually tell them apart.
So, like, I am a huge fan of different size models running across different things. Um, so one other thing that is also relevant to this talk is I spent a while at DataStax. So this is a database thing, like we were doing memory. So I ran open source and enterprise engineering. So I ran the DSC project for a-a-about a year.
Um, and I was responsible for, uh, all of their [00:11:00] existing business lines, hundred million dollar worth of business or something like that, as they were pivoting to Astra, which was their database as a service product. Um, so, like, that's a little bit about that. Um, so I left a robotic parking company that Chris helped us market, and I was like, "Oh, man, I'm really tired, and I wanna do something different."
Um, so I'm just gonna play around with this AI stuff. Seems kinda interesting, like maybe it'll go somewhere. Um, uh, and so I started working with some friends. We built, uh, like a RAG application. So we put a product manager in a box, um, as like everybody probably does a RAG application. Doesn't work extremely well, so you have to kind of be- expand beyond that.
We, uh, did-- we built a couple of apps. One of the apps that, uh, like Mark Cuban just tweeted about is Grounded Health, which is healthcare analytics, uh, using AI. So how do you, uh, actually reduce the cost of healthcare? Um, so as you can tell, I work on lots of large problems. And I also do some fun stuff. So, um, I play in a community band in the city, [00:12:00] and I di- didn't like any of the apps for, uh, practicing and rehearsals and other things like that.
So I was like, "Ah, just write my own." Um, so I have a little iOS and Android app, so you run it on my iPad and other things like that. So that's a picture of that in the corner. Um, so I, like everybody, started off with, uh, flat memory. So Rahul graciously enough gave me access to his research repository that had a whole bunch of skills in it when I was playing around.
Like, he and I went, and I think we had some wine, and there was a bunch of stuff going on. Um, and so we-- it was a linked, uh, file system store. We had a whole bunch of things in there. It was basic skills. It works-- worked great. I was like, "Oh, this is awesome." And I started creating a whole bunch of skills because, um, everybody who-- anybody who's ever worked with me knows that I have a book for everything.
I have a PDF of those books for everything. So I fed those into LLMs. I had them summarize it. I then used that with some context, and I probably have, you know, a thousand, fifteen hundred skills to do a whole bunch of different things. Everything [00:13:00] from security and stride analysis to graph and ontology modeling.
Um, and as, uh, JL has been through my, uh, personal, like, archive of things, he's like, "You have a music skill in here." Um, so things like i- it spans a whole bunch of passion. Um, started having a hard time getting my agents to stay on target because they started to look a little like me, a little ADHD. So I was like, "Oh, I need to kinda fix this."
Um, so I started, like, playing around with, like, what can I do to, like, only give parts of context around. Um, and at the same time, uh, one of-- I had a conversation with the chief product officer at DataSax when I was there. I was like- Databases in Java. Why would somebody do that? Um, I'm a-- I'm a-- Like, I wrote Java for years and years.
I-- Like, one of the startups I was at did Java before Java was one point O, and we had the first hundred percent pure Java application. I still was like, "Why would you write a [00:14:00] database in Java?" It's like, "Somebody should just rewrite it in Rust." And so in January, I was like, "Well, I have some time on my hands," like, "We'll see what happens."
Um, so I started prompting it with, like, how I like to do software development and test-driven development, like, all the t-tools and techniques from the beginning of time 'cause I've seen the beginning of time. 'Cause, like, I was-- like, started doing software development in the early days of the internet. And so I started pulling all the best practices together, and I, I didn't just say, "Rewrite, uh, Cassandra," 'cause that's a dumb way to do things.
'Cause if you've ever looked at the Cas-Cassandra code base, there's cycles, and there's, like, bad code 'cause it's open source. Open source is free as in puppy. So bring it home, choose your-- chews up your leg, pees on your couch. Um, so I built a bunch of skills to actually make it so that I could refactor the code and rewrite it and make a modular database, and I wrote it in Rust.
Um, so that's where this started. It started off as kind of a little bit of a side project. [00:15:00] Um, and along the way, like, I debated like, "Ah, do I do a database-as-a-service company?" And as I told JL earlier this week, I'm allergic to raising money. I was like, "I don't really wanna do that," 'cause in order to do a DBaaS company, you have to raise a lot of money, um, 'cause you have to run infrastructure, and you have to-- you're selling nines of availability.
So I was like, "All right. Let me play around with this MCP thing, and let me, like, play around with some agentic memory." Um, and then back in the nineties, I also had this passion for ontological systems. So how do you do computable knowledge stores? So how do you actually create a rule that can infer where this knowledge came from?
What does it mean? Um, so I started, like, putting all the things together in the database that I wanted 'cause I'd already written the Cassandra part. I was like, "You know what? I wanna put some Sparkle in here." How many know-- How many people here know what Sparkle is? Ooh, well, one person. So Sparkle-- How many people know who Tim Berners-Lee is?
Okay. After he invented the internet, [00:16:00] or after he invented, like HTML, 'cause this is the guy who invented HTML. He was like, "HTML is a shitty language to..." And pardon my, my, uh, language, 'cause, like I swear a little bit. Um, but, so he was like, "HTML is a crappy language for sh- computers to read. It's sprinkled with, like visual elements, it's sprinkled with layout."
Anybody who's written CSS is like, "Why would you ever do this?" And so he created this thing called RDF, which was a Rich Data Framework language. And it was encoded in things called a tuple store. And unfortunately for him, it was in XML, 'cause like nobody likes XML either. Um, but there was a bunch of databases that came about, and there was a standardized protocol, it was called SPARQL.
Uh, so SPARQL is an ontological web database that allows you to infer knowledge on top of it. Um, as far as I can find, no one fully implemented the protocol ever, because humans are crappy at reasoning on graphs. [00:17:00] Like graphs are hard. Um, graphs, like everything links, you get cycles, like how do you reason on a cycle?
All that other stuff. Turn-- And so I had this thesis, I was like, you know what? LLMs with their probabilistic, like inference, maybe they can help write some good rules on how ontologies get formed. And so I was like, "I'll put some SPARQL in here." How many people know what Datalog is? How many people know what Prolog is?
All right. So I got a couple, got a couple of hands in the back. So Datalog is a-- It's a derivative of Prolog. Prolog is a different kind of programming, uh, technique. Um, it's called reasoned reason. It's like a reasoning engine. So you, you, you don't write top-down declarative, you don't write object-oriented, you don't even write functional.
You just literally say, "This is this," and then you put it in a reasoning engine, and the rules run. And there's a whole bunch of optimization techniques that make this run re-- that make this run really fast. Um, and so I can say, "Data from this location is [00:18:00] knowledge," or, "Data from this location is financial."
And then anytime data gets put in there, it gets tagged as financial, and so I can query off of it, like, immediately. It's different than querying SQL. I don't have to write select star from data source and then go find it. I can just, like, it's just tagged. Um, so I added that to the MCP thing, and now I've got, like, more than...
And then everybody has to add vector, so I also added vector, so there's some nomic stuff in there. Um, so there's a whole bunch of stuff that, uh, this kind of come together. So still not really commercially viable, trying to figure out what, what's going on. Th-- None of this is still viable at all, um, I, I don't think.
So there's a SQL database with graph traversal. There's vector indexes, S3 backing, uh, inspectable substrate. There's an MCP native. I actually implemented MCP two point O for Revel, so maybe one of the first MCP two point O, uh, compliant things that are out there. Temporal facts. So you can do [00:19:00] temporal reasoning.
This happened before this. This happened after this. Um, you can fold, which is, like, the thing you do when you dream. You actually consolidate memories. Um, and then I added a task framework 'cause I'm- Doing a whole bunch of different projects for lots and lots of people. And then as we all need to do, your agents need to forget.
Like, how do they forget? Um, and so over the past couple of months, I've actually worked on some additional ontological stuff. Um, so the goal of this is to create knowledge that survives this session. Um, so for me, I lived through one of the Claude outages, and I kinda went through withdrawals 'cause as many of you here probably have, there's a little bit of AI psychosis that goes on.
Like, "Oh my God, what happened? Like, Claude went away." And so I was playing around with, uh, OpenAI, and I was like, "Oh, it just doesn't work the same, and it doesn't know as much about me." So one of the reasons that happens is, uh, and we were talking a little bit about this in the back, um, they're running m- memory servers for you, that they [00:20:00] control for you, and they know a lot about you, and they know a lot about your enterprise.
And so as you start to interact with these closed models and these closed companies, you actually are giving them access to some of your biggest, most precious things that are going on. Um, and you also then can't switch, so it c- creates your switching costs. Uh, it makes your switching costs very high.
And having come from the small language model world, I was like, "No, I wanna be able to move around." Like, I believe that local models are going to be where things go. Um, Apple has a whole bunch of stuff that is, like, about to land. So I went out and I bought, like, a hundred and twenty-eight gig RAM laptop from Apple.
Spent way too much money, um, just so I could play around with all the new models that have come out. So if you ever c- hang out on, uh, Rahul's Friday talk, um, generally I have a new model that I've played with that is, like, local on either that or my gaming PC, um, and then one or multiple of the, the other models.
But in order to make that work and make the, uh, the [00:21:00] process of that actually reasonably transparent, I needed a, a shared memory that I controlled, um, that was my memory, that it was my memory service and I could actually curate. So that's how this kind of came together. And so as I've started to do this-- to, to work with this, I've built a bunch of things, including building Ferrosa with Ferrosa.
So all of the stuff that you've seen here I build with my own tools. So I, like, have the MCP servers that actually help me build and they, and they can remember what the tasks are. Um, how many of you have ever had Claude say, "I decided not to do this. I will do it later"? It's frustrating as hell. So I wrote a stop trigger.
So whenever it does that, it write-- it creates a deferred task in the database that I can then do a slash what's next, and it'll cons-- and I actually have it go through, and whenever it does that, it will do a triage prioritization level in the LLM. So now it sticks that in the database, and all of a sudden I've got this task list that's auto-curated, and it doesn't forget.[00:22:00]
And that forgetting bit was like I would-- I was getting surprised because I was like, "Oh, this feature is done. Claude told me it was done." And then you go back, and you're like, "It didn't wire that? What the hell?" Like, you weren't done. Like, you really needed to continue. So this is where memory, um, is more than just, like, your session trajectory.
It's more than just the skills that you're in. It's more than just, um... It's like how do you take the output because these LLMs create a huge amount of data exhaust. How do you take that data exhaust and then classify it to make it useful for you as you go forward? Um, how many of you have actually gone and asked your agents, "What could we have done better?"
Yeah. So I actually implemented stand-ups with my agents at one point, where in the morning we would be like, "What are we gonna do today?" In the evening, like, "We should do a postmortem." So I would like compress a week into a day, and that's a really good way to interact with your agents because that-- part of that will help you [00:23:00] build better skills that will then make your agent better over time.
So implementing like extreme programming and rapid development types of things is really good. Um, and then my database started getting big. Like I think I have eighty thousand nodes in my database right now or something like that. Some of it is like testing 'cause I like have done some, uh, uh, evals on some of the big data sources.
So I had to actually figure out how to forget, but I didn't wanna just forget, I wanted auditable forgetting 'cause like why did I forget that? Was it, was it something that was retracted that was wrong, or is it something that expired, it's no longer valid because I changed that code, or is it, uh, other things?
So this is part of what, uh, I'm building and still building. Most of this, uh, lo-- uh, this stuff gets used fairly frequently. So layer by layer, I kinda went through it a little bit. So there's the Ferosa database that uses an LSM, uh, uh, database, so log storage. So it's a append only. Very similar to Cassandra.
It actually reads [00:24:00] Cassandra tables, so it's all SSTables. Um, five minutes? Oh, shit. All right. Uh, there's a memory service layer, MCP, and then there's the, the... The MCP layer that I didn't mention is called Forge. So it has a bunch of static analysis tools. So it does like breaks up your code and finds design tools.
Um, so the parts that make Recall trustworthy, like temporal facts, like there's an automated dream cycle that runs. Hybrid retrieval, so there's eleven signals that go in, including like a PageRank, and you can configure it so that there's a small language model that ranks your results and sends it back, so that it can re-rank automatically, and it will ask your agent like, "Were these results valuable?
If not, send me like, uh..." It's a, uh, a three-state, it's a tri-state. So it was not, not unus-- it was wrong, it was not useful, or it was very useful. So- Uh, then audible forgetting. So five verbs for every setting. So it's like these are the MCP tools that are there. Um, it's a very simple MCP tool. I've also implemented, [00:25:00] uh, uh, the back-- the backside of this is also, um, progressive disclosure.
So there's tier one tools and tier two tools. So I'm trying not to clutter context with too many tools that are there. So this is like these are the tier one tools. There's a bunch of other things that you can have your LLMs call, and they'll find them if, if necessary. Um, so when you're building a memory system, you kinda wanna think in this hierarchy.
So this is a common hierarchy for ontological build. So data is like raw, raw stuff that you only want to have s- uh, scripts go over. At the next level, you take those raw-- that raw data like this could be sensor data 'cause sensor data is super noisy. This could be, um, raw session data, trajectories, other things like that.
Run reports on it, and it creates information. So information is summarized data. Um, that, that information can be valuable or not. And then when you have your information analyzed by your agents, it can create what looks like knowledge. Sometimes it [00:26:00] hallucinates though. So the normal model is data, information, knowledge, and wisdom.
So what we put in the middle is kinda claims. So this is the artifacts that the LLMs produce that still has to be reviewed. So a human takes a claim and converts it to knowledge. Um, knowledge is stuff that you can share with other people. Knowledge might have an expiration time. Things that I tell you today might not be true next month, so you can actually expire it.
And then wisdom is timeless, and this is hand-curated. So this is where you want your skills and how you do things in your process, um, around building. Um, so that's, that's always true. Not to say that it can't be revised as you get smarter and get, uh, older over time, but wisdom is what you actually want to impart on others.
Um, and so like wh-what we're also building is I, I like-- So I believe that one of the things that databases have is like as your data gets put somewhere, you have gravity. This is why Amazon and GCP and Azure have so much market pull. [00:27:00] 'Cause as soon as you start putting data in there, it's really hard to migrate, and it's really costly.
So you don't need to move the data that was at the bottom level. You just need to figure out how to summarize it. So if you have a connector that like connects to your data, summarizes it into knowledge, you should control that knowledge 'cause that is like the, the information and the knowledge are critical often to your business, and you might wanna compartmentalize that out.
So what we're doing, what we're trying to build now is peer-to-peer sharing of information on up. And so, uh, uh, I think it was the next slide or slide up So there-- like with sharing, there's some trust in here, so I'll, uh, like, you guys can read this. Uh, so one of the things around knowledge is like, uh, if you were to share knowledge with me, I may, may or may not trust it, but I might wanna use it.
And so being able to say trust, but verify knowledge that's coming from other sources, um, this is the [00:28:00] whole, uh, when you ask Claude to open a GitHub, uh, repo, it's like, do you trust this repo or not? 'Cause you're literally letting them run untrusted code. That's the same thing with, uh, knowledge that gets shared because prompt injection can happen.
So you wanna actually put it in trust and put it in a sandbox. Um, so what we're working on releasing in the next couple weeks is supervising your agents in sandboxes with segmented knowledge all peer-to-peer. Um, and if you wanna see more, um, come find me at the end, and we'll, uh, I'll give you a little demo on my phone of some stuff I'm running at home.
And if you wanna join us and be on the, uh, 'cause we're doing iOS, uh, and Android, and there's a desktop app for Mac, Linux to come. If there's demand, maybe some Windows. But this is our Discord channel. If you join now, there's I think like only a handful of people in there 'cause it's been basically me in my basement coding away like mad, um, [00:29:00] uh, trying to figure out what we're gonna do.
So that's all I got
Sorry, I didn't leave time for questions today We can have one question. One Thank you. Sorry, the, the link is expired in Discord. Ah. All right, come find me and I'll, uh- That's why you have 12 people That's why I have... No, somebody else found the expired one on the website, so, like, I need to fire the marketing person, which is one of my agents.
So- ... uh, come find me and I'll give you the Discord link. What's your favorite, uh, automation you've done in the personal side of things? My favorite automation, I actually really, um- That's true ... so- I've managed a very large team of software engineers in my past. Like, I've managed like a hundred, hundred and fifty people.
And, um, I-- when working with agents, I find that I can get them to do what I want them to do without arguing about why it's [00:30:00] right. Um, and, uh, that to me is just like, oh, it's so freeing. Like, I don't have to argue about, like, why do you do test driven development? It's like, oh, it's too much work. My agents, like, sometimes they complain 'cause, like, Claude does complain about, "Oh, that seems really hard.
That's too much work." Mm-hmm. But I can be like, "No, just do it," and he's not gonna, like, go slam the door and walk out of the building, and I don't have to, like, be nicey nice. Um, so that's my favorite, like, bit is just, like, building tools to build the process 'cause I love building process. So building tools to build the process to make reliable software, 'cause I've also run large systems as well.
So, like, I ran, uh, BestBuy.com. We, we ran American Airlines and financial services when I worked at Verizon. Um, so, like, I believe that compute should be like a dial tone, so you need to make reliable software. Um, so the process of making reliable software is hard. So, like, building the tools to build the process has been the most fun.
Thank you. Come find me after. We can have one more question. [00:31:00] Oh, let's take one, let's take one. I think I can just... Hey, thanks so much for the pres- nice presentation. Uh, I just wanted to ask, like, what is your key approach to deal with hallucinations, uh, going up that? Uh, the key approach to hallucinations, um...
Oh, you're gonna give me the, the, the thing back. Um, AI is going to hallucinate. The question is, can you tell, and do you have guardrails enough that you catch it and push it back into place? So it used to be when we were all hand coding stuff, um, that you could get away without tests because humans were looking at it, and you did a lot of manual code review, and I would look at your code, you would look at my code, and like if you understood my code, and it s-seemed to meet the requirement, that was good.
Um, in the agent world, I believe that if you can't tell what it should do with a test, you don't know if it works or not. [00:32:00] Uh, so that's how I handle this. It's needed.
Yeah. All right. Thank you. Yeah. So that's how you handle halluc- hallucinations, 'cause they will, they will still go sideways. Um, one of the first things I did build was a harness, and if you try to build a harness, really hard because some of the models are distilled from OpenAI, and some of the models are distilled from, uh, Anthropic.
They call tools very differently, and so you have to actually do tool call mapping. So the way that Anthropic does this, so if you've ever read- if you've read the analysis on the Claude harness, there's a whole bunch of if/then/else logic in it. So it's algebraic logic that, like, catches the hallucination and converts it to, like, what it's supposed to be.
'Cause it's... The hallucinations are generally stable. Like, it won't-- it'll hallucinate the same thing, um, so if you can tell what it's hallucinating, and you can remap that to what you actually want it to do, you're good. Does that make sense? Yeah. Yeah. [00:33:00] All right. Thank you, Ben. And so our next speaker, uh, he has spent years, uh, taking, uh, emerging tech into, um, uh, emerging tech ventures to market.
Uh, he'll be sharing practical challenges and, uh, counterintuitive, counterintuitive lessons from building long-term memory. Uh, please welcome Arshan Moghadam.
Um, before I start, I just want to get an idea of who we have in the room here. Um, who here is not a developer, like, at the core of who they are? Guess we got a couple of hands raised. Right. Can I get an idea of the different industries we have here? Maybe starting over here, just shout it out. What industries are we coming from?
Insurance. Insurance, finance. Blockchain. Blockchain. What over here? Autonomous vehicles. Okay, cool. Finance. Finance. Another finance. Okay, awesome. [00:34:00] So before I introduce myself, I want to start with the thing that makes this topic particularly interesting to me and also our company. Uh, you can build a working memory system for AI agents in about probably less than ten lines of code.
Uh, extract some facts, embed them, store them, search them, done. And this is actually exactly why memory is one of the hardest problems, uh, for AI agents. And so why is that? Well, because everything breaks. When every-- when things break, they break silently. And this is what we discovered in production. So nothing throws an error, and you don't get paged at two AM.
So your agent just quietly gets dumber, leakier, and more expensive over time. And by the time you notice, the damage is already baked deep into your store. So I'm Arshan. Um, I'm not traditionally from the AI, uh, uh, AI industry. My background is actually in med tech and in pharmaceuticals. [00:35:00] Um, I work very closely with physicians and healthcare.
And one of the thing-- one of the challenges that we saw is that as medical professionals that are working with patients are relying on AI companions to remember the medical histories, what the consultations were, what medication they're on. If we don't have the right memories that persist at the right time for the right patient, it can cause harm and death.
And so this is one of the things that got me really interested in this space. And, um, I found myself at Memzero. And so now I help lead commercialization and partnerships at Memzero. Um, I sit right in between the engineers who build this stuff and the customers who break it in production. And what I'd like to do today is talk about what I've seen from real teams that are using our memory solution, memory solutions in general, and are doing it at large scale.
So by show of hands, who here has heard of Memzero? Okay. Who here is currently [00:36:00] building with Memzero? Okay. So a lot of people heard it, some people have activated, some people haven't. All right. For those of you who don't know, um, much about Memzero, uh, we're one of the most highly adopted, um, AI memory layers right now.
Um, we have over a hundred and sixty thousand developers, um, building with our memory layer, uh, over a billion API calls and getting closer to seventy thousand GitHub stars. So a year ago, talks like this, um, once started with, here's why your agent needs memory. In fact, when the founders started our company several years back, um, it was very hard to have a problem statement because we were really early into this space of memory.
Memory wasn't really a problem and a challenge back then. And now the conversation has, uh, you know, shifted, and people aren't asking this question as much. But rather the new question is, how do I get quality memory into my agents? And so in this talk, I'm gonna discuss some of the counterintuitive things that we had to learn and overcome at Memzero to be able to provide highly accurate memory, and we'll [00:37:00] define what we're talking about by accuracy, um, at very low latency, fast, and also at low cost.
And this is the triangle that we try to balance. And I'm sure people have heard of other memory solutions. They're trying to build their own. You can use different models that can help you get higher accuracy at the cost of latency or cost, or vice versa. And so the core problem is simple. LLMs are mostly stateless, as we know.
They retain almost nothing beyond the active session. And memory is how we add stateless. So not just across session, but across devices, across users, across LLMs. And there are three types of memory. So the first one is, uh, memory by duration, content, and structure. So I'm gonna give just a little quick background.
So by duration, there's two types. We have short-term memory, which we define as what stays in your context window, and then long-term memory is everything that survives outside of it. Um, in terms of the content, we [00:38:00] have episodic memory. So what happened, I went to this restaurant last week, for instance. Um, semantic, so I've gone vegan, or I like running.
And then we have procedural, like SOPs. So never use them dashes in my email drafts. I'm sure everyone can relate to that one. Uh, we have memory by structure. So, um, vector and/or graph. Um, we'll get a little bit into details of those two. And then we can also scope our memories. And so this is kind of a newer, newer thing that we're working on is this idea called custom scoping, where you have-- multiple people can have access to the same namespace.
How do we have the memories, uh, in that namespace be added and retrieved with having the right ontologies of who can add what memory at what time and who can retrieve what memories at what time? So the memory goes far beyond just with the individual person And so now that we have the basics down, I'm sure most of this room have the basics down.
Uh, let's talk about some of the counterintuitive problems that our team ran into, um, when building these systems. So the first one is [00:39:00] that the write path was actually more difficult to build than the read path. And so when we think about the conversation that's happening in the raw messages, we can extract memories from the conversation per turn or per several turns.
And the extraction that's actually pulling those facts, we wanna have the right, what we call custom instructions, on what we want to extract, what we want to ignore, and what we want to never store for security reasons. Based on how we configure these, the LLM behaves differently based on whatever objection or, um, uh, target we're trying to hit here.
And so the LLM here, if you use a heavier LLM, it's gonna cost more. And obviously, if you use a lighter one, it's gonna cost less. And how do those extractions actually impact the quality of your retrievals? 'Cause quality of your retrievals is really what we care about here, right? And the upstream piece of that is actually the quality of your extractions and your storage is what we found.
And so we categorize our memories, [00:40:00] we summarize them, we deduplicate them before they get into, um, our hybrid database, and we make sure that when anything goes into our database, it is extremely high potency of data per character or token, if you will, is deduplicated. And during this phase, we have memories that are similar.
They start to merge together, or they start to supersede. And so we focus a lot of our efforts in terms of what is our add function or our extraction. And so this is something that it took us a while to actually realize and come to because we thought, okay, like if we're retrieving these memories, then how do we build the right system for retrieval?
And it all happens with how you're actually storing the memories to begin with. So maybe it seems intuitive now, but back then it wasn't that intuitive. Uh, we have a, uh, hybrid database that we use. Um, so, uh, vector and also, uh, entity and, and keyword search. And we have found that those-- that hybrid system works best in terms of balancing [00:41:00] both accuracy, um, uh, cost, and latency.
So the read path is an interesting one. So after we essentially send a prompt, um, we do a search, and this is the, this is the, this is the second piece to this, where we have this search has to get ranked. So, um, Ben did a good job of talking about the idea of re-ranking or ranking. And so when we think about what's happening in what we call top K, so let's say you have your top K, um, equal to ten.
You're gonna retrieve ten of the most relevant memories that are scored, that we normalize from zero to one, right? The question comes, is out of these ten things, out of these ten memories that we retrieved, if we were to set a threshold at anything below point three, we're not gonna actually bring into our system, right?
And we don't get anything that's above a point three, and we have that as a hard, hard-coded line, then it's not gonna actually return anything. And so we have to think about re-ranking as a dynamic system, which was not something we thought about at the beginning, where you don't actually hard [00:42:00] code anything below this threshold, don't pull out.
But how do we actually have this threshold move based on the memories that are, that are ranked? So let's say we wanna pull a couple memories. We have one that's ranked point nine out of one, which is a high-- highly relevant memory, and the rest are all point threes, point fours, point fives. But our threshold is set at point four, and we have one that's a point five, and the gap is point four.
So statistically, do we want that to actually make it in or not? And so how does this window move? So that's something that we're working on Um, in terms of other, uh, uh, hidden, uh, memory rot problems that we ran into. Um, so this is obviously one. We have, uh, someone prefers a window seat and always books a window.
So do we wanna have multiple things that are trying to accomplish the same thing in a memory bank? The answer is obviously no. And so we gotta merge these things together. And so we have a, um, a good merge solution here that, um, finds things that are, you know, uh, semantically similar, and we essentially [00:43:00] gotta merge them.
So the next one has to do with, um, the concept of, you know, superseding and temporal reasoning. So lives in Tokyo in February and then moved to Osaka. You know, where do I live? We wanna retrieve maybe the most recent one, or maybe we don't. Right? And then finally, this, this idea of, uh, dreaming, if you will.
So we call these, uh, second order and tertiary order memories, which are the patterns and the insights that we can take away, um, from the actual memories. And so one thing I didn't get a chance to mention is my background's in biochemistry and biophysics, and I work a lot with neuroscientists. And one of the things that we try to think about when building these memory systems is how is it modeled, um, similar to the human brain.
And so when we sleep, and I don't know who's into philosophy or, or psychology, but a lot of things run through our brains when we sleep that is, you know, subconscious, and the joining of holes and the joining of parts. And this is actually one of the things that Ben shared, which you have the data, you have the insights, you have the knowledge, and then you get the wis-the wisdom, right?
And so what is actually happening from that perspective in our brains that when we wake [00:44:00] up, all of a sudden, the perspective that we had shifted. Something that we thought we were gonna approach something the same way, something happened when we were sleeping, when we were dreaming that shifted. And so we like to think about these phenomenons from a neurological perspective and bring them and model them from a, you know, a developer's per-perspective, if you will.
And that's how we kind of are thinking about this concept of dreaming, where you have second order and tertiary order patterns and things that you've actually, um, arrived to. So another thing that we found, which was, um, you know, that, that people are, are, are building with here is whose memories are whose and when and how.
And so we call this scoping. And, um, it's, it's, it's pretty, it's pretty in-intuitive when you think about it. How do we scope the memories to a certain person or to a namespace that multiple people have access to or to the agent? Now, what wasn't so intuitive for us is that if you perfectly scope and deterministically scope all the memories, let's say, to the [00:45:00] user, then the agent is not gonna actually get any of the agent-level learnings that it needs that need to persist between people who have the agent.
And so there's a fine line in terms of how you have to think about, um, a memory system that the proper memories are gonna be scoped to the user, not just all the memories, but what are the memories that are gonna need to get scoped to the agent? What are the other patterns and things that you want to persist?
So, um, again, we have a, a set of ontologies, uh, which we call custom instructions, um, that make sure that facts about the person go to the user and facts about the work, if you will, go to the agent. And doing that delineation, um, was not, was not intuitive for us when we were, when we were building these out.
So another big question right now is, you know, this idea of how do I optimize for accuracy, latency, or cost? And there's a lot of memory providers out there that are publishing a bunch of benchmarks. You know, we, we've heard of Locomo, Beam, you know, whatever. And what we have seen is that a lot of people get misled in terms of how they're actually [00:46:00] analyzing and reading these benchmarks.
Whereas you'll publish an accuracy number that's maybe ninety-six percent, but we don't publish what the latency was or what the cost was. 'Cause if we're gonna use a heavier model, yeah, we can get higher accuracy, but it's at a cost of latency and our tokens. And so it's really important that when you're evaluating a memory system to use, or you want to evaluate your memory system, you wanna think about what are the constraints that you wanna put ahead of time before you actually start building the memory system, and then build it to those constraints, such as what is gonna be my target latency or my target tokens, and then can I get my accuracy to that stage with a, with a smaller model, for instance, right?
And so, um, this was a, this is a huge challenge right now that, uh, we're working on. And, um, if you look at the benchmarks, we're, we're, we're one of the, one of the better ones when it comes to balancing these three. And just wanted to, you know, uh, put it out there that these benchmarks have a lot of confounds in how they're set up, and different people are publishing different numbers.
Even some of the numbers that Memzero might publish are not one for [00:47:00] one in terms of the type of benchmark that was used, um, in rela-in relation to other memory layers. And so it's really important to understand how these experiments are being done, what are the confounds, so that you can truly understand the, the numbers that you're getting and whether or not that's gonna come into reality when you're running, you know, um, these in, in production.
So another thing that was, um, interesting and what we have seen in practice is that the difference between, let's say, ninety-five to ninety-five point five percent accuracy, one or two points, might-- really doesn't make that big of a difference in, in production, but it could equate to two times the cost, right?
And so there's kind of a, you know, an asymptote, if you will, in terms of what we see in reality of where accuracy is good enough, gets the job done. An extra couple of points doesn't really make a big difference, but you have that, uh, cost perspective that comes into play. So just keep that in mind. So really like the five key type of takeaways of this really high-level presentation I have for you guys is, um, [00:48:00] the, um, focus on your extracted facts.
Focus on if-- whenever-- when you're, when you think about building memory layer, really think about how are you storing things in your database. That is gonna be the main thing upstream that's gonna impact and domino down all the way. All right? So push everything up there. We actually run our extraction models, um, asynchronously.
Um, it's the delays like, you know, maybe like point two or point four, um, uh, seconds, if you will. Um, decide the s- uh, decide the scope when you write, uh, not when you query Uh, route each fact by what it's about, not necessarily who said it. So this has to do with our, um, with our scoping. Uh, plan for background cleanup jobs early.
So when we think about like cleaning, for instance, you know, put that ahead of, ahead of time before you start building your la-- a layer. Think-- So always start where you need to end and then work backwards from there. [00:49:00] And then finally, um, set a latency and token budget per turn before you start tuning your accuracy.
I've got some people who want to take a picture of that. Um, so that's what I got. Uh, pretty high level. Happy to take questions. Um, feel free to hit me up on LinkedIn, um, if you wanna talk to maybe some of our more technical folks at the company. Um, if you haven't started using, uh, Memzero yet, it's really easy.
Couple lines of code to get integrated and to get started. We have open source. We have a platform. We have a free tier. If you want a thousand-- up to a thousand bucks in free credits, just hit me up and message credits. So thank you so much.
We'll take two questions, yeah Yeah
Hey, uh, yeah, thanks for the interesting talk. Uh, so could you talk about, uh, you mentioned, uh, merging, uh, you periodically merge memories. Do you ever, ever do the-- find the [00:50:00] need to do the reverse, like splitting memories? Uh, that's number one. And the second question I have is, uh, can you talk about, like forgetting?
Uh, is it-- is there some overall strategy for get-- for forgetting?
Yeah, good questions. Um, regarding the first one, splitting memories, come find me afterwards, and I'll direct you to the right person who can answer that one. Regarding, um, forgetting memories, so we call this memory decay. And so one of the algorithms that we run is understanding how often is a memory actually resurfacing in a memory search, and what is that relevance in the search?
Based on that, the weight of that memory coming back in a search slowly starts to decrease. So when we think about humans, right? Like, you know, uh, what shirt I wore today or whatever. Some of these things are not that relevant, and we don't recall that memory that often. And what-- the process that happens in our brain is actually the, the, the, the, um, the dendrites and the, and the neural connection that we have there starts to weaken over time, and it becomes more and more [00:51:00] difficult for us to actually access that piece of memory in our brain.
And so we try to replicate that with, with this idea called decay, where the relevancy score starts to decrease, decrease, decrease, and decrease un-- slowly until it's very difficult for it to actually retrieve. And you can toggle that on or off. Yes. Um, can you talk about the, the graph implementations in Memzero and how they are evolving?
Um, we were seeing earlier that you guys were using graph and now maybe not as much. So do you have any insights into that? Can you repeat the question? Yeah. So the idea was is that we were using, uh, uh, graph memory, and over the course of time, we started to get away from graph. We still have, uh, a form called entity, so entity linking.
However, it's not truly graph. Um, find me after the, after the call, and I can direct you to the right person who can give you a more technical, technical response to that question. Thank you. One, one more? Yeah. Yeah, yeah. Uh, do you have any [00:52:00] solution for memory conflicts? For example, if you have multiple users and one user say X and a second user say Y, um, what, what the brain's gonna do with it?
Yeah. So the question was is, do we have a system for understanding conflicts within the memory system? If one user says X and another user says Y, what actually happens in that regard? It depends on what architecture we're running here. So if we have a namespace where you have multiple users who are scoped to that namespace, the question comes is, does that memory, uh, supersede another one?
Or if you have a conflict, will they get referenced and shown in terms of from a dashboard, "Hey, there's a conflicts-- conflict in this memory system. Um, this user said this, this user said this. Which one do you wanna use?" Or do we just tag in the metadata that there's a preference of this user for this piece and a preference for this user and this piece?
Does, does that make sense? Yeah. Okay. Figure it out. [00:53:00] Yeah. Thank you. Thank you very much.
I feel like I went back to the, uh, course learning how to learn. Uh, Barbara Oakley, uh, she talks about diffuse mode and, uh, focus mode. Actually, it's all about, uh... I, I would highly recommend you guys going there, uh, and visiting that. And, uh, thank you for classifying the memory. I think that was the best classification I have seen.
Uh, yes, our, uh, next speaker, uh, is, uh, she's a developer advocate at YugabyteDB. Uh, she has a deep experience in distributed databases and cloud backend systems. Uh, she'll be discussing how the data layer handles scalable memory systems. And before I hand it over to Heather, I would just want to give a shout-out to Yugabyte.
Uh, I have worked with them, uh, I think it was eight years back when I was working at a startup, and, uh, that was a really nice experience. Uh, yeah, good to see you guys. Uh, okay. [00:54:00] Welcome, uh, Heather Downing
Thank you
We're really pleased to be able to host you guys tonight, and we hope to do many of these in the future as well, even when the topic isn't memory. It's okay to come. Um, yes. So welcome. My name is Heather Downing, and I work for YugabyteDB. But you also might be noticing this other logo running around. It is pronounced Miko.
I know it sounds like it might be Mecko, but it is Miko. Um, way back in the day when we were deciding names, um, there was all sorts of different names that came up, and this particular one is a combination of memory and knowledge. That's what that stands for. Got it. Um, just in case you were wondering.
Okay, so you've already heard some pretty good [00:55:00] talks about this, and I wanna make sure that you understand where we came from. So we are a very, um, huge scale type of distributed Postgres company. So that means that when you come to us, you already know you have a problem with a lot of data, right? So we don't have to do any convincing that way.
We're very specialized, though, in handling huge amount of workloads all the time, right down to the database layer. Our founders also worked and did the same thing at Facebook, worked for Cassandra. So data is kind of in our blood here. So we thought AI could be great. We're gonna be ready for this. And we brought it in-house, and we had all sorts of also very interesting experiences trying to get that done.
And so as we were building this, um, there were some lessons that came out of this. Um, so I wanna make sure that I share a little bit of what we discovered when we went towards making something that we felt was really reliable and really [00:56:00] useful, not just to us, but also eventually our customers. And this went GA in May.
So this is pretty new. Um, but it's not actually a pitch this time. This is just talking to you about how did we measure what good was and how do we benchmark this, um, and what were the problems we solved. Okay? All right. What we discovered is memory is very... I guess that's also missing there. Missing that memory piece.
Um, it's very useful, but it is not sufficient in the practices that we ran up against. Um, specifically here, you'll notice this context rot, um, graph there. Do you all know what context rot is in this case? Okay. So like a couple of you. What happens here is, um, as I continue to talk to whatever AI it is, it's very sharp in the beginning.
Um, quantifiably, it looks like if you had like one million token window, I think that was for Sonnet five, or I believe, or maybe that was Opus. Um, [00:57:00] I can't remember that one. Is that after as I would work with it, it would be very sharp in the beginning, but then it would kind of lose track of things or it would give you, uh, regurgitated stuff that wasn't useful to you.
And so you really can only use like a hundred thousand tokens to maybe a hundred and fifty thousand before you started really feeling it or noticing it on very complex workloads. You may not feel that for other things if you were just doing a long-running, uh, conversation, but we noticed that right away.
Uh, the other thing that we discovered wasn't really sufficient was the way that companies are structured for all their data and how to share things. Right now, what was happening is we were just trying to get quick URLs to chats to each other, and that was all sorts of problems with security team, and they're like, "Well, how do we know you're sh-sharing with the right person?
How do they-- we know they even are allowed to look at that information?" It was just-- it was very loose. And turns out we were not the only one. And so how do you share, um, like a coding [00:58:00] standard that applies to everyone, but then you have a team override because you guys are special, and you like to use a different tool.
Um, and you lo-like to have things be maybe more clean in your code, whatever it is. And maybe you also have a personal preference on things you like to do. So we found that there were definitely layers to preferences that weren't not just memories. They were kinda like we would treat them as policies, but in reality, how I use this shifts week to week.
So that means it became very organic information at multiple levels, right? That's what we ran into The other thing that we ran into, um, is an inventory. So what is in an agent's context? How many different surfaces do they need in order to answer my question or to do what they need to do? We found that conversations were pretty important, and maybe we should just stick them in an ad hoc Postgres database.
Like, we need to hold them for later because maybe the conversation wasn't actually just chit-chatting. It was more about an architectural [00:59:00] review or a bug or something like that, and we wanted verbatim exactly what was happening. We wanted every single keystroke that was done. So that's maybe one place we put it.
Then there was, like, that policy knowledge or the really big, huge ones we'd stick in a knowledge base or a wiki. We found that they also needed to actually have real, like, true memory, like Mensira talked about, um, in order to have the position of what the agent was seeing as well as us. We also needed to know what they thought and what they were doing at a tool surface level.
That also seemed to be important for us to do. And then we needed, like, more, more dynamic data, maybe even more than that. Just keep going, right? Uh, we needed them to also be able to tell us what they were doing in other databases. So you see the sprawl is getting quite large. Um, when you say, "Remember when you did X, which store do you go to?"
It really kind of depends on how you set it up, right? Your [01:00:00] architecture. But if you have no architectural setup at all, what happens? It'll go to right now, um, one of my five different harnesses will go to its own project context first because I gave it no context. I didn't tell it where to go. So it goes, "Hmm, maybe it was a conversation we just had."
And then maybe it will look in the knowledge base. So, like, the order kind of became, especially with organic information, became very difficult. And so then we're like, we need to b-build skills to tell them exactly what to do. But we found that where I put things a lot is not where Hari puts things a lot.
I use conversations a ton because every time I have a meeting, I save that, burn all the transcript in there, and then I never have to call that MVP server again, and it's just there, right? So now my agent knows about that stuff, but maybe one of my colleagues uses the knowledge base a ton because there's so many brand documents that keep changing in there and all their assets that do.
So when we're talking to each other and I want [01:01:00] my agent to see what's happening over there, we couldn't use the same skill, right? Because where our preferences are are different. All right, time for some math. I've got notes here, don't worry. Okay. All right. So let's say we have a hundred engineers, okay?
Five projects each. That's like five hundred contexts, right? Two or three people in each one, let's say. And then maybe that's like fifteen hundred. And that's just one snapshot. These things churn all the time. Depends on the size of customer that we had or what we were doing. Then you add support and sales.
Now you're heading into millions. And we ran this math on ourselves and, uh, honestly, a million kind of felt low when we were starting looking at individual like projects contexts, personal contexts. They were a lot. Um, and really remember, each context is really like four or five little like data stores, right, in different places.
There's also kind of a blind [01:02:00] spot. You can see at the top, um, that- There's nothing that really in between explains the plan. There's nothing there that controls any of this. It was just, we need to store the data. Now, as a database company, we don't have a lot of opinions as to how you want to build your app.
That is up to you. But this time, we discovered that I think agent context and memory is actually a data layer problem, so we need s- to solve this problem. So, um, a few months later, Nico started as RAG plumbing. That's what it started with. And the teams would load their docs into Postgres and search them by similarity, and that was it.
And then we were building a Postgres extension, so you could run like a whole RAG pipeline with SQL primitives against it. And then the same users kept needing other pieces like memories and then conversations and decision traces. And somewhere along the way, this thing turned into a context [01:03:00] engine, and it's one place that stores all these surfaces and serves it up of a, a particular slice to the agent.
Instead of the agent churning all those tokens, trying to look for the right thing and me saying, "No. Yes. No. Yes," that it was deferred this time to this engine. Um, in the closed beta, we ran five different tracks. So we had knowledge processing and retrieval, memory extraction and retrieval, conversation search, uh, agent and human collaboration, and decision traces.
Tonight, we're just mostly talking about memory and retrieval. Um, but that's-- it required more pieces in order to really, like, start to s- to save on your token spends. But once we started doing that, I mean, we got benchmarks, um, internally, and now we have some customers that are showing input savings up to ninety-three percent if you don't have any memory system [01:04:00] at all, and definitely more if you're able to rank results and then give it to the agent, which is what Nico does All right, when-- okay.
So the biggest problem with when you're building this is what result do I wanna give you when we're building this system? Um, if I give you too little context, then the model kinda makes up some. If I give you too much, then that's definitely a large token bill because I'm filling up your window. And if you're really confidently off, then it, you know, hurts two of these.
And, uh, so you really need all three in order to make this happen. The context window is not the context overall of your project of what you're doing. That is just one area that, that a-that agent is working with. So we thought, okay, how can-- if it's rebuilt on every call, how can we keep it as small as possible with high quality as possible so we're not spinning up so many new contexts [01:05:00] every time for this agent?
Uh, so Miko stores the, the whole context and then just delivers a very small, high-quality slice into the window, depending on how you set it up. And it has everything a project knows from the conversations you had with a, um, completely different department and not just a coder. That's important to say, is that these have kind of felt like they're really only for tech people.
But what we discovered is that I have an entire department who is very interested in how to retain context when they're talking to their potential customers or their financial data. Like, how do we talk the same language about something when we think very differently? And we found that your LLM knows you better than the other person does and knows how to regurgitate that information to you in a way that you can understand.
So it doesn't always have to be verbatim information. It can be interpreted in the way you need to understand it, or you can get it verbatim. It's up to you So these are the data surfaces and [01:06:00] primitives that we kind of landed on. If you don't know anything about SQL and you're trying to use Miko, then what do these mean?
The surface memory, we did try to benchmark, um, each one of these as we went. Um, that's just extraction of our conversations plus graph. Uh, most recently, we're looking at maybe changing that out. Um, and then we did do a benchmark against Long-Term Evol, but that was only like one piece of what this context engine is.
So we, we need to benchmark our other parts, right? So to know how good it was, how accurate- Mm-hmm. How much latency. And so for knowledge, we, uh, we had to build an in-house benchmark called Spectra that is RAG-based. And that one is going to-- we're going to be publishing that as well, so you all can use it for you if you want.
Um, and that's about the chunking, parsing vectors, and the re-ranking of that information so the highest quality one gets to you first. Mm-hmm. And then we also did another in-house benchmark, uh, against Babylon for [01:07:00] just overall context, and that was like the artifacts inside of the conversations that your LLM might create on the fly that where are they stored exactly?
They, they're stored inside of Miko because they can come along with the conversation and then be stored in an artifact area that gets retrieved right along with the conversation again in a completely different harness or a completely different chat if you want. Right. Um, so um, we had three problems that we noticed that we didn't get right the first time.
Um, we found that you have to serve in-- you have to serve like one enormous like Postgres engine to, to solve this and millions of little tiny ones, uh, at the same time, we discovered. Because if you let agents model raw databases themselves, they'll do it differently every single time. So then we had to think about that scheme and how that was going to happen on the fly.
Uh, or should we keep them all online? Um, and then two, [01:08:00] we discovered that, uh, you know, we, we used Apache Edge as our-- as this beginning of our graph memory, but it had no multi-tenancy at the time. And so we built like a multi-tenant, multi-tenant edge, and we called it Mage because it's-- that's fun. And we also discovered that, um, promoting or promoting me, you are just an isolated piece of context or piece of memory, and now I wanna share it with you.
So you are now being promoted to a shared memory. Before you were isolated to just that one agent and that one user, but now you're shared. Turns out that if you-- how you slice this in your system, holy matters. It's very hard to do this. Like, how would you like to set it up? Right now, we obviously do a lot of data infrastructure, right?
So how do you share across a database if I have like isolated clusters and isolated data? How do I now share that with you? Do I copy it to you? Like, do you have a pointer back to me? How do we do this? Do we have a ledger? There was a lot to [01:09:00] this. And luckily, this is stuff we love to do. So we were happy to help solve some of these infrastructure problems I don't got this one.
Okay. Um, so as we started, turns out, um, recall for PDFs were pretty bad in the beginning. And that was because we just kind of started with the basics when we first looked at it, and it- it turns out that different file types just needed different retrievers. So we would write different ones just to see what happened, or sometimes we would use another open source project off the shelf that was really good, and we would try for that.
Um, this is a lot of work that we ended up doing. So the good news is, it's more fun to show you stuff though. Um, we wanted to see how accurate we could be- For a lesser model. Um, if we could reserve that area of that context window, right, and, and it not be filled up with a whole bunch of noise, [01:10:00] and I think Mendria talked about that as well, then can that help your cheaper models be better at understanding what you need?
And we did find that that was true. And it wasn't-- I mean, it depends on your model, right? But we started noticing that some open-weighted models were starting to get better about-- specifically about accuracy, specifically about carrying out a task without having to keep checking back or not being able to run enough.
It was really interesting to watch. Um, and so I- we actually do have benchmarks right here. But, um, you can see me after. So yeah, pretty good amount of accuracy. And savings kind of scale with that. So everything that we have here is kind of against, like, putting the full conversation in your context window, right?
That's how we're-- that's we're benchmarking it against, um, versus just having your agent ask Meiko for the [01:11:00] information or to put it there. Um, it was pretty interesting how much we saved. Um, I think the biggest for me, I've used Fable a lot, and the moment I started using it with Meiko, I didn't run out of tokens, and that was weird.
Uh, I mean, it was a lot of work that I was doing, and I was wondering why. Is it because it was getting better just over time, more precise, and because I was not the only one working. I had colleagues working on the same kind of problem as the same time, and we would share the findings in the same place so that they could go back and look at it and go, "Oh, someone else has al- al- also found this thing," and be able to take that into consideration.
So now there's like a shared building and learning experience that was happening kind of on the fly. It was pretty neat. Um, yes. Well, anyway. Um, maybe you can see this one. We started calling it, um, Meiko Agent Memory, and it is, but that's like [01:12:00] maybe a third of it. Um, we're big open source fans. We have been open source like the whole time.
We use open source products as part of Meiko, but we decided that it's really context is what it is. And so what makes context up is not just memory, it really is all of the pieces. And I think the coolest part about what I like Meiko-- that Meiko does is that now I can share things gated as a human. Or if you feel so inclined and you trust your agent gateway, you can enable all twenty-three tools through NCP, and you can have your agent automatically promote things if you want to do like a bounded Data pack for them?
They can. If you want. If you want. Because we, we don't want to get in the way of in- innovation or I could tell you it's due. But that is definitely what I noticed is that I, I wanted to know what my agents were doing and how they were thinking. I could just go into my colleague's data pack and he-- I could see it.
I did an entire presentation, um, [01:13:00] putting it together with my CEO, and we never talked. Like, we put everything in the same data pack, and I was able to find what he was talking about and the visuals he wanted and what I was doing, put it all together, and we presented on stage together without practicing it.
And I thought that was so weird. But also, how cool is that? And we want you to try it if you can, um, and just give us feedback. We want you to kind of put your paces. So it's free. It's a hosted SaaS. Um, but we also like deploy things on-prem. A lot of our customers are also in like the banking regulated industry spaces.
Um, we wanna see what you guys do with it and how you guys like to play with it. Um, but what we definitely ran into, though, is that at first, you have to be opinionated about your architecture from the get-go. What do you want this to be? It becomes very bespoke, but does it have to be, does it have to be through an app?
Does it have to be through a deployed agent, or does it have to be through the data layer? And more and more agents will change, [01:14:00] apps will change, but the data lives forever, right? So we decided that was gonna be something we worked on. So yeah. And then I also have, like, the next part, whenever this comes back up, of what we're working on to also additionally help agents work with data.
Whenever... You ready? Okay. Yeah? Well, I, I mean, can press send, but- You can try. I mean, it's our office. We should know. It's a new office. It is a new office. You're doing great. Thanks. I do have bars you can look at. Yeah, please do that.
Um, yeah. Well, also- Maybe you can- Most of us who have laptops, maybe we can share the- Yeah, we can share our deck after ... the Zoom log-- Zoom link or something. All of us can share our deck. We're also in Zoom at the same time right now. It's like the ultimate in remote [01:15:00] meetups. Dude, I know. Exactly. That's-- Does that work, AJ?
No. We'll try something. Lately, it does feel like that, right? Do you have a question while he's getting set up? No. No? Yes.
And we have one question here Oh. What is your question? Um, are you doing any local AI stuff? Are we doing local AI? You mean with the, the model that's using our data layer? No, no. Um- Or are we putting the data local itself? See if
we can submit it once again. Oh, okay. Um, well, any, like, open source community stuff on the local, your local computer We, yeah, we have-- So Yugabyte is an open source database. You can run it in Docker container. Um, MiQo is [01:16:00] infrastructure, so if you have, if you have a really good gaming rig, and you want to spin up a bunch of RAG workers and clusters and stuff, then we can sort of replicate that.
But we are going to make with a light version of this so that you can run it locally. We just wanted to make sure that this could stand up to scale, like millions and millions of this, and it does now, so we feel pretty good about that. Um, but we almost always start with open source. We almost always just, like, put it on our GitHub.
So if you follow that, that is definitely coming. We just wanted to make sure that this could actually handle kind of scale because we're gonna, we're gonna hit so many agents all the time needing access data. Yay. Sweet. I didn't know anything. That's a good question. All right. Go ahead, Hari. So, um, I just wanted to say that the barcode that Heather shared was MiQo do-- MiQoData.ai.
It was. It was okay. So if you guys wanna check it out. Just... Cool. So this is gonna be a continuation to MiQo. [01:17:00] Uh, so we started the other way around, right? Because we are a database company, and with MiQo, we have knowledge, memory, context. So where is the database? So first question is, do you need a database?
And, uh, so like I think we've seen even other talks, so when you build this end-to-end system, you're gonna have your RAG, vector stores, you're gonna have graph, you're gonna have MD files, you're gonna have structured data, right? You're gonna stitch all of these together. Like, like, like even for packages, you know, we just want to-- Uh, to give you an example, I was-- I had a bad memory, so I didn't want to use GitHub issues.
I didn't want to use Jira. I want to do ten things. I was just putting them down in MD file, right? What is the next ten things that I need to do? So my task list was literally one big MD file. And after an MD file, when I got to, like, item number five hundred, I need to [01:18:00] put it somewhere. It's like, like we are a database, uh, company, right?
I want to put it in a table, right? So I want to create a table to store my task list. And then after that, I wanted to store more stuff. I want to actually store my, my application details. So we always come back and say, there is memory context and all of that, but you also have structured data that we want to store some of that, right?
So with MiQo, like I said, yeah, we have the data packs, we
have traces and all of it. There is no structured data. So now we are also introducing a s-- database within the MiQo data pack. So in addition to storing your memory and context, you also have Postgres. Right. Uh, this is like, uh, uh, it starts with zero. So you, like, uh, the ideal pays zero cost. So pay as you go, how many cores you use is how much you end up paying for.
It's isolated. Every data pack is a unique isolated PG, uh, Postgres cluster. So you're-- if you have two different [01:19:00] teammates using two different data packs, the data doesn't shadow work, right? So all the security stuff is in there. It's got branching, um, which is a hot thing with databases right now. And it, it works wonderfully well with agents, right?
Because like I, I want to experiment with stuff. I want to change things, but I don't want to touch my production database. I don't want to go mess with it. So if I just do an instant, uh, zero clone, uh, zero clone, sorry, zero copy clone, copy and write branch, I can do whatever I want or my agent can do whatever it wants to on the branch.
And essentially, if, if you are a database people, the example I used is I want to create an index You don't have databases or indexes. Creating an index isn't easy. It's hard. It can break things. It can break Prompt. So I ask my agent to branch, create my indexes there. If it works well, then yes, then I will use it in Prompt.
Otherwise, I'm not gonna touch Prompt, right? So branching turned out to be really useful, so we have branching as well. And of course, it's all in CP two dot [01:20:00] zero. Everything goes through MCP, otherwise Prompt doesn't like it. So any-- or any agent doesn't like it, so we have MCP in there as well. And, uh, like I said, we are primarily a distributed database company.
So how do we tie it all in together, right? So with, uh, you start at zero, so it's fraction cores, PS code. But then you start with one, and then you build one up, ten apps, hundred apps, and each app has a tiny DB. Each one isn't big, but now you have a hundred of those. So you scale up, and you have need a multi-tenant fleet.
Problem isn't the, the DBs themselves. Now you have a hundred DBs. So it's a hundred things that you have to upgrade, patch, uh, deploy, make sure they're, uh, up, they don't go down, right? And all of that stuff, right? So you need to handle fleet management. So let's also handle, uh, multi-tenant fleet management using agents again.
So we have a few in-house things for, uh, [01:21:00] migration, uh, performa-performance and stuff, right? And even fleet management. And eventually, when you... Not everything will get big, but if you do happen to go really, really big, distributed Postgres in a government does come in. Uh, and you can go from the half a core to ten core to three thousand cores, right?
So you can scale up globally, go distributed, grow multi-continent, continent if you hit that scale, right? And yeah, since we use the same underlying engine, it's all Postgres. Uh, we offer zero migration, no rewrite database scale up from zero to very huge numbers. And yeah, that was a very short intro to AMP.
Uh, we'll keep you posted on Discord when it's available for use. Um, that's the, that's the, uh, it's that Discord, uh, link here as well. So yeah. Yeah. So... Did we miss it? Yes. Any questions?[01:22:00]
Sure. Yeah. So my question is, have you had any problems with like scaling it on the infra side? Like, have you tried to scale Mico on like, uh, thousands of databases, for example? And, uh, do you scale it in the same way as you scale typical relational database like sharding across application, or did you use some special approaches?
Okay. So, so we have two, uh, things, right? When sharded and non-sharded. So you start without sharding, that's just regular Postgres. You can keep scaling up. The biggest Aurora box you can get is like one hundred and twenty-eight cores or something. Mm-hmm. So you can get up to one hundred and twenty-eight cores.
Once you get past that, you have two options. Either you have to do sharding in the application and Postgres, and with MySQL, there's Vitess, there's Postgres, a few things, but it needs app rewrites, right? With [01:23:00] Yugabyte, that's our type pitch, is that it is Postgres that you can use without rewriting your app.
So the sharding is done within the database. Mm-hmm. And th-that's not-- That is just, um, like we have built over ten years. Okay. And that's, uh, better and better. Maybe I should repeat the question. Uh, the question was, I guess, how to scale from, uh, small ones to a large sharded system. Okay. So basically you don't want to rewrite your app.
The DB can scale for it without you having to touch your app, because app rewrites are the hardest thing in the world. Okay. Okay Any other questions? Yes. So my understanding is Miko is an MCP that plugs into an LLM. So there's this idea that the more, more MCPs you add, the more context gets loaded. Mm-hmm.
Have you seen any patterns with that? Like, has Miko actually added to the problem, or how has it actually reduced the actual context loaded-ness of the, of the [01:24:00] window it's off? Yeah. So I can repeat the question. Uh, so Miko is an MCP, and now we're getting an MCP per. So adding more MCPs is making the problem worse and adding to more context and more confusion.
Our way to work around it is just use skills. Um, you just-- If you-- As long as you use the skill to say, uh, this is what you use Miko
for, this is what you use your Slack plugin for, this is what you use your GitHub MCP for, it seems to get it right. And without a skill, w- there's ten things that does the same thing, it gets confused. But once you put the skill in, so what we are doing for coding, every one of our, uh, Git, uh, Agentic MP file has at the top saying, "This is the Miko data pack ready," and all the context is in Miko, right?
So as soon as it-- The first thing it reads is the Agent MP, and it's all there in Agent MP. So that gets stored, and anything in Agent MP is sort of preserved throughout the session. So [01:25:00] we feel that doesn't lose it because of that. Yeah, yeah. You want to add something to that, yeah? Yeah. Well, it depends on h- which version of MCP, right?
'Cause sometimes now it can just be, um, it not loaded in your window by default. Uh, so that's part of why it depends on how you set it up. It's very flexible. So like I can use it in like a, like Cloud Mobile or desktop with no skill. It works. You just tell it, "Hey, I need this." Um, it doesn't have to load the tools on the fly.
You can just load one of the tools, three of them. It depends on how you want. It's very flexible. If you use our installer, it'll use the hooks framework to enforce every five turns to push conversations you're having and what you're doing up for you if you don't want to think about it. But if you're a control freak like me, I'm like, "No, I don't want to burn anything unless I really want to burn something, and I really want to look something up."
So I would like a separate little tiny one and say, "This is, this is it." And I turned all the memory off by default in all of my harnesses. I only use this, and so it looks for it first. But it doesn't [01:26:00] have to know all of the tool calls. It just says like, "Just do a search for me." And then we do hierarchical search inside of Miko, and then we return that.
So it doesn't have to have all of the tools all at once active. So we haven't discovered that it-- the results that it returns blows the window. It's usually quite precise. Is that he-helpful? Yeah. That's helpful. Thank you. Okay. Oh, so that sounds like you ma- you force it to make a call. Uh, you force Claude to make a call to Miko every time- If it needs it Okay.
So it decides or you-- you're kind of telling it- It's up to you. So you can use- Can you repeat the question please? Yeah, sorry. The question was like, it sounds like you're forcing your LLM to use this. I-- It depends on what you do. I've also written this with like strands agents. So if you want to deterministically decide when you want to load things into context, you want to control the agent life cycle, you can like run your own local agents with it.
Um, it's kind of up to you. You can decide to encourage it with a skill. You c-- [01:27:00] Or you can kind of deterministically go in early as a coder in the early part, in the deterministic loop outside of the inference, and you can also choose to do it then as well. It all kind of depends on how you set up your data pack.
Like I do like an end of the day one, but I know that the VP of product, he wants the every five turns because he doesn't want to think about closing out his day. He wants everything in there all the time. It's very flexible for what you want. But if it's the end of the day- Mm-hmm ... then, uh, that means everything you're working on at that moment has-- the whole day has to be in your context window that Claude is managing, not Nico, right?
Uh, yeah. At that time, it can-- it's kind of up to you as to when you want to push things out and when you want to pull them in. It's completely up to your, to your choice. Okay. Yes. Because if we had said this is the only way, it would be wrong in three weeks because somebody else would have a different architectural style.
And the purpose h-here is persistence, right? The idea is memory is part of contextual persistence, so it's kind of up to you as to when it goes in there. [01:28:00]

