James Ward: [00:00:00] The Python code that gets written for code mode is honestly, like, hard to review. Like, like, you know- Mm ... anytime that the agent is writing code mode stuff, y- I look at it and I'm like, "I, it doesn't... Like, that is gobbly-goop code I don't understand." Yeah. Looks good to me. I hope, I hope it works. And then, like, the next day I get an email that was like, like, "Hey, um, my agent is telling me that the markdown docs on AWS are, are trying to do an injection of- Oh, no
James Ward: something potentially malicious." I'm like, "Roll back, roll back." Oh. We didn't... Our matrix wasn't big enough.
Demetrios: James, you're gonna be giving a talk at Agent Con North America on October 22nd and 23rd in San Jose, California. I would love to know what you're talking about.
James Ward: Awesome. Yeah, super excited to be there. Yeah, so my talk is gonna be about some of the things that kind of come later in [00:01:00] your MCP creation journey, because, you know, it's easy to build an MCP server and throw it out there.
James Ward: But then, you know, as it's out there, uh, as I've built some MCP servers, you know, discovered that there's some things that you're gonna want, like, down the road, you know? Uh-huh. And so this is really exploring some of, some of those sorts of things, and so I can go into these, some details on what, what those look like.
James Ward: But yeah, that's the general idea is, like, like, you know, not day one with MCP, building MCP servers, but, like, day 10, you know? Yeah. These are some things that you're gonna have to think about down the road.
Demetrios: So this is, like, the MCP Builder 201 course.
James Ward: Maybe even, like, 301 or 401. It's- Oh, nice ... yeah, it's... Yeah.
James Ward: It's, you know, we're gonna get into some more advanced patterns and maybe not quite 401. We'll just call it 301. Oh. Yeah, yeah. And for MCP- Don't undersell it ... like, server creators. Yeah. Yeah, yeah. This is- It's for those- You know, if you're not building MCP servers, like, not the talk for you.
Demetrios: Yeah. It's very [00:02:00] deep into the MCP server creation.
Demetrios: Okay. Well, talk to me about some of the anti-patterns of building MCP servers or the-
James Ward: Yeah. Well, really, we're gonna focus on the patterns for these, so, you know. Okay. I'm sure there's many talks on the anti-patterns, but we're gonna talk about the patterns for, for, you know, what, what your MCP servers are likely gonna need to do down the road.
James Ward: So one example of that is that oftentimes the agent, when it makes calls to an MCP server, you as the MCP server have no idea, like, okay, this call fired off to my MCP server, this call fired off to my MCP server, this fire... Like, you have no correlation across those different calls, and so there's no, there's no way for your server to know, like, okay, this is, like, one agent session, you know, based on one prompt- Hmm
James Ward: from the user. Um, and then how do you actually then correlate those together? And this can be really helpful for understanding, like, how is my MCP server actually getting used? Is, is it, is it producing the [00:03:00] best results that it can for the agent? And this is the kind of observability stuff that when you really wanna tune your tool descriptions, your, your, um, data, uh, schemas, um, your input parameters, all those sorts of things, when you really wanna get that thing dialed, you're gonna need this kind of information to be able to see the journey that a agent is making, you know, with your, with your MCP server.
James Ward: So there's- Yeah, the- ... a way that you can actually do this and, uh, and we'll talk about, you know, all the details in the session about that, but
Demetrios: There's almost this performance aspect of it where you've got the system performance, but then you've got the actual, like, is it responding and making it very agent-friendly in a way too, right?
Demetrios: That's right. So there's that- Exactly ... performance side of things.
James Ward: Yeah, I think that a lot of it does come down to possible areas of optimization, and that's what you're really trying to discover. Because generally the agents, what they do is they decide they want to call [00:04:00] an MCP server, so they make the call, get the response, send it back to the LLM, and they keep doing this, right?
James Ward: You know, you've seen the agentic loop just going in circles, you know, trying all these different tools, exploring around. And, and oftentimes this can be really inefficient because all these round trips to the LLM, you know, take a long time and are expensive. And so maybe what you're trying to discover through this observability is, do I need like a aggregate tool that's gonna take potentially 10 tool calls and roll that all into a single tool call that the agent can make instead?
James Ward: Mm-hmm. And so you're trying to, through observability, discover those places for optimization where you can make the agent more efficient. And you really only know that when you can create correlation across requests, so then you can look at the, you know, all the things that happen in a given session.
Demetrios: Well, it is funny how you're mentioning you've got your agents that are hitting the MCP server various times, and unless you really take it a step further and try to figure out, [00:05:00] is this the same session or is this a different session, then you're not really going to realize that this agent was trying to do X.
Demetrios: I made it 10 times more difficult because I had it making 10 different calls.
James Ward: Yeah.
Demetrios: And if I can abstract that up a few levels, recognize its intent, and just create the server so that that intent is only one call, that's going to be a much nicer experience.
James Ward: Yep, absolutely. Yeah, and then likely in a real system, what you're gonna do is build out evals.
James Ward: And so you're gonna take the observability data that you've uncovered, and you're gonna ultimately get to a set of evals that then will give you actual data around, okay, use-- the agent was trying to do this without this, like, aggregate tool or whatever the solution you come up with is. Here's how many turns it took with the LLM, here's how many tokens it took.
James Ward: Mm-hmm. And then, and then you set up another arm in your eval [00:06:00] that is, okay, now do it with this, you know, new tool that's this aggregate thing. And then you see, okay, did the agent choose that path? Did I have the right tool description so it decided to go the shortcut instead of the long way? And then did I use less tokens, less turns, those sorts of things.
James Ward: And so they could definitely influence the eval system that you, that you would hopefully have in place around your MCP servers as well.
Demetrios: One thing that we've been seeing a bunch of at all of these different MCP Cons, MCP Dev Summits is the advent of an MCP gateway. Have you thought about adding that into your talk?
Demetrios: Is that one of the patterns that you're going to discuss?
James Ward: I-- We won't talk about gateways. I'm sure there'll be many talks about gateways. You definitely can use a gateway to, to start to, uh, potentially do these optimizations automatically. And one of the things which maybe we'll talk a little bit about in, in the talk is this idea that [00:07:00] if you know, if you know input and output schemas to your MCP servers, which input schemas that you have to know those, but output schemas are optional.
James Ward: When you know- Right ... actually both and you put something between the agent, like a gateway, but something between the agent and your MCP server, then you can actually automatically discover the right optimizations that you can do. And so this is- Hmm ... you know, you take a-- Instead of doing it manually with correlations and evals and all that kind of stuff, you can actually have a self, uh, self-optimizing gateway in the middle.
James Ward: But I think to do that well, you do need to turn on output schemas because you really need to understand the shape of the things and how they fit together and like, okay, this shape is gonna be able to go, uh, into this form. I know the output schema of that, and then I know that the next sequence of things then takes parameters from that thing and feeds into this thing.
James Ward: And so, um, so maybe not all use cases require those [00:08:00] output schemas, but I think when you have them, you could have a more automatically, um, updating, optimizing gateway in the middle.
Demetrios: Yeah, it's funny you mention that because Sean Smith at the last Agent Con in EU, uh, in Amsterdam last week, he was talking about that, like the self-optimizing MCP servers where they just recognize, "Oh, well, we could be better.
Demetrios: We could be better here," and it is a little bit more-- The whole-- His whole take was it's much more nuanced than you would think, but it is possible, and I think that kind of is echoing what you're saying right now.
James Ward: Yeah, and I think we're, we're, we're just beginning to uncover s- uncover some of the approaches that we could use for this.
James Ward: Um, the-- We took kind of a tangent generally in the industry where we said, "Okay, MCP, the way that it gets used in this like, like very heavyweight, you know, back and forth process with the LLM, uh, it's really inefficient." And so this was like this [00:09:00] backlash against MCP, like, like, "MCP wastes all these tokens."
James Ward: And- Mm-hmm ... so many of the, the use cases that I saw where people were making this point, like the MCP is inefficient, what they were doing was they would, they would-- Let's say they had a tool that was, uh, get customer by customer ID or something like that, and then they would go into their agent and say, "Okay, give me all the customers that, uh, that I've emailed in the last 60 days or whatever."
James Ward: And so then they show that the agent's like, you know, does an initial query, and then it goes through and calls get customer, you know, 100 times. And of course that's gonna use a lot of tokens and be slow. Yeah. And the solution to this is very obvious. Just have a tool that is get customers, you know, and takes a parameter that is all the IDs that it wants.
James Ward: And it's like, it's not MCP's fault that you like have inefficient tools that are like making the agent do all these round trips. Like that can be solved, and we don't have to go to code mode and these really, you know, complex, complex solutions. It's like, just look [00:10:00] at the, how your tools are being used and discover the ways that you could, uh, you know, create more efficient tool usage.
James Ward: It's not MCP's fault- That's right, Billy ... you like haven't done that optimization.
Demetrios: There's so many pieces, especially in the beginning, where we were exploring and we're trying to figure out the best ways to do it, and you recognize that, man, we just, like, built a lot of pretty crappy MCP servers. Yeah. And MCP was the scapegoat.
Demetrios: Right. It wasn't the people that built the MCP servers that were the- Yeah ... scapegoats.
James Ward: Exactly. Yeah, and I mean, I've built those terrible MCP servers, and it's like, we just didn't knew- know better, you know, initially. It's like, cool, like I, I- Yeah ... was able to connect my agent to this thing. Awesome. Okay, that's step one.
James Ward: Yeah. Right? And then, then we gotta get to, you know, step seven, where we add the, the efficiencies and make things more optimized and, you know, I, I think it's- Yeah, how many servers- It's okay that we go on that journey of like, what, like what do they say in coding? [00:11:00] Like, like, I don't know, something about how you, you, you, first you make it work, then you make it fast, or something like that.
Demetrios: Yeah. Yeah, and then you make it cheap.
James Ward: Yeah, that's right.
Demetrios: Hopefully.
James Ward: Yeah, hopefully.
Demetrios: Some of us never get to that
James Ward: third step. Fewer tokens. Yeah. Yeah. That's right. Yeah. So I think there's- But how many- ... there's lots of areas that we can, for improvement, and, and yeah, I hope that my talk will help people see some of the patterns that will help them get there.
Demetrios: How about with the new MCP release in July, that 7.28, AKA MCP 2.0 Having to now support two different MCP servers, and I've heard talk from folks where they said, "Yeah, the migration was seamless. It was easy. We bumped everything up, and we're good." And then I've heard other talks where it's like, that was us.
Demetrios: We can enforce all the migrations we want, but what happens when I have a gateway and now I'm pinging to external MCP [00:12:00] servers that are not following the new update? So now I have to simultaneously support both of them. Is there any patterns that you've been able to decipher from that?
James Ward: It's, it's definitely a challenging area.
James Ward: Um, it-- we won't go into it in my talk, but I, I do have firsthand experience with exactly this, and it's, it's hard. Um, it's h- it's mostly hard on the library, like gateway and library author side of things. So I maintain a, an implementation of MCP, and so I had to deal with this, and I use that implementation in a variety of MCP servers.
James Ward: And so I had to, to be able to support the dual era MCP, do the like negotiation of the protocol, both on the server side and the client side. And, you know, thank God I like had AI helping me do this, and I had, um, now the MCP, uh, compliance test suite is really helpful for, for getting the agent to, you know, know what success looks like and be able [00:13:00] to iterate and get there.
James Ward: And so, but, but once the library and the, the like gateways and those sort, sort of things solve the, the negotiation and the dual era stack and all that kind of stuff, like once it's solved in the libraries that you're using, then as an MCP server or client, like you, like it's totally transparent to you, which is nice.
James Ward: Like- Nice ... like you as the MCP server author You don't, you don't have to think about the actual underlying protocol. And, and as long as your, the library that you're built on is doing the right thing, then, then you should be good. Um, so so yeah, I think it is, is nice that most, most MCP server authors really shouldn't have to deal with the pain of, of the underlying protocol, which is the way that it should be.
James Ward: And I'm, you know, thankfully the MCP, uh, contributors have been able to, to put the burden where it should be, which is on the library author side and not paying the, the actual server creators with, with this stuff, so.
Demetrios: What are some other [00:14:00] patterns that you wanna highlight?
James Ward: I don't, I don't know if we'll go too deep in the talk, but one of the things that, that I've been exploring, um, in my day job, so I, I work at AWS and working on agent experience, so how do we make our, the experience of AWS agent-friendly?
James Ward: And- Right ... one of the things that, that we're looking at is, um, right now we have like a code mode MCP server that enables, essentially just lets the agent throw some Python at it that's like sequence these AWS API calls together and do all these things. And it, it works, like the agent can write the code and, and do all that.
James Ward: Um, but I think there's a better option, and it relates to what we'll talk about in the talk of like, like creating MCP servers that, that can, um, provide better optimization. Uh, but then also, uh, one of the downsides of the code mode [00:15:00] is the user, um, like verification of what's gonna happen. Like when you're, when you're having your agent talking to AWS, like you wanna be careful, you know?
James Ward: You wanna like probably- Yeah ... review like, like what the agent is, is gonna go do on, on your infrastructure. And so, uh, so the Python code that gets written for code mode is honestly like hard to review. Like, like, you know- Mm-hmm ... anytime that the agent is writing code mode stuff, y- I look at it and I'm like, I, it doesn't-- Like that is gobblygoo code I don't understand.
James Ward: Yeah. Looks good to me. I hope, I hope it works. It looks good to me. Send it. Yeah. Send it. And, uh, and so, so what I'm, what I'm exploring is can we essentially have a low-code mode that is not full Python code, but is a very constrained DSL for w- for the sequence of operations that the agent wants to perform?
James Ward: And then can we make it so that the human can look at it and understand what's gonna happen? [00:16:00] And can we make it so that the schema for that, the grammar and schema for that is validatable so that the agent can then do a step where it says, "Okay, here's the code mode that I, you know, here's my plan. Now let me validate that against the actual schema that, that ava- is available so that it doesn't actually try to do things that aren't gonna work."
James Ward: 'Cause what often happens in code mode is it gets, you know, into step two of 10 and, and fails because it tried to access something that, some property of something that wasn't actually there. And so, you know, this happens all the time. And so- Sure ... can it actually do a validation step against the schema and get like, like very clear feedback on like, okay, you know, you wrote this low-code, you know, DSL that you tried to access this property, that property doesn't actually exist.
James Ward: Here's what you should do instead, and then the agent can try again. And then by the time that the agent has ver- verified the code, the human has reviewed it, boom, we send off this like sequence of operations. And so, so that's part of what I'm exploring is how do we, how do we enable this, [00:17:00] this, uh, kind of better code mode.
James Ward: And, um, it relates definitely to the, the, the talk that I'm giving on like As an MCP server, how could you potentially expose a low-code mode that doesn't require sandboxing, you know, allows you to interpret the, the actions and, and then as the, like, potentially the fallback for, for when the agent's like, "Okay, I, I see the, all the things that I can call, but let me just write code."
James Ward: If you could- Yeah ... offer this, like, low-code approach instead of a full code approach, um, you know, hopefully that would be a better agent experience. But yeah, so that's part of what- Dude, that's so creative ... I'm exploring in my day job.
Demetrios: You know what it reminds me of is back when LLMs first came out, there was a guy that came into the community and he was saying, "Look, I was able to get LLMs," and this was before, like, the coding use case became very [00:18:00] popular.
Demetrios: And what he said was, "I was able to get LLMs to create their own code, but it, it's not any code that we as humans know." It's basically just saying, "If you were to create this, you know, like, here's what I want to be done. How would you, in a almost like pseudocode way-
James Ward: Yeah ...
Demetrios: explain it?"
James Ward: Yep.
Demetrios: Yep. And so the LLMs were able to talk to each other in a way that was like this pseudocode that they understood, and when you read it, you kind of were like, "Ah, I guess that kind of works."
Demetrios: Absolutely worthless for, like, compiling or any of that, you know? Yeah. But-
James Ward: Yeah ...
Demetrios: for the agents, well, they weren't really called agents back then. It was just straight LLM calls. For the LLM calls, they were like, "Yeah, I understand. I get it."
James Ward: Yeah, and that's, that's exactly what I've been, been looking at, is not the agent-to-agent case, but, but if you give, if you give the LLM a language grammar, and you give it a schema, and you give it a goal, it [00:19:00] actually is really good at coming up with the language that you've told it to fit things into.
James Ward: And- Yeah ... uh, you know, not, maybe not surprisingly that it's a language model, so, you know, if you tell it, "Here's my grammar, here's my schema, give me, do this," then it can, it, it's actually pretty darn good at that task, so yeah.
Demetrios: Have you seen the website 2027.dev?
James Ward: No. No, what's that?
Demetrios: Okay. So these folks, uh, I love what they're doing.
Demetrios: They're basically evaling how agents work with the tools that are out there- Nice ... and how well the agent experience is with- Oh, that's cool ... all these developer tools. So they'll kind of like- Nice ... see how fast is the agent able to do what you normally do with these tools. So if it's like Railway- That's awesome
Demetrios: it's Spinup and- Yeah ... you know, host a website, or if it is- Yeah ... I think they were, they were using like Snowflake. It's like, how [00:20:00] does the agent work when it's trying to write different queries? And- Oh, that's cool ... they give, uh, they basically have a whole thing. I think it's like benchmarks.2027.dev- Oh, definitely check that out
Demetrios: or around there. It's- Yeah ... relevant
James Ward: to my day job for sure, but-
Demetrios: Exactly.
James Ward: That's why- 'Cause like you gotta have the data. It's like you just don't know- Yeah ... until you have the, the data. It's like, does any of this make anything better? And it's like- Yeah ... I don't know. Like, like you gotta have the evals, you gotta have the data, and then you can actually know.
James Ward: Like, and, and then you-- There's very easy... See, here's a really cool thing about where we're at with this, is that we actually can get data like never before. Because when a developer is just like on their CLI or in their IDE doing things, whatever, there was no way as a, like vendor to measure, you know, their productivity essentially.
Demetrios: Yeah.
James Ward: Now, very easy to measure token usage and how long something takes, right? And so now with evals, we can get really good data on user task, what tools [00:21:00] did it use? How long did it take? How many tokens did it take? Where does it, was it, was it successful? And we can now have, get data around all that, which is absolutely amazing.
James Ward: So true. So I'm hopeful that- So true ... all this results in so much better end user developer experience. Like agent experience is just, you know, saying that, okay, we wanna make users and their agents, uh, more productive. And so now we can measure that and see are we actually doing it, which is exciting.
Demetrios: Well, the fun part that I've been thinking about a lot is how you can optimize the act of improving the agent experience.
Demetrios: So I, I always go down these rabbit holes of like, you know, if you're looking at the benchmarks and you're looking at the companies or the products that are very low in the agent experience score, why is that? And oftentimes it's not as clear as like, oh, just they don't have llm.text on their website, because right now mostly everybody has that.
Demetrios: It's [00:22:00] things that you wouldn't necessarily expect, and you also have to almost like work with the agent to see why is this hard for you? Where are you getting stuck right now? No,
James Ward: totally.
Demetrios: And-
James Ward: Yeah, so that's a task that I, I do a lot is I take a transcript from an eval and I give it to the LLM and I say like, "Root cause analysis on where, where it went wrong."
James Ward: Like, like why- Yeah ... why did it choose that tool? Why did it give that tool that parameters? Like, like, you know, and, and it can do it, like which is amazing that you can take a transcript from an eval and actually get really good root cause analysis on it
Demetrios: Are you-- 'cause a lot of times where I get frustrated is some of the thinking steps get obfuscated from the user.
Demetrios: How are you doing it? Is it with open source models, or is it through the labs? Like, how can you see some of these thinking steps?
James Ward: Yeah. In my case, I, I use different agents and different models, and that's part of the, the matrix of things. And so some of [00:23:00] the- Uh-huh ... some of the models and agents that I'm using do expose the thinking steps, and so probably in those cases, I'm able to, you know, get a better, uh, root cause analysis of, of what happened.
James Ward: But that's not universal, and so it could be that- Yeah ... you know, this agent and this model that is not exposing thinking, like I'm not able to solve its problem on why it went wrong. But, but then you kind of hope that, okay, this, if I solve it for this one over here, then hopefully I solve it for this one.
James Ward: But that may not- Close enough ... be true. And so again- Yeah ... like, like with the evals, we can get the data. It doesn't necessarily mean that we're, we're gonna be able to get to the root cause and solve the, the issues. Like I was just, uh, I've been playing with Jev the last, um- Mm ... few days, and was trying to- You and everybody
James Ward: root-- Uh, yeah, me and everybody. It's so fun. So probably talking about Jev in the, the agent con talk, um, 'cause you know, we know it, it's required. Um, but- Yeah ... but I was trying to like g- like root cause a, uh, something with, with Jev, and it took writing hundreds [00:24:00] of evals to, to try to get to root cause, um- Wow
James Ward: of what was going on, which luckily with Jev, it's, you know, they're inexpensive- Fast ... fast and cheap and all that. Fast and cheap. Yeah. But, but, but yeah, it was like, like, uh, this was like, like hundreds of actual, um, s- like, like ARMs in my eval that I ran many times. So it was, you know, I was burning through- Wow
James Ward: you know, like, like four cents of Jev, um, usage to do all that stuff. It's so cheap. Yeah.
Demetrios: Exactly. Four cents. That's a-- Yeah, exactly. That's hundreds. That's a whole lot of Jev.
James Ward: Right. Whole lot of Jev.
Demetrios: Oh. Yeah. Yeah. And you know, like the other piece that I wonder about with Your job and the agent experience is how much are you evaling the harnesses and how much are you changing harnesses and what you're doing on that level too, because on one level it's the models, but then on the other level...
Demetrios: And what folks would call like the inner harness, the outer harness, all of that fun stuff.
James Ward: I do [00:25:00] almost all of my testing in a harness, and the reason is, is that in my case, the, the experience that I'm trying to make better is the user that's in a, in a code assistant, in a harness. And so I need the system prompt.
James Ward: I need the loop that the harness has and- Mm-hmm ... and so that, uh, almo- I, I would say almost universally my eval testing is in a harness. Um, and then I have to have many arms for, you know, different models and different thinking levels and, you know, it's just like- Yeah ... the matrix gets massive real quick.
James Ward: But I, um- That's what I'm thinking ... fun experience, I, I, um, I learned a lesson on this just recently. So we, we, um, on AWS documentation, we have the markdown versions of the documentation, and we, we have the, uh, agent toolkit for AWS, which is, you know, the, the agent tools and CLI setups and MCP and all that kind of stuff for, for AWS.
James Ward: And we wanted to [00:26:00] put a line in the markdown documentation that said like, like, "Hey, you know, the, we have, we've got this agent toolkit for AWS and, uh, you can set it up and it'll make your life easier." And so I, I did some evals on that, like, "Okay, let's, let's add this to the documentation." And this was not on the live site, but, um, just on like a testing site.
James Ward: Added this, this like line in, and immediately the agent was like, was like, "This is, this looks like an injection attack." And so- Oh, no ... like, like classified, classified me just like adding this line into the document. So I'm like, "Okay, I'm glad, like glad I didn't add that to production." But then I wrote all these evals to be like, is there anything that we can put into the markdown to just like let the user- Yeah
James Ward: know like, like, "Hey, you know, you, it'll, it, you'll have a better experience if you use this agent toolkit for AWS." And so had, you know, hundreds of different variations of different, different markdown snippets to include. Ran these through, um, [00:27:00] a s- a, a small matrix, and this is the important one. I've ran it through a small matrix of, of, uh, harnesses and, and models, and we thought we got to something that, that looked good and was like, you know, the evals are like, "Yep, the agents aren't flagging it as malicious or, um, distracting or an advertisement."
James Ward: And, you know, sometimes the agent surfaces it to the user and says, "Hey, you should use the agent toolkit." And so thought we had something good, so we rolled it out to production, this little, uh, snippet, and it was, you know, it was very carefully, uh, created to, to not trigger all the things. So we rolled it out and then like the next day I get an email that was like, like, "Hey, um, my agent is telling me that the markdown docs on AWS are, are trying to do an injection of-" Oh, no.
James Ward: something potentially malicious." I'm like, "Roll back, roll back." Oh. We didn't, our matrix wasn't big enough. And so- Yeah ... so we rolled that back, and now we're working on, um, hopefully something that the, the [00:28:00] agents will be more happy with. So, so yeah, you have to be careful with these sorts of things. That is exactly what I mean.
Demetrios: Uh, 100%, because it's like, wow, that you, you're trying to help, you're trying to make it much more valuable, but then if you're not thorough in your evals and that matrix- Yeah ... and you're not trying to get as much data as you can before you put it out, you can have these potentially difficult situations.
Demetrios: Yeah. And that's just for, again, like I'm looking at agent experience on so many different levels, and as you are, I'm sure, with the models, the harnesses, and then you're saying, "Well, we can also optimize these MCP servers."
James Ward: Yeah, yeah, yeah. Exactly, yeah. There's just so much now, and so it's, uh, but yeah, it's, it's hard 'cause the, the matrix or the matrices are getting massive for what you really need to, to test.
James Ward: And then there's, you know, all the combinatorial things. It's like, okay, the user has the MCP [00:29:00] server, they have the CLI, they have the MCP server, not the CLI. Yeah. They have, you know, like, like and then what documentation pages get fetched and like it, it gets really massive to... So yes, we can get the data, and it could cost a lot of money to get the data because, you know, all of these evals and this giant matrix of combinatorial things, like that you really should be doing, um, those are all expensive LLM calls, so yeah.
Demetrios: Again, enter Jev to the rescue.
James Ward: Yeah. Um, I, I think it's an area that will be really interesting to see. The-- Jev helps just on the judge side, so it doesn't necessarily help on... Like, I still need to like actually execute this task in an agent. Yeah. And until, until those agents switch over to Jev or something like it as being the, the, the like loop harness part of it, then we're still gonna use a lot of tokens just [00:30:00] to run a test in a harness, like actually have the, the harness do something.
James Ward: So, so the, so on the judge side, we'll, we'll save some by, by having Jev, you know, take a transcript or whatever and judge it. Um, but, but, uh, on the actual execution side, we're still, still gonna be in the harness, so still.
Demetrios: I, I saw one pattern. Actually, I, uh, I was talking to some folks last week, and they were saying, "Look, we have decided that we're going to expose five tools maximum on our MCP servers.
Demetrios: The four are going to be like the most common use cases," like you were talking about, "and then the fifth is going to be basically a tool search so that it can go a little bit deeper and reach those ones that are not exposed on that first glance." Have you seen any creative ways... I know that Cloudflare got popular, whatever, when they, I think they were calling it Code
James Ward: Mode, right?
James Ward: Did their, like Code Mode. Yeah, yeah.
Demetrios: Yeah.
James Ward: [00:31:00] Yeah.
Demetrios: With the two tools
James Ward: Yeah, so it's, it's a really interesting time in figuring out, like, what is the right architecture, and of course it, it depends. So in the case of AWS, we have something like 16,000 APIs or something like that. Like, we can't model all those as a tool.
James Ward: Like, like, that's just not even a viable option. Like, you're-- They're- Mm-hmm. Like, I think you'd break every agent on, on the planet if we, if we had 16,000 tools that got returned, you know, when you said list tools on your MCP server. And so that's just not even an option. And so we initially, with our MCP server, we had a tool that was, like, execute API call or something like that.
James Ward: W- which is, uh, was essentially like, like making a CLI call. It was in the shape of a, of a AWS CLI call. And th- that worked okay, but what we found through looking at these things was that a lot of times we needed sequential operations. Uh, [00:32:00] we needed to provide the agent a way to do sequential operations, so we added something like code mode, the code mode tool.
Demetrios: But
James Ward: then- And then we actually deprecated the, like, single call one, where, like, the agent can just write a-- You know, if it only wants to call one API, it can write the Python to do that. And so-
Demetrios: Yeah ...
James Ward: that's where, that's where the AWS MCP server is today, is it essentially just has just code mode, and that's because the API surface is just too large to cover.
James Ward: But that's where I'm exploring the idea of the DSL with, with the grammar, with the schema, you know, all those validation with something that's human reviewable as, as being a low-code mode code mode. So, so yeah. So currently exploring that as a, as an alternative, but at least today it's just the MCP server's just code mode.
Demetrios: And have you thought about, uh, trying to scope down different MCP servers so potentially you just have an MCP server for DynamoDB, and it's [00:33:00] really good at that?
James Ward: Yeah. We, we have looked at that, and there is some of that that exists. We have, we have a bunch of-- We've got our main MCP server, and then we've got all the community or, like, like, I don't know, broader ecosystem MCP servers.
James Ward: And so those are up on the AWS Labs GitHub, and there are more point-oriented MCP servers. And so that does exist. I think, um, it's not totally clear, like should those get rolled in somehow and do like, like- Yeah ... provide a search across those? Because part of the challenge is discoverability then. It's like people- Yeah
James Ward: install the AWS MCP server, and then, like, then they, I, I don't know. I think I'd assume that then I've kind of got everything I need, but I, but I don't because then there's these other MCP servers. And so, so there's I guess this balance of like ease of use, discoverability, functionality, not, you know, bloating things too much, and, you know, for the people that don't wanna use [00:34:00] Dynamo, um, then, but they get the Dynamo tool.
James Ward: Like I don't know. It's, it... Like having such a big API service definitely creates some additional challenges.
Demetrios: Yeah, 100%. That's what I was just thinking, like how do you optimize for that? And it really-- I haven't-- I'm still on the fence on what I think about this idea of we're gonna have the one agent to rule them all, and it'll be able to just kind of sprawl its other agents and build the agents on the fly and all that fun stuff, or we're gonna have very, like, domain-specific agents, and we're gonna have thousands of them or hundreds of thousands of them, right?
Demetrios: And- Yeah ... uh, it's probably not one or the other. It's some kind of a blur in between. But when I think about- It's
James Ward: like monolith versus microservices. It's like- Yeah ... uh, both. Yes. No.
Demetrios: Yeah.
James Ward: Yeah. It depends.
Demetrios: Yeah. It's like monolith or microservices. [00:35:00] Yes. That-
James Ward: Yes. Exactly.
Demetrios: I want them. So- Yeah ... but I, I do think, like in a way, when you tell me about some of these difficulties of the MCP server and having a gigantic AWS MCP server, then if you created smaller MCP servers for the different services, you would want those domain-specific agents so that they could have those specific MCP servers that you could call that agent and it would know like, "All right, cool.
Demetrios: I'm just like the DynamoDB agent, and I know exactly what I need to do here."
James Ward: Yeah. Yeah Yeah, at least today our MCP server is not itself an agent. But you, if it did become one and then we expose like agent as tool, then, then yeah, then maybe there's some other things that we could do there. Um, also, you know, a year ago, tool search tool was not something that even existed.
James Ward: And so now with, with [00:36:00] tool search existing, you know, that opens up some additional opportunities. Um, and then, and then maybe l- like you're saying, like having a tool search built into the MCP server and, and could we make that work? Um, and then the skills and skills over MCP or maybe another- Yeah ... piece to this as well is like, like maybe we use skills as the orchestrator to then be able to know what tools to call or something like that.
James Ward: And so, um, so yeah, I think we're, we're s- there's still a lot to, to kind of figure out what the most optimal patterns are around all of this. And so it's fun to do all the exploration and figure out like what's actually gonna work for our users and make their experience great, and how can we- Yeah ... eval it and get data and- Yeah
James Ward: yeah, there's just, there's so much fun stuff that's-
Demetrios: And not get prompt injection warnings when you do.
James Ward: Yeah. That's right. Yeah, yeah. And try not to break things in the process.
Demetrios: Yeah. Yeah.
James Ward: So
Demetrios: that's great. There, there was a big one that was kind of like a reoccurring theme that I've [00:37:00] seen across these MCP dev summits that we've been doing, and that's how folks are doing auth with their MCP servers.
Demetrios: Have you thought about that? Because I imagine, like you were saying, you don't wanna m- be so willy-nilly with your AWS account. It can get out of hand really quickly if you give it the wrong permissions and the wrong authentication to the wrong people.
James Ward: Yeah. Yeah, it definitely is an area that, that we're, you know, trying to figure out what the most optimal options are.
James Ward: Uh, w- the AWS MCP server does support OAuth now, which is cool 'cause you OAuth, you, you give it an account that specifically has the permissions to be able to be, like, used from MCP, and then you, you likely w- uh, a experienced AWS user would then lock down that particular role to only be able to do certain things.
James Ward: It's like, maybe I only want read operations. And- Yeah ... um, and so with [00:38:00] IAM and AWS, you can actually set that up and be like, "Okay, the account that I'm giving the MCP OAuth server can only do these things," and that would be, you know, a good approach to, generally to, to lock things down. When you've OAuth'd from an agent, there is no concept in agents of, like, multiple profiles, which maybe that's- Yeah
James Ward: a future need on MCP. But that's one of the challenges is that most of the time when you're working with AWS, you've got different profiles and, you know, this profile, you know, can do this, and this profile can do these other things. And you, uh, uh, from the AWS CLI, you can toggle those profiles. And so we do have a way, whether you're using AWS CLI to toggle profile, or there's a proxy that we have for the AWS MCP server that allows...
James Ward: that supports multiple profiles. So when you're using OAuth, you can't toggle profiles, but when you're using this, like, SigV4 proxy for the AWS MCP server, you can actually toggle profiles. And [00:39:00] so it, yeah, it's a, a complex space when you get into auth and security, but I think we're trying to make sure that we're supporting the different use cases of, you know, what users need and how they work with AWS and how they wanna maintain security and least, least privilege and all those sorts of things, so.
Demetrios: Well, and sometimes it's just knowing that I can do that. Like you said, I, I didn't realize I could do it, and I didn't read the documentation or I didn't see that update, and then I hear you say it now and I go, "Oh, sweet. Well, I'm gonna do it that way." Like, I didn't realize I could do it with the CLI.
James Ward: Yeah.
James Ward: Yeah. Yeah, and there's skills that point to the CLI and tell it to use the CLI. And so, yeah. It's, part of the challenge is that because things are moving fast, we have lots of different ways to do things, and so- Yeah ... what is best for every user, it depends, you know? And some users, like, they just want the easy route, and so MCP OAuth, like, like, great, like, go that route.
James Ward: Other peoples will wanna, you know, manage their skills manually and use the AWS CLI and [00:40:00] use profiles and, you know, have a more, more, you know, advanced use case. And so, so yeah, as usual, with complex things, it depends. And so we're trying to help, you know, make sure that people have what they need and can do things in the ways that work for them.
Demetrios: Well, James, dude, I'm so excited for your talk and- Thanks, Tim ... I appreciate you coming on here explaining it to me. For anybody that's listening, you can come and see James give his talk at Agent Con North America in San Jose, October 22nd and 23rd. And as a special gift, I realized I've got a discount code for you.
Demetrios: You can use the- Nice ... discount code COMMUNITY25 for 25% off, and the ticket prices are gonna go up at the end of this month, September. So go ahead and get that before they go up.
James Ward: Yeah. Oh, and I should do a shout-out. So I'm on the technical committee for the Agentic AI Foundation, and we've- Yeah ... got these work groups which are great ways to get involved into kind of specific topic areas, [00:41:00] and we're gonna, in San Jose at Agent Con, be getting the work groups together and, you know- Nice
James Ward: I'm looking forward through the technical committee to be interacting with the work group members. So yeah, like shout out, like join work groups, like they're, uh, public on the Agentic AI Foundation website now. Join the work groups. Join us at Agent Con. We'll get some good in-person time there to talk through all the, the fun areas, you know, of, that agents relate to, like the things we've just been talking about.
James Ward: So yeah.
Demetrios: Excellent, dude. Well, I look forward to it. I'll see you soon.
James Ward: Sounds good. Thank you.