Welcome back everyone to day three of
the Kaggle and Google five days of AI
agents intensive course. I'm Smitha
Khoulen, senior developer relations
engineer here at Google Cloud, and I'm
here again with Anand Nagaraja. Anand,
welcome back.
Thanks, Smitha. Hi, everyone. Welcome
back. I'm Anand, the founder of this
course and
an evangelist and product manager at
Cloud.
Take it away, Smitha.
Good to see everyone here again. Drop in
the chat what actually really surprised
you the most from day two. I'm also
curious if anyone went down the rabbit
hole on the managed MCP servers after
yesterday's session.
Let's start off with a quick overview
before for anyone new who's joining us
today.
So, throughout the week you're going to
be getting white papers, companion
podcast, hands-on code labs, daily live
streams, and AMAs just like this one.
And also, an optional capstone project
at the very end where you can compete
for Kaggle certificates, badges, swag,
and recognition across Kaggle and
Google's social channels. So, if you've
registered, the content will actually
land directly in your inbox. If not,
everything is going to be on the
Kaggle's learner portal, and also it'll
be announced in Discord.
Also, huge shout-out to all the people
who put everything together, the Google
researchers, the engineers who wrote the
white papers, the speakers joining us
this week, and also the Discord
moderators who have been answering
questions non-stop.
With that said, let's actually get into
an overview of day three.
So, day three, right? Yesterday was
about how agents actually reach
outwards,
tools, other agents, payments. However,
today's the opposite direction. How can
an agent actually manage what it knows
without falling apart as you ask it to
do more?
So, the problem is probably one you've
already hit.
Um the obvious way to make an agent more
capable is to keep adding to the system
prompt. More instructions, more
examples, more tools.
It works for a while and then it stops.
You know, the agent will pick the wrong
tool, maybe forget earlier instructions,
or even start hallucinating. So, the
white paper today calls this context
rot. And then that's actually goes it
goes really in-depth into how you can
avoid stuff like that, right? And the
answer the paper actually proposes is
something called agent skills. So, the
format itself is really simple. It's
just a folder with a markdown file
called skill.md plus optional scripts
and reference files. That's it. It's you
know, clever part is actually the
progressive disclosure. The agent only
sees a tiny bit of the metadata for each
skill at startup. So, maybe 50 tokens.
And then the full instruction only loads
when the task actually matches. So, you
can have 50 or even 100 skills installed
and you know, still keep that context
really lean.
That's also why this format spread is
really fast and it's not locked to any
vendor in particular. The same skill
actually works across multiple different
uh IDEs, Antigravity Cloud Code, ADK,
any sort of other agentic tools.
Basically, anything that supports that
standard.
With that said, Anant, do you want to
walk us through the white paper in more
detail?
Thank you, Smitha. So, yes, everyone,
welcome back to today's white paper
overview where we are focusing on agent
skills, the the para- the core paradigm
uh outlined in the white paper. So, over
the last two days, we looked at a lot in
our two white papers, everything from
redefining the software development life
cycle, and also exploring how open
protocols like MCP and A2A, as well as
the other
protocols and interfaces that we
discussed like A2UI, connect models to
tools and other remote agents, and a lot
more as well.
Well,
in today's white paper, we're going to
look at agent skills. So, if MCP is the
hands of your agent, agent skills are
the playbooks. In this white paper, we
define the new architecture primitive, a
self-contained skill folder as we have
outlined centered around a single
skill.md file accompanied by scripts,
references, and assets.
Now, we also highlighted the high
elegance of progressive disclosure, as
we have highlighted as well,
by loading only lightweight metadata
initially and fetching detailed
execution guidelines and scripts
strictly on demand, we can give our
agents dozens of capabilities without
causing any of the context drop or
blowing our token budget beyond what it
is set to be.
This pattern lets a single
general-purpose agent dynamically flex
into specialized roles offering a highly
maintainable alternative to complex
multi-agent setups, which can be hard to
maintain and trace.
Now, but the question that
often arises when we discuss this is,
how do we actually prove that these
skills work? Right? I mean, sounds nice
theoretically, but do we really not need
multi-agents? So, we also discussed the
evaluation of skills focusing on trigger
failures, output errors, and tool
trajectory analysis. In some ways, um
slightly aligned with how we
eval
multi-agent setups as well. We explore
We explore the differences between
simple success rates and rigorous
consistency uh, such as pass at K, and
then we also showcased Google's Agent
CLI, which hopefully some of you had a
chance to play around in the code labs,
illustrating how you can scaffold, test,
and deploy skills into the runtime
environments. Finally, the last part of
the conclusive part of the white paper,
we entered the meta skills territory,
looking at how agents can watch
successful execution traces and
autonomously write or optimize their own
skills.
So, uh, that's a summary of what we
covered in the white paper, and we also
have some incredible incredible labs
built, which my colleague Polong will be
covering in the latter part of the live
streams, but let's move on to the guest
QA of the Usmita.
Awesome. Okay, let's get into the QA.
We have some amazing guests today. Uh,
we have Devanshu, uh, Gabriella, Julia,
as well as Tanvi. Thank you all for
making the time. Uh, Anand, over for you
over to you for the first few questions.
Thank you. So, the first question is for
you, Gabriella and Julia. So, as public
registries for agent skills scale, what
security and verification mechanisms are
required to ensure imported skills are
portable and safe across different model
families for coding? We see a lot about
security risks, uh, when we
uh, security injection risks and a lot
over the news cycle. So, what's your uh,
opinion on this?
Yeah, um, so I really love this
question, and thank you to the person
who asked it because I do feel you. I
feel like I'm drowning all the skills
war. They are appearing everywhere. And
what is happening is like a skills, the
format is very portable by default. So,
it trans literally everywhere, but it
seems like over the time this runs
everywhere quietly turn into runs safely
anywhere. So, uh, think out there like
around thousand public skills, I found
out that roughly one out of eight had
critical vulnerability vulnerabilities
like hard-coded secrets,
funding home, prompt injection, all of
that talking to the body.
So, but I want to think about what are
the main two problems of this because
there are two problems wearing one
single code. So, the first one is an
easy one because a malicious script is
malicious everywhere. So, you handle it
the way you handle any dependency. You
scan it, you check for secrets, you
check who published it, the provenance,
etc.
The second problem
is a little bit more complex because
it's where this portability, this beauty
of portability bites back. The same
skill can be perfectly safe from one
model,
but it can be very dangerous on another
model because every model has a
different backbone. Some models
follow pushy instructions to eagerly.
Some other models have different
guardrails and stop it. So,
behavior doesn't travel the same way
than text and your verification step is
not on the skill only, but also on the
model level that um
that is using that skill. Now,
what you can practically do today and
what we have seen uh organizations that
we've been working with are adopting.
So, the first one is a registry
approach. So, similar to what firms
adopt, every skill gets gets a trust
tier. For instance, the built-in skills
ships with the agent it's like not
audited. The official from the projects
on catalog are also not audited.
The trusted from publishers like Google
and traffic, they are also trustable,
but the community one, this long tail,
is what gets can check and even can get
blocked. Another pattern organizations
are adopting and I think it's a it's
personally I I I I it will make my life
way easier if we all adopt this one is
similar to what Nvidia is doing which um
I think around May 26th they launched
verify agent skills and basically it has
three pieces is the skill factor which
is looking for code vulnerabilities,
agent specific prompt injections, um
cryptography singing knowing what
changed after um this skill was achieved
and the last one is um the skill card
which is a machine readable file that
tells you who made it, what access has,
what are the limitations. So,
the idea is very simple is like to trust
the info that arrives with the skill
like
um is like not only having the skill but
like this uh
this artifact that makes you feel like
this is like uh secure and this is not
like a standardized today but we are
seeing this need and organizations
approaching it a very good way. So, I
will leave you with one prediction is
like the skill registries that will win
are not the one with the long listing
but are the ones who uh which treat um
is uh evaluation suites and safety card
as a first class cheapable part of the
skill itself. Um so, I don't know if I
want to make clear it like a skill is
part of the context engineering that the
agent gets, right? So, many of these
things can be propagated to other
things. So, it's not something new
something that the model receives. So,
we can adopt also many things that we've
been adopting ever since LLMs uh emerge.
Um I will stop here and and I'm also
curious to hear what you have to say
Julia.
Yeah, absolutely. Thanks Gabriela. Um I
think you did a great job covering a lot
of the details at the high level I would
say probably there's one unifying
principle for me that I talked to all
the organizations we work with about
which is don't ever let safety fully
depend on the model or on the set of
skills. You want to have capabilities
and skills that are explicitly in host
and forest make instructions explicit
rather than implied and put hard
guarantees that we're used to like
Gabriella said in the beginning from
sandboxing, scoping permissions, egress
control, etc. Around the outside so that
you can also switch out different models
and test them along the way and make
sure you're doing that before you
publish them or before you kind of
publish any any agents or any skills
live. So I think those that is really
the the key thing at the end of the day
to treat this in many ways and
with some of the principles that we had
around classic software engineering and
leverage those
around around this area and don't ever
delegate it fully to the model and
assume that it's going to go kind of
down that path.
Amazing. Thanks a lot guys.
It's I I love both of these perspectives
and you know I just gave me the idea
like an app store for skills where
everything is verified is definitely
something and verified across different
models.
That's definitely where I
see
something a must do in the community
as well in addition to the custom
controls which you mentioned that we
have to put on our side. It's still good
to know that something is secured and
tested and there's no malicious prompt
injection in there.
Perfect. Amazing. Shall we move on to
the second question?
So this one is for you Debanshu.
How can file-based skill patterns such
as the skill.md be leveraged to help
future foundation models dynamically
retrieve and follow multi-turn run books
natively.
Thank you for that question, Anand. It's
an excellent question.
Now, there are three aspects to this.
First, looking at how skills dynamically
retrieve knowledge, we have to solve
context drop.
Feed an AI a massive manual today, and
it ignores instructions that are buried
in the middle.
The skill.md pattern fixes this via
something called progressive disclosure,
acting as the AI's procedural memory.
Instead of loading everything up front,
the model keeps tiny metadata footprints
in active memory,
retrieving only the specific
instructions it needs at the exact
moment.
Second,
this is how they will follow multi-turn
run books natively,
using DAGs, or directed acyclic graphs.
Instead of accumulating an execution
history in a single bloated prompt, the
DAG controller handles handoffs by
passing schema references between
isolated nodes via a message bus.
The model finishes step one,
flushes the old instructions to protect
the attention capacity,
and cleanly swaps in the new capability
profile for step two.
Third,
this structure will help future
foundation models evolve through meta
skills.
Looking ahead, models will author and
optimize their skills themselves by
observing successful workflows to draft
new run books.
However,
any assumption that this will be a fully
autonomous
thing right away is flawed. Anything an
AI writes
must start as a draft.
Without human spot-checking, a model
might optimize for the wrong metrics and
make the library worse.
Maybe I would like to add something here
because we've seen this pattern in the
past.
What happened is like some models the
beginning was using like tools like
Google search as external tools as a
part of the environment, right?
Eventually we found out customers
organizations really using it and they
don't want to select this a tool. So
what happened in the newest versions of
the models was like it was natively
trained this model so when the user
asked something it was automatically
like calling a tool like in this case
for instance Google search. So I think
we might think that this is a still a
research question whether we want the
same to happen to skills that
automatically the model is trained to
also learn this like all the skills by
default calling okay I have this set of
things that I can use and I think this
is a research question and this is very
new but it will come also how the people
is using how is adapted
and
we've seen this pattern before with
tools.
Makes sense. Yeah.
Awesome. With that said I think we can
move on to the next question.
All right.
Perfect.
All right, this question is for you
Tanvi. In complex multi-agent crafts,
how can state passing across distributed
skills be optimized without flooding the
models context window as a shared
messaging bus?
Thank you Anant and I think Dibyanshu
and Gabriella also touched on these this
question a little bit.
I would like to share a story probably
from a customer. So we were once
shipping a graph of skills that was
basically flawless in the demo.
Everything works well but it was
silently being degraded in the
production. So,
everyone said it must be the model, but
it wasn't. So,
the problem was that the full output was
taking
every node was taking the most context
and sending it it as a prompt to the
next one. And by the fourth hop,
basically, the model was rereading
everything in the upstream and the
context wrote really sad in, right? So,
the fix for this problem was that we
moved the state of the prompt into a
file bus. So, moved the entire state, so
decoupled the state. And the basically
the graph controller will own the state
now and not the conversation. So,
execution history will be
um
will stop accumulating anything inside
the model now. So, this is this was the
one way which we solved this. And the
second
solution was pass self references, not
the values. So, instead of sending the
entire JSON payload of the information
to the next node, we were passing the
pointers instead to help with the
reference of the information. And the
third way was to protect the model's
attention was
by having heavy data sitting outside of
the text input. And the context really
stays small and the model's capacity
only
looks into the reasoning of the
preserved context. So,
these these were the three ways that we
actually fixed this. And the second kind
of a second order benefit which we get
out of this was that the system became
debuggable, right? So, the system was
sitting on the disk.
It was inspectable instead of like
having uh somewhere there is a deep
conversation in the fifth hop, the whole
um transition actually changed the
setup. So, in the end like I would like
to add that the context window is not
really a database, so let's not add
everything in there and uh let the model
pass the handle to the system which can
uh do a lot more than just an LLM
context.
And some ways I think um
skills it's just prompt engineering
combined with dynamic loading. So, all
the best practices with which apply to
prompt engineering and dynamic loading
of things like only load what's
necessary.
Yeah.
Make sure that uh
the graph is defined to set the right
decisions get made. Uh
all of that uh should be applied. Uh
thanks a lot and making things
debuggable is especially important in
non-deterministic existence.
So, like on the using
Can we that we can summarize this like
pass pointers
uh and not like data, like not context?
Just pass pointers.
Exactly. Yes, that's the main point
here. Yes.
I think the a follow-up question to a
lot of what Gabriela and Tanvi and also
like Julia has been have been saying is
um so, imagine you have a lot of skills,
right? But
you have a question that maybe the user
is trying to solve, a problem the user
is trying to solve, and multiple skills
can actually be used to solve this. Uh
there's no actual right answer. So, how
do you kind of test which path the agent
should go down on? Or what's kind of the
most accurate one if there's multiple
skills which can help solve something?
I have
I have Yeah. I have a very because I've
been thinking a lot about this. Guys, I
feel like we are being
thinking and fixating too much on doing
self skill improvement. I don't believe
on that. It should be a skill
library improvement because this is a
actual a very important problem. You can
have two skills that they have very
ambiguous description that they are more
or less doing the same and they are
colliding. And honestly, can be a skill
we both have faced that and it was like
horrible. So, if you I will suggest this
you take it don't consume all the
skills. Like know which skills you have
for your project. Do optimization of the
skills for your project at the library
level, not only at the skill level and
have evaluation data. Like I think we
cannot address enough in all these five
days how important is evaluation at
every part. But I think if you have
different skills do not optimize at
skill level, optimize at the library of
skills level.
So, so are you telling
are you telling me that skill is not all
you need?
[laughter]
Skill is all you need but control it.
Like I mean like other ones like
I mean I I don't want to say it as
strong as I say it here Julia. Over to
you.
No, I was I was just going to say I
think that in reality you have to think
about it in terms of the whole system,
right? In terms of the whole the the
whole all of the pieces rather than just
a single component within it because
every time you're looking at just a
sliver and and a component it can always
collide with another with another piece
which is I think the thing that that
that people are kind of
calling out here and I think that's
really what it comes down to. We talked
about this in the very very very early
days of agents. When Patrick and I wrote
the first version of this agents paper
where we went down the path of
essentially
every single time you pass something
off, you're going to lose kind of some
some kind of
you're going to probably get some kind
of performance degradation, right? And
so you want to look at this whole this
this from an entire process and then
break it up into pieces and but then
bring it back up into into that entire
process because that's the only way
you're really going to get a truly
performant system out of whatever you're
trying to leverage and build here.
I would add just one more point is that
yes, put it in the system and also have
a routing system of when and where we
need which skill and then route it
towards that skill.
So, it seems like even skill management
is going to be a huge thing.
Or probably still already is.
Yeah, prompt management is skill
management. All right, let's move
I wonder if it'll just be skills or
actually go into something like tasks
and and things like that when you're
when when you're describing the goal
rather than just the individual
component primitive level, but yeah,
agreed.
Perfect. I'm glad this question was so
deep. Let's move on to the
[laughter]
All right, thanks and a really great
context in all of those answers. We're
going to now head on to our first
community question from what scale
does the token size of progressive
disclosure metadata trigger context
route? And is hierarchical routing
emerging as a viable solution to manage
it? Anvy, over to you.
Thank you, Smitha. This is exactly what
we were just talking about routing and
absolutely. So,
this is a major worry and it's natural,
right? So, um
it's always on the metadata. So, let's
say 100 skills is only 5K tokens, for
example. And if we do the math, then we
we are safe if we have 1,000, right?
But, the context route is not really in
the tokens, but it is like a gradient
which
is which will be a distractor for the
catalog. So, if you have a catalog of
skills, then
uh all of the others which look alike as
the first one, then these are all
distractors. So, routing really breaks
uh with let's say 100 skills and not a
few thousand, for example. So,
uh, the catch here is that the overlap
is the issue, not the size. So, everyone
who is reaching for a semantic
retrieval,
they might look for trade-offs with
misrouting for, let's say, you know, it
may happen silently with a recall
problem.
But, um,
what we are trying to fix here is
emerging through hierarchical routing,
right? So, instead of having a flat list
of everything, you have a router pick
top top picks and then,
uh, give a capability profiling to your
skill and then, only the top 10 or, for
example, 15 skills are exposed to the
metadata, right? And then, always the
context is bounded by the profile. So,
you don't need to send the entire
library,
um, to the prompt. So, that's how,
uh, you know, it can be solved through
hierarchical routing. And, um, in the
end, like, I think how I would, um,
summarize this is,
um,
the route isn't about how much you are
actually loading, um, it's about how
many skills are actually
looking alike when you load them.
Yeah, that's perfect. That's actually
exactly what we were discussing, too.
So, this is a great question.
Um, on to the next one,
we have from Praneeth.
When skills, MCP, and tool calls are all
available, how do you design the
boundary so that skills handle the
know-how and MCP handles reach without
overlap?
Um, Gabriella and Julia.
Yeah, let me get started. This is
another philosophical question. So, and
especially like a very system design
question. Um, the other day I was in the
library and was like a book called find
your why or something like that. And and
that made me think every decision that
you do today with the engines, you have
to think why. Why I'm making this
decision? There are so many primitives
today out there and the interesting part
is like you can do hacky ways. You can
put like a skills inside agents and PD
or you have you can put a skill system
prompt. So there are many things that
you can do technically speaking or like
agents and MCP. So the putting
Technically speaking, you can do all of
that, but you need to ask yourself the
question, why I'm doing this? What is a
skill? What is the MCP? And in this
context specifically of what you say,
you have to ask yourself the question,
is this knowledge or if this is access?
So access is anything that touches the
outside world, databases, API for
pulling car information, MCP job, this
tool to access the weather, etc. Whereas
knowledge is everything about how to
think about the work, which is which
tools to call, in what order, what to do
with the results, what are the edge
cases, what are the gotchas, what is the
process to get this um
HR um journey, etc.
But what I want to make very clear is
like what we see in today across the
organizations is like you can create
these solutions or plugins like binding
these two, three, or four different uh
primitives or you can like bind skills
and connectors, MCPs, and sub-agents. So
if you want to do this, you have to ask
what is the purpose of this? For the
skill, what is the connector, and what
are the sub-agents? So you can say,
"Okay, the skill,
I'm going to put how to write earnings
update, how to run a diligent analysis,
etc. For the MCPs, I'm going to put like
API for credits to extract information,
and the sub agent will be the
delegators. So, I think the theory
is very clear what the boundary for the
skill is like the sides, processes,
gotchas, and the connector is the
acting. So, I think when you create
this, you have to ask your question,
"Why I'm doing this? What is the purpose
of the skill? What is the purpose of the
MCP?" And have this clear separation of
concerns, and you can build beautiful
plugins, you can build beautiful
solutions with clear separation of
concerns, and and it just gets elegant
and easier to maintain, etc. And etc.
So, I think Julia, do you want to add
something?
Yeah, I'd love to. So, I have a pretty,
um, maybe overly simplified framework of
of how I I how I I put these things
together, but I think about tool I think
about each of these things as a layer
that you're answering a different
question with. So, the tool should
answer, "What is the single action that
can be performed?" That means it's a
verb that can do a stable contract like
send a message, list files, etc. etc.
[snorts]
MCP answers, "What can I reach?" It's
kind of the,
for lack of a better term, wiring that
connects all the actions to a real
external system. So, you think about the
connection logic, the authentication
data access,
um, all of all of those kind of
components. And then the skill is the
how do I go about doing something? They
are the know-how, they are the
conditional logic, the right sequence of
steps, your domain conventions, your
soft policies. They have no actual
capability, um, in them in that sense.
Um, so I think what I've been telling a
lot of people in in this is, if you can
do the following kind of quick test to
see if you've drawn the line, that is
usually a pretty good set of boundaries.
Um, at least we we've we found that in
the in the organizations that we've been
working with. If you're able to delete a
a skill's instructions, and the model
can still technically perform the action
even if it's clumsy, i.e. in the wrong
order with missing edge edge cases,
etc., then the boundary is pretty clean,
but if deleting it removes the ability
to do something, you've kind of
leaked some of the capability pieces
into the know-how and you should push
that down to the tool or the MCP server.
This is great. I think a a lot of people
have this question. Yesterday we also
got something similar on when what's
kind of the fine line between MCP and
A2A. Uh for skills in particular,
do you have kind of a heuristic for when
something that maybe started as a
a skill should actually get promoted
into its own MCP tool? Has that ever,
you know, come up in use cases?
I think so
I would generally segregate that a skill
I would I would a a skill the MCP is the
connection, right?
Is in in, you know, oversimplified terms
the connection and that's where I would
kind of draw that boundary. It's the
connectors that Gabriella referred to in
in in her structure in her analogy.
And the skill is kind of that knowledge
and know-how and I think I think of the
two of them
used together in many senses. I would
actually say a skill probably
becomes an MCP when you need to add like
the connection to an outside world for a
specific skill or something along those
lines.
Right.
Yeah, I mean what one thing I've seen is
like if you like have SOPs or something,
some people start putting into the
references into these things and then
they maybe here you have to ask, is this
the right question? Like do you I want
like because this also has to be secure
and we've been talking about it. What is
the right way? That's why we I go back
to why I'm making this decision. So I've
seen that because it's very simple, but
in enterprises, in organizations,
simplicity is not always the right way.
You have to ask, why I'm doing? Is this
the right way? How can I reach, how can
I scale, how can I do all of that? And
if we just go back to the strict
definition, why we want to scale, how we
want to do it, I think we find the
answer. And for me, simplicity, like
have everything short, elegant, is the
is the right answer. It's the separation
of concerns.
Yeah, and I would I would suggest that
Yeah, I would suggest that, you know,
this is actually a hard problem, right?
It's not easy, you know, having that
know-how and
making sure that you're concerning and
differentiating the what from the why. I
would ask, you know, our viewers to also
look to the community. There's a great
open-source library where, you know, we
can like check out what like the
community is thinking about skills, you
know, how are they approaching it?
And use that rather than build your own,
right? So, I think that's something we
should lean on.
Awesome.
All right, let's head on over to the
third question from the community. This
one's from Prasad. When building a
real-world AI application, how do we
decide whether to use a single agent
with multiple agent skills attached to
it or a multi-agent architecture? And
what are the key factors that influence
this decision? Julia, would you want to
take this?
This goes so much into the same
conceptual ideas that we just talked
about around skills, tools, and and
things like that. I think your default
here, and I think we've actually talked
about this in previous courses, your
default here is start simple and get
more complex as you hit boundaries. So,
you should start with a single agent
with multiple skills and split only when
you have some form of a specific problem
or something along those lines that
starts to force you to.
Every boundary that you add adds a
handoff or context, it adds latency,
failures become harder to trace, all of
kind of those things become a reality.
So, therefore, I would always say start
simple, and this holds true kind of in
in in all of these things. Start simple
as you're learning, as you're kind of
going through this process, and then go
into the tools that could become more
complex.
The key factors that I would think about
here are things like the context window
pressure. So, um again, we go into
bloating of of the context window and
and and the permanent context kind of
expanding,
uh tool count and how many tools you can
use in a um in a in a reasonable kind of
manner, and then ultimately selection
accuracy for some of those tools.
Usually, they can use 15 to 20 tools
well. It the accuracy starts really to
drop around 50 and with with most model
models today.
So, you're really starting to look at
kind of all of those pieces, and that
are some of those are some of the key
things that I would start to look at and
think about as I as as I start to
expand. But, in many in many cases, what
we're seeing in the in the industry
today is really start simple, start with
the the the components, and then start
to break it up as you start to hit areas
and boundaries because the more that you
put into the harness, and the more that
you put into the structure, the more
they'll likely have to change also as
you as as um as the model intelligence
starts to continue to grow.
I would like to add something here
because I'm very passionate about this
topic. Um in the sense like the beauty
of a skill in my point of view is like
they also came to change our mind, how
we were thinking about architectures.
And some architectures do not
necessarily need a multi-agent. So, for
all of you who read it what the white
paper, there is a beautiful organization
real-life example we have there where we
are discussing when multi-agents, when
uh
one agent with multiple skills. And I
agree with Julia, like you have to start
simple. But, again, you have to question
what is this about because multi-agents
are very good when you have parallelized
things, when you have communication
between agents, when you have all of
this, but imagine that you have
100 SOPs how to do something. Do you
want to create one agent and then
hundreds of agents? Imagine like every
time you add a new sub agent, you have
to add a new deployment. Like it gets so
huge. So, if this is very clear that the
agent when to use SOP one, when to use
SOP two, when to use SOP three, why
don't you use fast these like
skills in a skill way that the agent,
only one agent can do this. So, you are
simplifying the deployment, the agent
itself, and the scalability because like
adding a new SOP is only another empty
file, and not another agent that is
being deployed. So, I I I I think we are
coming from the same the simplicity, but
also the why like do do you need to ask
these questions like the
parallelization, all these things where
actually makes a lot of sense to have a
agent agent over over skills, but I will
I will let you read the white paper and
find this example that I'm mentioning
here.
I
I would just add one last thing to
Gabriella's point that I I would totally
ask this question and also ask if
the flat
systems are scalable. So, if they are
not, then we have to think about, you
know, if the flat catalogs are really
scalable. So, yeah. Then choose
routedness and in different ways.
[laughter]
And maybe on top of that, it's not just
about when you when you think about
scalability, also think about latency.
Are you building something that is a
live agent where you need to have an
immediate response,
otherwise it kind of irritates the user,
or are you building something where it
can operate with a higher latency and it
doesn't matter?
And and that becomes an another really
really important distinction on on what
you're you're you're choosing for that
architecture.
Well, thank you all. That was actually
great. And also as Gabriella pointed
out, do check out the white paper. I
think it does a great job giving you
kind of system design ideas on when you
should be, you know, sticking with your
single agent with skills framework or
moving to multi-agent. Um and then with
that said, we're going to go on to our
final community question from Tinawal.
Um
how does the anti-gravity
framework reconcile long-term memory
with skill versioning to prevent
reasoning hallucinations
when an agent actually iterates on a
past workflow using a newly updated
skill with altered execution boundaries?
Uh Debanshu, would you want to take
this?
Sure. Thank you. Appreciate appreciate
the question. Um actually relying on an
agent's long-term conversational memory
to manage execution boundaries is an
architectural antipattern.
Uh speaking for from experience, if you
try to reconcile past memory with a
version skill in the same prompt window,
you're almost guaranteed to produce
broad and compounding hallucination.
The white paper outlines how
anti-gravity handles this actively,
avoiding that reconciliation.
Instead of doing reconciliation, they
decouple state using DAG orchestration.
These are the exact mechanics that
prevent reasoning errors.
Instead of letting the model rely on
accumulated execution history to
understand like a past workflow,
the system extracts the payload from the
text input.
State is passed between nodes via file
message bus using structured schema
references.
Second, when the agent needs to execute
an updated workflow, it loads the
capability profile.
This acts as a swappable version control
bundle
that dictates the skills, tools, and
boundaries for that specific node.
Finally, to prevent the old execution
boundaries from bleeding into this new
logic,
the orchestration layer performs a hard
reset.
It unloads the previous system
instructions and flushes
stale variables before swapping the
version profile into memory.
By utilizing the strict teardown and
rebuild process, it is
the framework forces the model to
execute against the constraints of the
updated skill MD
rather than attempting to synthesize a
hybrid from that conversational memory.
Maybe if I want to add a clarification
here is in this case that the question
for this question when the agent gets
this wrong, it is isn't hallucinated in
the usual sense. It's not inventing
anything. It's faithfully following a
procedure that no longer exists, right?
So the agent is sent wrong, but the
memory is a stale. The memory is the the
problem. So to fix this is to treat
memory and skills as different things
with different life cycles. Memory is
episodic, what happened, when happened,
etc. And skill is procedural and
version, how we do it now. So if you
think how we are doing now, this skill
has been updated, you are like you know
like assuming this is a very outdated
skill. So you know that this is the new
version of this skill. This is the
skill, the source of truth that you
should use today, not on the memory. So
the skills are the source of truth,
version and current. Memory is the hint,
never an instruction. Pin versions to
the workflow the state, make every
boundary change visible. And one thing
that is very important is fail loud. So
if the agent
thinks like there is a discrepancy, tell
the user, "Look, this is the memory.
This is what I'm doing it.
This skill is already out of version."
So make part the versioning part of the
agent visibility. So the agent can
explain this to the user and give
options how to proceed and the user is
aware.
Um, so
uh, so I think this is the pen
regardless of anti-gravity or any
framework. This thing is something that
you will have when you have like
memories and skills and workflows, etc.
So, we I would recommend to follow this
um, this approach.
I love the concept of failing loud.
Uh, just to follow up on this, how are
they bunching and Gabriel? Well, how do
you think about, you know, kind of
backward capability uh, compatibility of
skills, right? Like should a V2 version
of a skill always be safe to swap in
with the the version one or is that
unrealistic?
Um,
if we were to think through, you know,
from a very ground up principle, I hope,
you know, that's what the case is, you
know, and unless like you're sort of
coming up with like a version of skills
that are
battle tested, right? And you're
swapping that with like a V2
version that is battle hardened, right?
I think that would that would create
like
sort of behavioral differences which you
didn't have with V1. But maybe V1 was
like failing
under certain conditions which were not
very apparent. Uh, but I would recommend
to always use the latest version
of the skills so that
uh, you have like a crowd sourced
information that we were talking about,
you know, in the repos.
I think there is not a single right
answer. What if you were having a
security vulnerability in the previous
skills? So, obviously there is no way
you want to go back to that, right?
Obviously, like no way you want to go
back. But what if the new version has
the security vulnerability? So,
obviously you want to go back. So,
that's why I think like failing loud
from the agent goes more than telling
the user is trying to find its
discrepancies. What is the difference?
What is the What does it mean actually
if I move from here here here we we like
it's just like getting more explicit on
what is happening between the versions.
And I think this is like just code
versioning. This is something we can do
to that.
Perfect. Thanks a lot everyone for your
insights.
Yeah, thank you to all the guest
speakers for coming on here to answer
all these questions. Really appreciate
the in-depth answers that we got and all
those, you know, nice conversations that
we had.
Okay.
All right. Now we're going to move on
into the code labs. Today we have two of
them.
Um, and they're actually going to take
you from using a skill to building
agents that use them. The first one will
walk you through how skills work inside
Antigravity. The second one is where it
gets really hands-on. You'll use Agent
CLI together with ADK and Pulong is
actually going to walk us through both
of them. Over to you, Pulong.
Hello everyone. I'm Pulong. Um, I'm a
developer lead for Cloud AI Google and
I'm super excited to talk to you through
some of the code labs today.
Um, so the first code lab, as Smith
mentioned, is on using skills, authoring
skills, and installing some pre-built
skills in Antigravity. So
in in this uh first code lab, you'll see
that you can create skills. And skills,
as was mentioned by the folks, um,
they're essentially just a skill.md
sitting within a folder. And it could
optionally have some scripts, some
references, some assets as well. And
here's an example of a skill. So you
would have a skill within a folder and
this skill is a database schema
validator. It may have a, you know, it
definitely needs to have a skill.md uh
file within it uh along with some front
matter at the top that describes the
name of the skill, the description of
the skill, and what the skill should
actually do. So this code lab will walk
you through how to create your skills in
Antigravity, uh the breakdown of skills.
And you also learn uh across a different
variety of skills uh in terms of
different levels of complexity, if you
will, as well. So there will be skills
that you'll be using in Antigravity, and
all that it will do is just have a a
prompt. It's just a set of instructions
of how you can, uh, you know, run
through the skill to help you with
formatting your your Git commits. And
then you also have other skills in which
you're actually going to add a little
bit of determinism into your skills. You
know, sort of as, uh, Gabriella and
others had discussed, the, um, the the
trajectory or the evaluation of skills
is is a really hot topic. And so, there
are times in which skills will be a
little bit non-deterministic because
they're essentially prompts. So, you may
want to think about sometimes you want
to enforce certain very deterministic
behaviors, um, within the skill, and the
database schema validator will be an
example of that where you'll be actually
calling upon a Python script to actually
run as part of the flow of the skills.
So, this is a really interesting one to
to look at. Um, and so, each of these
skills, uh, that you'll be learning in
this first codelab, will go through a
different variety of different kinds of
ways in which to to to build skills and
and and discover how the different
resources, uh, within the skill, uh, can
help you add more capabilities to have
the skills essentially run.
You'll also learn a little bit more
about, uh, Agent CLI, as Smitha just
mentioned. This will enable you to, just
using prompts, uh, create agents in ADK
that you can use to help you with
process automation or to provide agents
that you can, uh, have customers or your
users, uh, use directly as well.
You'll also learn about installing agent
skills, uh, directly in Antigravity, as
well. And so, this will make up the
first codelab that you'll be going
through uh today. Um, and so, as an
example here, uh, here is a skill. Um,
this is a Markdown file. You have your
front matter at the top, and then you
have some of the instructions below. And
in this one here, because of the
deterministic behavior that you might
want to have to validate a particular
schema, and you always definitely want
to run it in this particular way, then
you might have it in a code file like in
this Python script that will run as part
of one of the instructions steps within
this skill.
Okay. And then so the the next code lab
is actually a deeper dive into agent CLI
and the whole agent development life
cycle with ADK. And so for those of you
who joined us in the previous
Kaggle 5 Days of AI with Agents course,
you had authored ADK agents
with with code, right? This time this
code lab doesn't have
code that you'll be writing. You'll be
writing everything through through
prompts in Antigravity using the skills
and commands that are made available
through agent CLI. So you'll create a
project and then in this project you'll
be creating an agent. This is a customer
support agent and what it'll be doing is
to classify if the user query is related
to shipping because it's for a shipping
company. And then if it's related to
shipping, answer some questions. So
it'll look something like this. So I've
asked it, you know, where do you ship
to? Probably only ship within the United
States. Okay. How can I track my order?
And then it will, you know, say
we'll send you an email to track the
link. Okay, super simple. So you'll be
able to create this agent just using
prompts in agent CLI with Antigravity.
And then so I'll show you kind of like
how that sort of works. So here I am in
the Antigravity IDE and I've simply just
run like one of the final commands in
that code lab to test my agent which you
saw just earlier in the other screen.
And what I did was I said, "Launch the
local development playground for my
customer support agent." And you'll see
during the 26 seconds that it was
running, it viewed the skill, the Google
agent CLI workflow skill, which is part
of agent CLI
to understand how to run or launch the
local development playground. So it
knows the commands it needs to run, it
runs the commands, and then it launches
the the web interface for me to be able
to actually
see it in localhost, which is what you
see here. So, you see this sort of
end-to-end journey of building agents
just through prompts using Agent CLI,
including linting, including testing,
and
command line execution as well. And
then, of course, you can discover some
more things that you could use with ADK.
The skills that you'll be learning
throughout these code labs, you can also
serve inside the ADK agents themselves
as well. So, you know, you have skills
that you can use as part of antigravity,
you have skills that you can use within
the agent as well. So, you can think of
how you could share skills across when
you're live coding, you could create
skills, and you can even create agents
to to allow other people to use those
skills.
So, those skills could be same and
common across different sort of facets
and different layers as well.
So, I have I have listed a couple of
bonus challenges here in addition to
these two code labs,
especially based off of this wonderful
discussion from all these guest speakers
here.
The first one is
So, as you go through these code labs, I
want you all to be thinking about what
are some some manual tedious tasks that
you kind of have to do as part of your
work or part of your your studies or
part of your your daily life that you
might want to think about automating.
So, repetitive tedious things that you
might want to think about automating
that maybe you could convert into a
skill. So, the way that I would approach
it is in antigravity, try to do things
manually, and then ask antigravity to
convert that into a skill, to create the
skill for me based off of that.
So, okay, so that will help you create
some skills that will help you in your
day-to-day life.
And then, the second thing, which is a
really hot topic today, was on
evaluation. So, how do you how do you
think about evaluating the effectiveness
of your skills that you've just
authored? So, that's going to be a big
question for your particular workflows.
And then, finally, as Gabriela said,
the the why, the the big why question.
So, um
making sure that, you know, your skill
is solving your problem, not just making
sure that the skill works as intended,
but also solving your actual problems in
your day-to-day uh
lives, I think that's going to be uh the
biggest question of all and and thinking
in and seeing how skills fit into that
uh sort of larger equation. So, uh those
are my sort of soft bonus challenges to
these uh two code labs. Hope you have a
wonderful time with the code labs. I'd
love to see you all in the Discord uh
server. And uh that's it. So, back to
you, Smitha.
Thank you so much, Fulan. I highly
recommend running both of those code
labs end to end. You're going to learn a
ton from that. And then also, if you get
curious on how you can give your agent
capabilities to use Google services,
Google also has its own set of skills to
empower your agent to do exactly that,
and I'll be leaving those resources in
the description box below. With that
said, let's move on to the pop quiz.
Perfect. By the way, I was just
thinking, skills are like macros they
used to be in the olden times, uh like
record actions and repeat them, as Fulan
mentioned. So, that's how I would think
of them. All right, pop quiz. Uh first
question for today's pop quiz, take your
pen and paper, is
under the open standard defined at
agentskills.io, which file must be
present in every skill folder? This has
been discussed already in the white
paper and this live stream. Is it A,
skills.md, B, config.yaml, C, run.py, or
D, manifest.json?
The correct answer will be shown in 3,
2, 1, and it's A, skills.md. Talked
about a lot um
in the white paper and this live stream.
Uh let's move on to the next question.
How does the progressive disclosure of
agent skills optimize context windows?
Is it
through A, by loading the entire skill
body scripts and assets, B, by keeping
only metadata in context and loading the
skill with MD
strictly on demand, C, by
compressing all instructions into vector
embeddings, or D, by using vector
database similarity instead of token
window. The correct answer, think about
it. We discussed it a lot. Correct
answer will be shown in 3 2 1 and it's
B. We discussed a lot. It's more about
on-demand loading and routing and all of
those things we discussed earlier.
All right, moving on to the next
question. What phenomena occurs when
dumping too many static instructions
into an LLM system prompt leading to
degraded performance as the input size
grows? Is it A, overfitting, B, context
rot, C, epistemic drift, or D, token
compression?
So, think about it and the correct
answer will be shown in
3 2 1 and it's B. It's context rot. The
rest do not apply. There's no drift
happening. There's no ML model being
trained here and you're not definitely
not
compressing tokens.
All right, let's move on to the next
question. What is required for an agent
skill to graduate to the action allowed
tier?
Is it
A, simple peer model LLM
and an LLM is a judge email,
B, an offline run passing once, C,
golden data set review and manual spot
checking, or D, full adversarial red
teaming and sustained access across
multiple runs email runs?
Think about it and
the answer will be shown in 3 2 1 and
the correct answer is D. It requires the
full suite of things before it gives
gets that access privilege.
All right, to our last question.
What is meant by the engineering
practice of shifting intelligence left
by skill design?
Is it A,
moving runtime logic from the LLM prompt
into standard scripts, B, increasing the
overall system prompt size, C,
generating code before writing tests or
specs, or D, allowing the model to
autonomously design the database
schemas?
Think about what left could be here.
All right, the correct answer will be
shown in 3, 2, 1. And the correct answer
is A, where you move it more towards
um uh
LLM logic more towards the scripts
towards scripts.
All right, that brings us to the end of
our pop quiz and our session.
Um
Smitha, do you want to
Yeah, awesome. Thank you, Ananth. Quick
wrap off before we sign off. Day 4
assignments are going to be dropping
shortly, and tomorrow's topic is
security and evaluation, which honestly
might be the most important session of
the week. We've spent 3 days giving
agents more capabilities, right? Tools,
interoperability, skills, memory.
Tomorrow is about how you keep all of
that reliable and safe in production.
Uh another thing, keep the disc
discussion and going on in Discord. The
mods are active and also as a reminder,
questions get selected from Discord, and
you can stand a chance to win Kaggle
swag.
Um and also get started on the code labs
if you haven't already. Try writing a
custom skill of one of your own and see
how it changes the way your agent
actually behaves. Also, if you've missed
day one and day two's live streams, they
will all be linked in the description
box below. Thank you, everyone, and see
you tomorrow.
Thanks, everyone. See you tomorrow.