Imagine giving a a super smart AI a
million pages of your company's standard
operating procedures. You'd think it
would become your absolute best
employee, right? Like capable of doing
anything.
the dream. You just feed it everything
and assume it's going to work perfectly.
Right, but the science shows it actually
gets dumber. You hand it this massive
encyclopedia and it acts like someone
who just, you know, threw a giant
instruction manual into a blender.
It's a really fundamental
misunderstanding of how these models
actually process information.
I mean, we we assume more context equals
more capability, but in reality it just
creates this massive attention dilution.
And that is exactly what we are tearing
apart today. Welcome to the deep dive.
Today we're looking at the day three
white paper from the five day of AI
agents vibe coding intensive course by
Google X Kaggle.
It's a really fascinating read.
It is. And you know, if you're a
technical Kaggle competitor looking to
squeeze every drop of performance out of
your agent architecture, or even if
you're just intensely curious about how
AI actually executes complex tasks
behind the scenes, this deep dive is
absolutely for you.
Yeah, there's a lot to unpack here.
We are uncovering why dumping data into
massive AI context windows is, frankly,
a recipe for disaster. And why this
simple lightweight primitive called an
agent skill is quietly replacing the
entire bloated multi-agent ecosystem.
It's a massive paradigm shift. I mean,
we spent the last few years building
these incredibly complex networks of
like manager agents and worker agents.
whole swarm thing.
Exactly, trying to brute force
specialized behavior. But agent skills
completely sidestep that complexity.
They turn a single general purpose AI
into a hyper specialist,
but um only for the exact seconds you
need it to be.
Let's start with the physical anatomy of
this, because when I first read the term
agent skill, I pictured some uh really
dense compiled binary code or a deep
neural network weight adjustment.
Oh yeah, like some complicated
fine-tuning process.
but it's actually shockingly simple at a
structural level.
It really is. At its core, an agent
skill is literally just a folder on your
file system.
Wait, just a folder?
Just a folder.
Inside that folder, you have one
mandatory markdown file called skill.md.
And you can think of this as the brain
of the skill.
Okay.
It contains the metadata. So, the name
and a very precise description of when
to use it, along with the core
instructions.
Yeah.
And then, around that file, you have
optional subdirectories that handle all
the heavy lifting.
Right, I noticed the paper breaks down
those subdirectories into three specific
types. There's a scripts folder, a
references folder, and an assets folder.
Yeah.
Why do we need all of those if the
skill.md already has the instructions?
Because you really want to separate the
what to do from the how to do it
deterministically.
Oh, I see.
The scripts folder holds actual
executable code, like, you know, Python
or bash scripts. If the agent needs to
calculate a cryptographic hash, you
don't want the language model guessing
the math.
Definitely not. You want it to just run
a script.
Exactly. Python script that does it
perfectly every time. Then, the
references folder is for dense domain
context, like a 50-page PDF of tax
codes, which is way too long for the
main prompt. And the assets folder holds
things like JSON schemas or email
templates.
The whole folder structure isolates the
moving parts.
And the paper outlines two distinct
paths for how these skills actually get
written in the real world.
Path A is driven by subject matter
experts. So, I'm picturing
um an HR manager who has an onboarding
guide, or maybe a compliance officer
with a massive run book.
Right.
They don't need to learn Python. They
just translate their existing human
instructions into this skill.md format.
Precisely. It really democratizes agent
creation. You don't have to be a
software engineer to build a skill.
But path B is different, right?
Yeah, path B comes from the developer
side and it's much more dynamic. A
developer might watch an agent
successfully stumble through a really
complex multi-step workflow.
Okay.
And they take that successful trace,
clean up and crystallize it into a
reusable skill folder. Next time the
agent doesn't have to burn compute
figuring it all out from scratch.
It sounds a lot like a restaurant
kitchen.
A restaurant kitchen, how so?
Well, the skill.md metadata is
essentially the menu out front. The
agent, our chef, reads the menu to know
what's available.
Okay, I like that.
chef doesn't actually plug in the
blender, which would be our scripts
folder, or read the exact grandma's
recipe out of the references folder
until a customer actually walks in and
orders that specific dish.
That is a highly accurate way to
visualize it.
The technical term the white paper uses
for this is progressive disclosure.
Progressive disclosure?
Yeah. You only reveal the complex
mechanics of a task when the task is
actively triggered.
But wait, a lot of developers listening
to this are probably screaming at their
dashboards right now because they
already use MCP, the model context
protocol.
Oh, sure.
Or they just keep a giant agents.md file
in their project root. Aren't those
tools already doing this exact thing?
It's a common confusion, actually, but
the distinction is absolutely critical
for system design.
MCP gives an agent reach.
Reach?
Right. An MCP server acts as a bridge
connecting your agent to an external
system like pulling a record from
Salesforce or, you know, querying a big
query database.
Okay, so it connects to the outside
world.
Exactly. A skill, on the other hand,
gives the agent know-how. It's
procedural memory. It's the step-by-step
logic of what to do with that Salesforce
record once you actually have it.
So, the skill might say, "Here are the
five steps to process a customer
refund." And step two might be, "Use the
MCP tool to check the database."
Exactly that. They compose beautifully
together.
Mhm.
And regarding the agent.md file, that is
global, always on context.
Like project-wide rules?
Yeah, like telling the AI, always write
your code in TypeScript. Skills are
strictly on demand.
That distinction brings us right into
the token economics of this whole
movement. Because if a skill is just a
smartly organized folder, why did this
specific architecture explode in
popularity, especially with the vibe
coding crowd? You hinted earlier at the
issue with dumping a million pages into
a model.
Yeah, it solves a massive kind of hidden
crisis in AI engineering known as
context drought.
Context drought, sounds gross.
It is, computationally speaking. For a
long time, the industry's naive
assumption was that larger context
windows would just magically solve all
our orchestration problems.
Just throw more context at it.
Right. If you have a 1 million token
window, just dump every tool, every
instruction, every API spec into the
system prompt, and let the model figure
it out.
But the paper references that famous
lost in the middle study, and also a new
2025 study from Chroma research,
they proved that across 18 different
frontier models, performance degrades
silently as the input grows.
Silently is the keyword there. You don't
get an error message, it just gets
dumber.
Wow. So, even if the window can
physically hold a million tokens, the
model's accuracy, its ability to
actually retrieve and use that
information, drops off a cliff long
before the window fills up.
Because the attention mechanism
literally gets diluted. Every extra
token you add competes for the model's
focus.
Yeah.
If your system prompt is cluttered with
your instructions for 50 different
tools, and the user just says hello,
Right.
the model is wasting compute ignoring 49
irrelevant tools.
So, progressive disclosure fixes this by
keeping the noise out. If I have 50
workflows, I don't load them all, I just
load the metadata, the trigger
descriptions.
Yes.
And the paper notes that might only be
about 4,000 tokens.
When the user asks for a specific
workflow, that single skill triggers and
adds maybe 2,000 tokens for its active
body.
The other 49 stay completely hidden on
the hard drive.
You're taking 15,000 tokens of active
heavy context and compressing it down to
maybe 6,000.
The white paper shows this achieves up
to a 98% reduction in active context.
98%?
Yeah, it's huge. The model stays
incredibly sharp because it's only
looking at what actually matters in that
moment.
Hold on, let me just I'm doing the math
here. You said the metadata for 50
skills is 4,000 tokens.
Roughly, yeah.
If I'm an enterprise with 40,000 skills
in my library, just reading the menu of
metadata is going to blow out the
context window anyway. How does
progressive disclosure survive at a
massive corporate scale?
That's a great catch. At that scale, you
don't even load all the metadata into
the prompt.
Oh, you don't?
No, you use a lightweight retrieval
system, a a rag pipeline.
Okay.
So, when the user types a prompt, a
really fast embeddings model quickly
scans the 40,000 skill descriptions,
grabs the top five most relevant ones,
and only passes those five to the
agent's context window.
Oh, wow.
It's progressive disclosure taken to the
next level.
That completely reframes how I think
about routing.
But, let me ask you this.
If we can just dynamically swap skills
in and out of single agent, does that
mean the traditional multi-agent
architecture is dead?
Do we not need a manager agent routing
tasks to a swarm of specialist worker
agents anymore?
Multi-agent systems aren't dead, but
their role is getting much, much
narrower.
Okay, when do you still use them?
You still need them for genuine
asynchronous parallelism.
Like having three agents research three
different companies simultaneously.
Or when you have differing security
postures.
Security posture?
Yeah. So, if you have an HR agent that
needs clearance to read salaries, and a
marketing agent scraping public
websites, you keep those as separate
agents, so the marketing bot can't
accidentally leak payroll data.
That makes perfect sense from a security
standpoint.
almost everything else,
Mhm. the swarm approach is dying.
Let's say you're a logistics company
with a hundred different process
variants for shipping freight.
Right.
Maintaining a hundred different
sub-agents with a hundred deployments,
memory stores, and complex routing
layers
is just an operational nightmare. One
general purpose agent dynamically
loading a hundred different skills from
a folder is vastly more elegant.
Which sounds amazing until you realize
the danger zone we're entering. I mean,
if anyone in the company can write a
skill just by typing up a markdown file,
it sounds like a recipe for absolute
chaos.
absolutely can be.
If you're listening to this and you've
ever watched your agent just completely
ignore the custom tool you spent three
hours writing, this next part is going
to hit close to home.
The reality is quite stark. The white
paper references a 2025 study from
Skills Bench that looked at real-world
agent deployments. They found that 19%
of poorly designed skills actually
caused the agent to perform worse than
having no skill at all.
Wait, worse?
It actively subtracted baseline
capability from the foundation model.
That's wild. And a Versel study found
similar results, right? A 56%
non-invocation rate, meaning more than
half the time the agent just ignored the
skill completely. So, how exactly are
these skills failing so
catastrophically?
Well, the paper outlines four specific
failure modes. First is trigger failure.
Your description's too vague, so the
wrong skill fires, or the right one just
stays quiet.
Okay.
Second is execution failure. It triggers
correctly, but the internal instructions
are messy, so it produces wild tool
calls or hallucinates data.
Right.
Third is token budget failure. This is
when the skill author dumps a massive
unoptimized PDS into the references
folder, blowing up the context window,
and ruining the agent's short-term
memory.
And the fourth one is regression, which
honestly feels like the most insidious.
You add a brand new skill to the
library, and because its trigger phrase
is too similar to an existing skill,
it hijacks the routing.
Yes. You break a system that used to
work perfectly just by adding something
new.
Exactly.
This is why the white paper heavily
mandates evaluation-driven development
or EDD.
Right, EDD.
Before you write a single line of your
skill.md file, you must write three JSON
evaluation cases.
I love this concept. You have to define
the input, the expected tools to be
called, and the expected output format
completely up front. If you don't know
what success looks like, you really have
no business writing the instructions.
Yeah, exactly. You are testing the skill
in isolation using a single skill
sub-agent pattern.
Okay.
You give a blank agent only this one
skill and run the JSON evals. But here's
the critical part.
How you score those evaluations matters
immensely. The paper dives deep into
trajectory scoring.
I want to stop and unpack trajectory
scoring because the paper mentions modes
like exact, in order, or in order.
But like, if the user asks for a result
and the agent delivers the correct final
output, does it really matter how it got
there?
It matters more than the final output
itself.
Really?
A 2026 analysis by Latitude found that
if you only score the final output,
you'll pass 20 to 50% more test cases
than if you score the trajectory,
meaning the specific sequence of tool
calls.
Wait, if it's passing more cases, isn't
that a good thing? We want it to pass.
Not if it's stumbling into the right
answer through disastrous methods.
Uh.
Let's make this concrete. Imagine a user
asks your customer service agent to
refund my last order.
Okay.
The agent gets confused, pulls up the
user's profile, accidentally deletes
their shipping address, but then somehow
successfully triggers the Stripe refund
API.
Oh, wow. That's terrible.
Right. If you were using output-only
scoring, the test passes. The user got
their money. But the trajectory was a
complete disaster.
I see.
Trajectory scoring looks at the exact
sequence of actions.
For a read-only skill, like summarizing
a public document, maybe any order is
fine. But for an action-allowed skill
that touches production databases, you
absolutely need exact or in-order
trajectory scoring to catch that deleted
address.
So, if the model is just the engine
block, I wouldn't let a brand new
unproven engine drive my kids to school
without testing the brakes in a
controlled environment first. There has
to be a staging area, right? Like a
sandbox where these skills run without
actually pressing the buy button in the
real world.
You're anticipating the demo-to-deploy
gap.
Right.
Confidence always peaks during the local
demo, but it completely shatters in
production. Teams immediately blame the
model for hallucinating, but it's rarely
a model problem. It's an environmental
fit problem.
Yeah.
And this brings us to one of the most
staggering statistics in the entire
white paper involving Claude code.
Yes, the reverse engineering stat. This
completely blew my mind.
It's incredible. Researchers mapped the
entire infrastructure of Claude code
version 2.1.88,
and they discovered that 98.4%
of the agent's code base was purely
operational infrastructure.
98.4%.
Yeah. Things like permission
classifiers, session storage, retry
logic, context compaction pipelines.
Only 1.6% of the code base was the
actual prompt reasoning loop.
1.6%.
The AI part of the AI agent is barely a
rounding error. That proves your point
completely. The foundation model is just
a commodity. You can swap the engine out
tomorrow if a cheaper one drops, but the
agent skills, the guardrails, the
infrastructure, that's the steering
wheel and the transmission.
Exactly. That is your durable
proprietary asset, which is why we have
the graduation ladder.
Right, the staging area.
To your earlier point about a sandbox,
skills don't just go live. A new skill
starts at the read-only tier where it's
evaluated safely by an LLM as a judge.
If it passes, it graduates to
draft-only. Here, it can generate
outputs and draft emails, but a human
has to manually review and approve them
before they go out.
And to get to the final tier, action
allowed, the paper says it requires a
metric called pass to the K. What does
that math actually look like?
Pass to the K measures sustained
reliability.
Let's say a skill succeeds 60% of the
time on a single run.
That sounds okay for a quick prototype.
Sure.
But if you require it to succeed eight
times in a row without a critical
failure, so pass to the power of eight,
that 60% success rate drops to about a
1.6% chance of clearing the hurdle.
That is ruthless.
It's brutally effective. It exposes
inconsistent skills long before they
touch production data.
That is ruthless, but you know,
absolutely necessary. Now, earlier we
talked about Path B developers capturing
an agent's successful workflow and
turning it into a skill. We are
basically letting the AI write its own
instructions here.
Yeah, we are crossing into the frontier
of meta skills.
Meta skill.
The white paper buckets these into four
categories. First is authoring, where
tools draft a skill based on a simple
human prompt. Second is assisted
authoring from traces, which is Path B.
Third is improvement, where an agent
runs bounded automated experiments to
optimize an existing skill.
Right.
And fourth is library evolution, where
the agent analyzes, say, a month of chat
logs, notices a recurring problem, and
independently proposes a brand new skill
to solve it.
Okay, I have to step in here. Letting an
AI write, update, and evolve its own
operating instructions sounds
terrifying. What's stopping it from
optimizing for some bizarre metric,
over-fitting its own triggers, and
quietly destroying our entire library
over a weekend?
The white paper strongly agrees with
your paranoia.
Good.
Meta skills fall apart completely
without that pristine evaluation suite
we just discussed.
The absolute non-negotiable rule is that
anything an agent writes must enter the
library at the draft tier.
So, AI can never instantly deploy its
own updates.
Never. A human must be in the loop for
the first few edits. I mean, an agent
might optimize a skill's description to
pass a specific test, but accidentally
use a trigger phrase that steals traffic
from another critical skill.
Now,
a human reviewer can spot that
regression instantly.
But wait, keeping humans in the loop for
meta skills solves the library
destruction problem, but it doesn't
solve the traffic jam. If an agent uses
three different skills in a row to
complete a complex task, aren't we just
dumping all those intermediate outputs
back into the context window? Doesn't
that trigger the exact context draw we
just tried to escape?
Yes. That is the orchestration
challenge. You cannot use the context
window as a database.
Okay, so how do they fix it?
Production systems use DAG
orchestration.
DAG? Woah, directed acyclic graph sounds
like a terrifying college math course.
Are we just saying it's a one-way street
so the AI doesn't get stuck in an
endless loop passing files back and
forth?
That's exactly what the acyclic part
means. No cycles, no infinite loops.
It's a structured map of dependencies.
And to pass the data between these steps
without polluting the context window,
they use capability profiles over a file
message bus.
How does a file message bus work in this
context?
Okay, so instead of the first skill
printing a massive 10,000-line JSON
output directly into the chat history
for the next skill to read,
which would blow up the context window.
Right. Instead, it saves that JSON to a
hidden file on the disk. Then, it just
passes the file path URI to the next
skill.
Oh.
The LLM's attention is perfectly
protected. It only sees data saved at
file X.
The golden rule here is to shift
intelligence left.
Shift intelligence left. Explain that.
It means relying on deterministic
software constraints rather than prompt
engineering.
Don't write a system prompt yelling in
all caps,
"Always validate the zip code."
Because it's just hoping the LLM will
pay attention.
Exactly. Shifting intelligence left
means you write a rigid five-line Python
script in your scripts folder that
validates the zip code, and the LLM
simply calls that tool.
You rely on code, not vibes. I love
that. So, moving from the mechanics to
the real world,
how is this actually being used in the
enterprise right now?
Looking at the ecosystem, there are
apparently over 40,000 public skills out
there.
Google has their official repository at
github.com/ Google skills. How do you
even navigate that safely?
The paper gives three firm heuristics.
First, prefer first-party skills. If you
are querying BigQuery, use the official
Google BigQuery skill, not a random
community one. Second, pin your versions
aggressively so an unexpected update
doesn't break your DAG pipeline. And
third, audit the code before adopting.
Remember, a skill can contain executable
Python scripts.
Oh, right.
treat it like a supply chain software
dependency.
Let's ground all this theory with the
case study they included in the appendix
B. They looked at a massive home
improvement retailer. I really want to
know exactly how a company of that size
uses this architecture.
They use it to distribute knowledge.
Historically, an AI team at a retailer
tries to build one massive monolithic
retail assistant. It inevitably fails
because AI engineers don't know the
nuances of plumbing codes or lumber
grading. With agent skills, the
ownership is distributed to the actual
domain experts.
Let's walk through a specific customer
query. Say I go to their site and ask,
"I want to tile my shower. What do I
need, and can I get it delivered by
Tuesday?"
How does that route through the system?
The agent dynamically loads three
separate skills in sequence. First, it
triggers the project guidance skill.
This folder is owned and maintained by
the retailer's actual trades team.
Okay.
It loads the expert logic about
waterproofing substrates and grout
types. Once the agent figures out the
method, it drops that data into the file
message bus.
And moves to the next node in the DAG.
Exactly. It triggers the materials list
skill, which is owned by the pro
merchandising team. This skill
calculates the exact square footage and
wastage.
Awesome.
Finally, it passes that list to the
delivery window skill, owned entirely by
the fulfillment team, which runs a
deterministic script against the
inventory database to check Tuesday
availability.
The brilliance of this governance
structure is just staggering.
You aren't asking the AI engineering
team to become experts in plumbing
regulations. The plumbing experts write
in a simple markdown file and the AI
just reads it when asked.
And that distributed ownership is what
creates a company's strategic moat.
Right.
A competitor can buy access to the exact
same base foundation model you are
using.
Right.
But they cannot buy the codified,
version controlled institutional
knowledge of your best employees.
What a journey.
We've gone from unpacking the basic
folder structure of a skill.md to
understanding how progressive disclosure
defeats attention dilution and context
rot. We looked at why trajectory scoring
is vital to catch an agent that stumbles
into the right answer dangerously.
Yeah.
And finally, how to use a DAG and a file
message bus to deploy a distributed
library of institutional knowledge
without causing a massive traffic jam.
It's a profound shift. We are moving
away from trying to make the model
inherently smarter and instead focusing
on making the environment around the
model highly structured and
deterministic.
And for everyone listening, don't just
let this be an academic theory. We
strongly encourage you to actually build
this. Jump into the code labs provided
in the 5-day of AI agents, Vibe coding
intensive course by Google X Kaggle.
Absolutely.
Try out the practical implementation of
these agent skills yourself. The paper
gives incredible advice on where to
start tomorrow. Find an expert in your
company, record them doing a manual
workflow, draft a skill.md,
rate your three JSON evil cases, and
ship it to a read-only tier. Just start
small.
Getting that first successful skill to
fire perfectly is a lightbulb moment.
You'll never want to go back to massive
system prompts again.
It really changes everything. To wrap up
our deep dive today, I want to leave you
with a thought to mull over, circling
back to that home improvement retailer.
If agent skills effectively digitize and
version control the procedural memory,
the exact how-to knowledge of your
absolute best employees,
what happens to a company's culture and
its operations a decade from now when
its corporate subconscious consists of
hundreds of thousands of skills that
outlive any individual human worker?
That is a very big question.
Something to think about. Thanks for
joining us on this deep dive. We'll see
you next time.