Welcome back everyone to day five, the
final day of the Kaggle and Google's
five-day AI agents intensive course. I'm
Smitha Colin, senior developer relations
engineer here at Google Cloud, and I'm
here one last time with Anant
Navalgaria. Anant, welcome back.
Hi everyone. Happy to host you for the
last
uh conclusive day and
uh
excited to see what you build for the
capstone coming up.
It's hard to believe we're already at
day five. Uh drop in the chat where
you're joining from today, and if you've
been with us all through the throughout
the week, drop a five in the chat. Would
love to see how many of you actually
made it through this whole intensive
course.
Wow, I'm seeing quite a lot of diversity
across the participants.
Yeah.
Awesome. So, let's actually do a quick
overview for anyone new joining us today
as a final reminder.
Uh for everyone else throughout the
week, you have been getting a lot of
things coming at you. So, you've been
getting white papers, companion
podcasts, hands-on code labs, daily live
streams, and AMAs just like this one,
and also the optional capstone project,
which we will be talking about at the
end of this live stream,
which, now that we're at day five, is a
thing that I really want to start
thinking about already. So, you can
compete for a Kaggle certificate,
badges, swag, and recognition across
Kaggle and Google's social channels.
Awesome. So, with that said, I also want
to do a quick thanks to all the people
who put this together, the Google
researchers and engineers who wrote the
white papers, and all the speakers who
joined us this week, and also the
Discord moderators who have been
answering questions nonstop for five
days straight. We genuinely could not
have done this without them.
And I would uh especially like to call
out uh the Kaggle team, also the Kaggle
marketing team, and the product
management team, who have been doing a
fantastic job in keeping the lights in
the show
uh proceeding smoothly. Thanks a lot.
Amazing.
So, let's actually get into day five.
Um this is the one which kind of ties
the whole week together.
Uh here's the arc we've been on. So, day
one, we vibe coded our first app. Day
two, we plugged it into the world with
tools and protocols.
Day three, we taught it to specialize
with skills. Day four, we made it safe
and observable.
And today, we answer the question that
you know, that's been hanging over all
of it. How do you actually take this to
production at enterprise scale without
it falling apart the first time you do
any type of changes?
So, vibe coding is incredible for
getting to the prototype in minutes, but
the same property that makes it really
fast is also exactly the same thing
which makes it fragile in production.
Because if the code is disposable,
what's the actual source of truth? If
someone needs to regenerate the system
tomorrow with a new model or a new
framework,
or even a new compliance requirement,
what do they generate from? So, the
answer that the white paper actually
proposes is something called spec-driven
development. Treat the code as
disposable, but treat the specification
as a durable artifact. So, the spec is
what gets versioned, reviewed, and
reasoned about. The code is just one
possible implementation of it. So,
Anand, why don't you walk us through the
white paper in more detail?
Thanks, Smitha. So, yes, as Smitha was
mentioning,
uh welcome back everyone today, and
we'll be looking into spec-driven
production grade development in the era
of vibe coding,
uh especially using the enterprise scale
patterns uh that we discussed in the
white paper.
So, throughout This is a small recap of
the last 4 days. So, throughout the last
4 days, we have started at day one,
where we kind of saw how the software
development life cycle is being rebuilt,
and we moved on to day two, where we
discussed established standard
connection protocols such as A2A, A2UI,
etc. Then, in day three, we looked at
packaging procedural skills to automate
some of the uh the process and make it a
lot easier. And then in uh in yesterday,
we looked at the things of security and
email aspects of it such as active red,
blue, green security teams, and much
more. Now, as we close in this white
paper towards the week on this white
paper, we must establish a vital rule.
Even though the course has the name byte
coding in it, byte coding is not byte in
production. To build software at scale,
and that's what we looked at in the
white paper as well, we must transition
to spectrum driven development. Uh as
Smita mentioned, it's the it's the part
that
should be maintained. And we showed in
the white paper how writing a rock solid
behavioral specification in Gherkin BDD
format makes code entirely disposable,
allowing agents to regenerate entire
projects accurately.
Now,
uh we also looked at how we can map out
the file hierarchy of where your rules
should live, from global project
definitions in agents.md or the markdown
file, and to specific configurations in
local, say, Gemini.md file, depending on
um which uh LLM model you're using
behind the scenes, and then down to the
task specific specifications in your
specs directory. Next, we covered the
code review bottleneck, and we mapped
out a spectrum from tier tier one
managed reviews to tier three custom
runtimes, showing how a graph database,
for example, X a graph can analyze
multi-million line code bases for impact
mapping.
Then towards the last part of the white
paper, we look we built an enterprise
policy server or we look into how we can
build an enterprise policy server that
performs structural role validation and
semantic safety checks actively blocking
tool execution if the agent attempts to
leak unmask unmask PII personally
identifiable information.
Uh to enforce this, we talked how we can
implement dynamic context resolvers
sanitizing tool arguments on the fly
with secure placeholders. Finally, we
concluded the white paper talking about
the softer aspects of
the white coding era about team culture
addressing developer burnout and
mitigating the dangers of approval
fatigue. And from my point of view,
uh not doing
something
in the industry which is called token
maxing where we focus on the the wrong
metrics than the end business outcome.
So, yeah. That's it. This is our final
blueprint for taking AI engineering
across your organization. Let's move on
to the exciting QA panel, shall we,
Smitha?
Yeah, sure. And also, thank you, Anand.
Same tip one final time. The podcast is
a really good entry point before that
you read the white paper. And honestly,
of all five white papers this week, I
think this one is one of the you know,
like it really tackles what people are
facing
right now in the industry. So, great to
start on and also really check out the
podcast. All right, with that said,
let's actually get into the QA.
And we have some amazing hosts today.
So, we have Ankur,
Antonio, Ilya,
Lee, and Omar.
Thanks for making the time everyone,
especially on the last day. Anand, over
to you for the first few questions.
Thank you, Smitha. So, given that is
last day, we have
one more question
than usual to kind of close off the
exciting week. The first question is for
you, Ankur.
For massive code bases with millions of
lines of code, why does standard
Rag fail and how do knowledge graphs
help? Also, what recommendation does
Google Cloud provide to make
long-running agents easier to build and
manage?
So, Anantya, when you look at our recent
work with Siemens, where they're
modernizing a massive industry legacy
code base that spans hundreds of
millions of lines,
you will see exactly actually what we're
recommending from an engineering
standpoint.
You cannot expect a standard monolithic
AI agent to handle a long-running
complex task. And so, if you give a
giant open-ended instruction, the
context will get muddled, it will it may
exhaust its memory, it starts
hallucinating. And it
And so, at Google Cloud, we approach
this by starting with a rock-solid data
foundation and combining it with a
highly structured decentralized agentic
architecture.
So, first, it all starts with a data
layer with the Spanner graph. And for a
long-running agent to stay on track, it
needs to understand the relationships
inside the enterprise data.
Code and documentation, they're not
flat, right? And so, they're deeply
interconnected webs of modules of
dependencies, rules, and so on.
So, with Spanner graph, we model the
entire structure using a graph query
language, and we layer in vector search
via Spanner's ANN, approximate nearest
neighbors algorithm.
This ensures that the agents aren't just
guessing based on keyword matching, but
they have a precise, highly grounded
three-dimensional blueprint of the
system that they're working on.
Now, once that data foundation is locked
in, the architectural pattern we
recommend through our
ADK,
the agent development kit and agent
platform is something we call the
slicing the elephant.
Now instead of building one massive
super agent to handle a months-long
modernization project, the ADK lets you
break that ambiguous goal into a highly
coordinated set of specialized micro
agents.
And so every step in the sequence is is
is a small, tightly scoped, and
predictable
task.
So we
as an example, we deploy a search agent
that uses span a graph to map out the
dependencies.
An architectural impact agent that could
be used to run simulations to predict
side effects before a single line of
change is made. A task breakdown agent
that could take those findings and slice
the remaining work into a context-rich,
bite-sized task.
And finally the coding agent itself that
executes just that one specific piece of
the puzzle.
By grounding your workflows in this
robust graph structure like a span a
graph and then using ADK to chunk
long-running processes into specialized
task, you protect the model's context
window and ensure enterprise grade
reliability. And crucially
because it's broken down this way, it
allows teams to also organically weave
in the human in the loop verification at
every key transition.
So like in summary, if you're looking to
scale a long-running agents in your
organization, don't build a monolith.
Build a domain aware network of micro
agents grounded by a graph database and
then use
ADK to slice the elephant.
Oh, thank you for that very very deep
answer, Ankur. Yes, and
I kind of miss the bygone era of just
putting a rag and all the micro agents,
multi agents, spec driven.
Lot has emerged over the last years, but
thank you so much for this.
No indeed, yeah.
Amazing, thank you.
All right, should we move on to the next
question?
All right, so this is for you Antonio
and Ilya.
When PR pull request volume scale
exponentially due to background coding
agents, what team workflow structures
are needed to avoid human approval
fatigue and blind merges? I think we saw
a lot in the news about millions of
codes being generated, entire OS
operating systems being generated
in a few minutes. You can imagine
there's a lot of PR, so would love to
hear your point of view on this.
Well, thanks for the question. I would
say that the the risk of fatigue is is a
real risk in the sense that you need to
adopt a number of things to avoid the
situations where your team is your team
is actually burning out.
Um, one principle that people should put
in front of themselves is that avoiding
the situation where you want to review
everything. And indeed what you want to
do is to have AI assisting you in all
the process. So, uh, in a way you should
think about the the idea of reviewing as
a risk
a way of estimating the risk, right? So,
there are
probably the first level of review is
everything should be automated. Whenever
you try to to implement you know, small
things like
fixing a typo or dependency minor
dependency in in in the code. This
should be completely automated, all
right? So, AI can help on this, can
bypass human entirely and can auto merge
when you pass CI integration. Then there
is another layer, let's say middle
medium risk layer
where essentially
these are things that you want to have a
human intervention, but you don't want
to have every single PR going to human
immediately. Perhaps you want to bundle
them so that you have a like a kind of
batch that digest every
once per day.
And then at the latest layer, the last
layer is the layer where you really want
to have human
uh reviewing things.
Uh and uh
in reality, this is what we do in
Google. We have every time you know,
like a a new CL, a new change
change in the code is actually
submitted. A new PR is submitted.
A first phase is
uh our models are actually doing the
first phase of review and everything
that can be automated is actually taken
care of of
uh
from the models themselves and then the
things where we really need to have
human inter in the loop are going to our
peers, our colleagues. I find that the
only way to uh prevent the burning out
of teams is to implement a serious set
of
of layers or or risk estimation. Perhaps
one more thing I want to say in all of
this, you really need to have a test
that is working well and
all of us we generate tests also with
AI, but you need to be in charge of what
these tests are doing and and have a
clear sense that your test is actually
not flaky and reflecting what what is in
the code.
And plus one to what Antonio said. Um I
would say the best kind of investment is
an investment towards building tools
that building that foundation so that
you can further benefit downstream from
this process. Um one approach I found
really well is actually creating those
recursive layer of adversarial review
uh so that we push the human review, we
start to
engage with the human as late as
possible for as fewer times as possible,
right? Uh so like you can imagine an
agent saying asking maybe uh the the the
the agent or the human that make the PR,
uh,
can you clarify this bit or does it
really make sense for for the final
product to have this code change? Can I
reproduce an error in case there is a
bug?
Um,
and then finally,
whenever you are making this maybe you
have a PR, maybe you're reviewing it,
maybe there is something to learn. How
do you then take all these learnings and
make sure that
all these background agents that raise
that are raising PR against your code
base can actually,
uh, improve and benefit from the
existing reviews process that you put in
place. So like some kind of a positive
flywheel that you can generate from from
the review process.
Yes,
a positive flywheel is something we
discussed in the yesterday's
the session as well and to add on to
what you're saying, adding a lot of
adversarial layers to minimize human
intervention is important, but we also
discussed adding too many layers. Make
sure you evaluate those layers that
they're adding value so that you're not
just burning tokens and you get a crazy
bill at the end of the day.
[laughter]
Yeah, and that's why doing testing
deterministic test, let's say unit test,
integration test actually can help you
reducing that bill, let's say.
And I have confidence in in, uh, yeah.
Awesome. Thank you, Antonio and Elia.
Really fantastic answer.
All right, then on to our next question.
All right, and this one is for you, Lee.
So if a robust BDD compliant
specification allows entire code bases
to be cleanly regenerated in minutes,
how does this redefine how enterprise
teams manage technical debt and legacy
migrations?
Yeah, thank you.
For the question. Let me break this down
in in four points.
First of all, like the life of a
developer like is changed nowadays. Like
we don't have that emotional attachment
to our code and files. Like before in in
the past like pre-five coding era, when
we would write code, we would spend
hours on perfecting the architecture or
we spend like many hairs on your head on
fixing that bug. And then
I don't know about you, but as a
developer then if then at the end
my code would be scrapped from what code
to production. Yeah, I would get annoyed
like I spent so much time on it.
Nowadays that's not the case anymore,
right? We we write a spec and that
generates code. That makes code
basically disposable. The the source of
truth here is now the the spec.
And um
that also means that our daily routine
changes as a developer like cuz we would
normally like spend hours of our time
like not necessarily in writing the code
but more in like reading into API
documentations or doing like syntax
debugging or figuring out like well I
tried to do it the way how the API
describes it and it doesn't work. I
think about it like how much time did
you actually spend hands on keyboard?
That's that's different now. You should
write a spec once. Yeah, and it and it
yeah, generates all the code right away
for you. And that means for legacy
applications, yeah, we no longer need to
write manually line by line and do this
translation nightmare. You write the
spec once. If I decide to have it
generated in Python, I can do so. If
tomorrow I decide to have it generated
in Java, yeah, that's fine, too. Like it
can happen. We can do that quite quick.
It does comes with a shadow side like
producing thousands lines of code before
lunch time. Yeah, that doesn't mean that
that is what
I mean
that means that you're bringing the
bottle back to the downstream of the
integration because right now like you
need have like cultural changes and and
discussions within your organization
because yeah the way how you refuel the
code might
and you might focus much more on testing
and writing these specs
than ever before.
Amazing. Yeah, definitely. The part
which I really relate to is the
emotional attachment to your code.
Um
uh
Yeah, we are all developers, right?
[laughter]
Uh especially when when you name even
the the function names and the modular
decomposition. Now it can all be
regenerated. So yeah, we are definitely
in a very different era. Thanks a lot
for your detailed answer. We really
appreciate it.
Awesome. Then let's move on to our
question. So this one is for you, Omar.
Uh so are current open-weight models
capable of running complex multi-agent
workflows locally or do we still need
proprietary models to act as the central
orchestrator?
Yeah, that that's a it's a great timely
question. Just 2 months ago we released
Gemma 4.
And when we release open models, our
focus is to release models that people
can actually run in their hardware,
right? Like they can run in their own
Pixel phones, for example. They can run
the models in their laptops, in their
local gaming computers.
At the moment the local models are quite
good for simple agentic tasks.
Some multi turn function calling. Of
course, the context of local models is
not as large. So if you want to operate
over a very large code base, if you want
to work over very complex agentic
scenarios, most likely you will want to
use
the most capable models, which would be
Gemini, for example, for doing very
complex agentic stuff.
Another setup that is interesting is
also hybrid inference. That means that
you may want to use Gemini as the
orchestrator that is calling sometimes
smaller Gemini models or maybe a local
models where something can be fulfilled
directly on device.
So, it's an interesting space. It's
evolving very quickly, but we have been
quite amazed by what we can do. You can
use ADK with Gemma. You can
use Gemma also via Vertex. You can run
Gemma in your own phone. So, yeah, many
exciting areas.
Very cool. And I would like to make a
couple of points there on the
the the the local version of Gemma,
which is released, which I forgot the
exact name, but the one which can run on
your local devices, even your local PC.
Even the multi-modal version, that was a
real game-changer. I think it kind of
transitions a very capable multi-modal
model running literally in a regular
everyday PC. So, really great work on
that
from your team.
And another point,
a lot of what you mentioned about hybrid
inference.
I'm thinking doesn't it remind you of
federated learning, where you used to
have one big model central and used to
kind of give way to the like kind of
combine.
I'm wondering if a paradigm like
federated
inference and LLMs is coming through.
Yeah, yeah, it's a it's an interesting
space and there are many different ways
to approach this, right? So, I just
mentioned that you can use Gemini as the
main orchestrator that can sometimes
call Gemma. You can also do it the other
way around. That means that you can have
a local model on device that determines
when the the the user prompt should be
fulfilled on device or when it should be
sent to a server-side model. And when
it's sent to a server-side model, which
one do you want to use, right? So, for
example, if you want to ask a model,
"Why is the sky blue?" most likely you
don't need the most intelligent model
out there. But if you want to operate
over a code base that has
a million lines of code, lots of very
complex repositories, It's you will want
to call a for an orchestrator agent that
then can help determine how to separate
the task and do like this more complex
agentic task. But definitely it's a very
exciting moment. Uh
the model is multimodal, relatively long
context for on-device. It can understand
images, videos, audio for the smaller
checkpoints. And again, it's designed to
be a model that is friendly towards uh
on-device developer uh devices, not uh
not super expensive hardware, right? So
pretty much if you have a phone, uh if
you have a laptop, you may be able to
run one of these uh different size Gemma
models.
Amazing. Wow.
Thanks for sharing that, Omar. Uh
yeah, let's move on to our
uh next question.
The community
Thank you, Anand. Really great questions
and context. Uh let's actually move on
to the community questions. So the first
one is from Savio, and they're asking,
"How do we prevent formal specifications
from inheriting the very technical debt
and bloat they are actually meant to
replace? And as we shift towards
spec-first workflows, what are the
emerging patterns for spec
modularization? And can AI agents
eventually handle spec refactoring to
catch logical contradictions before any
code is actually generated? Uh leave you
on to take this."
Sure. So thanks, Savio. Uh it is a fair
concern. Um I typically prevent this
debt by treating my documentation uh
modeler like my code.
Um yeah, it depends, of course, but if
it's like a fresh new project, uh
typically how I structure this is I
would would start with a technical
design. We all work with the technical
design. And typically I convert this to
markdown, and then I store it somewhere
in my code base, probably in the spec
folder. And uh this sensor serves as a
master guide. Uh
so the coding agent will understand the
overall high architecture level like the
architecture and how everything is
connected and how everything flows.
And then from there I start like
building like the the specific features
down into the smaller specs like the
Gerkin style uh behavior driven specs.
Grouped uh by the how it reflects the
the software structure.
And then I uh typically would hyperlink
uh to these smaller files from the
master doc.
I also link to all the
uh
configurations and and
and data schemas. So typically like when
I use natural language that's for for
specs, everything that's related to
behavior, everything that's related to a
contract uh like the inner output from
an API or uh database schemas or JSON
objects. That's what I typically write
either in JSON or in YAML or or maybe
whatever the language the the database
uses.
I link to that separately. And then to
tie it all together
what I typically do is I would create
like an overarching uh Gemini uh dot
markdown file. And in that markdown file
I write like this system prompt which is
basically my super prompt that tells my
agent uh that whenever I generate code
it will always need to update my specs.
It will also always need to generate new
tests because the more tests I have like
the more uh
I'm not I feel better like assured that
that whatever I generate is is right.
And then uh I typically also let it
update like my change log and my read me
files. Uh yeah. So So that's how I would
typically do it.
I I really like that you're also telling
an agent you're having an agent which is
actually updating the spec as well as
you make uh changes to the code and
test. That's awesome.
Uh let's move on to the next question,
also that from Savio. Congrats, Savio,
for getting two of your questions
answered. Um so,
if spec-driven zero-trust development
essentially positions the developer as a
technical architect of agent workflows,
rather than a coder, how do we prevent
the human review bottleneck from
becoming a source of cognitive atrophy,
uh where engineers lose the deep system
intuition required to validate
agent-generated blueprints in the first
place? Um Ankur and Omar, do you want to
take this?
Yeah.
Omar, do you want to start off?
Yeah. Yeah, I think it's a interesting
question. I think all of us are learning
the industry is evolving very quickly.
So, there are many paradigms and
patterns that are changing. Uh so, it
will be very interesting to also see how
things look in a year from now.
Uh one way to prevent cognitive atrophy
is to shift the developers' focus from
reviewing uh every single line of code
to more on auditing the testing
footprints and how the system what the
system is doing, the logical assertions
of the system, right? So, if you need to
read 2,000 lines of code uh that were
generated in a couple of minutes, and
you then you need to uh spot every
single edge case, that would usually be
very error-prone. Uh like, humans will
commit errors if you need to review
5,000 lines of code, right? So, instead,
I think we need to shift to review
behavioral tests, right? Uh whether the
agent generated the right things to
prove its worth. So, if you are able to
review and approve assertions a bit more
like test-driven development in some
ways, you can validate that the system's
contract is accurate, not necessarily
all of the syntax within the system.
Uh I do think, I'm going back to my
first point, the developer ecosystem is
going to evolve a lot in the next couple
of months and next years. The developer
tools are going to evolve to show more
semantic diffs. So, rather than seeing
every single line of code that was
changed, maybe this change modified this
data prevention policy or maybe this
change uh
caused this effect, right? Rather than
raw code diffs. Uh and in that way the
engineers will still be very engaged
with the system uh design level. Uh but
yeah, it's an interesting lesson and I
think all of us as an industry are
learning.
Yeah, and I would say that look, in
addition to the the the how that Omar
talked about, like I think the why is
important to understand as well. Like
keeping this knowledge of the underlying
system, I believe is actually quite
imperative. Whether you think about
fundamental issues around security,
around system reliability, scalability,
I do think it is important to have an
increased intuitive understanding of
what is the system itself that you're
working on. And
and another good thing about exactly
what Omar was saying, like with with
with things like semantic diffs which
will probably be coming up
in these tools more
more broadly, I think understanding how
your PRs are having a change in the
underlying system, tracking those, as
well as uh of course like going back to
your standard practices around having
the spec docs, the design docs, PRDs,
like these are things that uh we would
still we should still maintain so that
we can keep on explaining what the
underlying system is both to ourselves
but also to the LLMs.
Uh thank you both. Omar, you also
mentioned that the developer ecosystem
is going to be changing really fast. Uh
do both of you have any tips for junior
engineers who are actually coming into a
world where they rarely have to write
code anymore? and how do they kind of
get that experience to become, you know,
as we say technical architects of agent
workflows?
Yeah, from my point of view, uh one of
my recommendations for everyone in the
space at the moment is build, build
things, build in the open, and build
collaboratively.
So, the more you can collaborate with
others, the more you contribute to open
source, the more you get engaged with
others people systems, with other
companies projects, the more you can
understand how different systems work,
how your agents, or how you can
grow in an agentic way, in a way that
you can effectively collaborate across
the broader ecosystem. And I think just
this part of building in the open,
building with others, building
collaboratively,
uh allows you to grow this
muscle in which you understand how
agents work, which are their
limitations, what they excel at,
uh and of course as the agentic systems
evolve,
uh
the scope and the capabilities of what
you can do will also grow, but you're
growing in a agent agent way, but still
with that deep understanding on how
systems work.
Yeah, and I would say that there is uh
no sort of way around building because
that's because these systems are
evolving in these underlying
technologies are evolving so fast. So,
whatever you learn right now, and
whatever you have
used to build today,
uh will probably not
uh still work or will be obsolete in 6
months. And so, yeah, you need just need
to keep on being on
uh being on top of these.
And and in the spirit of building, what
you guys mentioned, our capstone
project, where we you get to build
collaboratively with others, is coming
right up, uh and a fun challenge after
the capstone project as well. So, stay
tuned, everyone.
Awesome. Okay, let's go on to our next
community question from Izadiar.
Um now that Google is restructuring
production readiness and expanding
automation in Google workspace. Do we
ever return to using A2A or another
protocol or approach to coordinate
software development with other
departments for organizational
automation workflows? Antonio and Elia,
you want to answer this?
I'll start on this.
So, absolutely.
I would say it depends on whether you
want to consume some something
pre-packaged like in a workspace or you
want to build, right? With this about
building here, you want to build like an
agent that allows you to do things.
I would say A2A is super relevant when
it comes to interaction between agent to
agent communication, right? So, and
especially when those communication span
across different teams, different
departments for which each team is
responsible for a subset of those
agents. Let's take for example the case
where you maybe want to
you want to do like a PR code reviewer
agent, yeah? And of course, this is not
something that our workspace automation
can do automatically.
And I would say in this specific case,
definitely like building on top of A2A
allows you to have way more flexibility.
In such such situation when you have
like a PR code reviewer, maybe you want
to have as part of it
like a compliance checker, yeah? And you
happen to have a team that specifically
does an agent for compliance checking.
So, how do you
allow that your agent which is a PR code
review to interact with this compliance
check agent? You do it through A2A. So,
you will actually establish a network
communication so that the your PR
reviewer agent is capable of consuming
the compliance check and all the best
practices, all the domain knowledge that
the compliance check brings in and so
that you can then finalize a result. In
this case here, you will be responsible
for maintaining the PR reviewer agent
versus like the different team that will
provide to you just the compliance check
expertise and
and task. So, definitely still super
relevant.
Well, plus what would you say that is I
think that one way of thinking at this
is really depending on the type of use
cases you have. Like workspace perhaps
is very suitable whenever you have a no
code situation because the ecosystem
workspace is very much into no code.
But then when you talk about agents,
there are also situations where you have
design patterns, right? So, agents can
be also considered in terms of design
patterns. And whenever you talk about
communication between agents, A2A is
is a very good choice.
Um two examples just to clarify.
Before we were discussing about a
situation where an agent is actually
routing requests to multiple models and
you have a trade-off between cost and
quality and performance.
That's a situation where you essentially
want to use A2A because it's very simple
to implement that particular
design pattern with A2A. Another
situation is perhaps when you have a
number of agents that are working in
parallel and so the the requests are
dispatched to these agents.
Another situation where A2A is really
very very useful. I wrote a book about
this where you talk about design design
patterns. Every time there is a
communication between agents
and and the type of communication is
non-trivial, I would say that A2A is a
very good choice.
Absolutely. Especially because sometime
like it's a protocol, so at the end of
of the day it's like a language that we
are establishing between two agents and
like you don't want to reinvent the
wheel. You don't want to reinvent
another way of communicating between two
agents, so you can rely on top of A2A to
leverage a set of
built-in best practices.
Absolutely.
Amazing. Okay, let's head on to our last
and final question from Sammy.
In large-scale multi-agent development,
how is architectural consistency
maintained across agents working on
different parts of the same codebase?
Omar and Antonio, do you want to take
this?
Yeah, I think it's a very relevant
question as well.
It I think the key is separating the
architecture planning from the
execution. So, when you build your
Yandex system, you want to have an agent
maybe that is in charge of, yeah, what
is going to happen, designing the plan,
and then that plan may be uh offloaded
or delegated to other agents, right? So,
uh rather than letting every single
coding agent to decide where to write
files, how to write them, uh this
central architect agent, let's say, will
own the whole dependency graph of the
codebase. Uh
so, when a coding agent needs to make a
change, it may first submit like a
structural plan. The architect may say,
"Okay, uh
this goes against or this does work
well." So, it will analyze and
understand and then a scaffold for the
change and then hand it over to the
coding agent to actually execute. Uh
and then you can also uh sandbox the
coding agent so it doesn't have access
to everything necessarily. So, it will
just fill in the blanks within an
approved uh
pre-approved structure uh
in such a way that it will not suddenly
cause the architecture or the patterns
of your codebase to just drift in
different directions.
Well, plus one on what you said, Omar.
Uh Ankur, before you were saying that uh
we are all learning what is working
today probably in 6 months will be
evolving. One thing that I'm uh
I'm doing at this point is uh
I believe on this idea of architect, but
for me it's more like a squad of of
agents that are playing this role. So,
there is an agent that is in charge and
co-editing the spec as a
uh Lee, you were saying.
Uh there is an agent that is actually
co-editing the test
um the test suite. And
again, Lee, you were saying you increase
the confidence on this.
Uh and then I have agents that are not
only writing the code, writing the test,
and write the documentation. But this
thing is still in this in this squad, I
have also agents that are managing the
infrastructure. So, the world deploy
at this point is after having the
studied the artifact. So, code, test,
documentation, the world deploy is
actually run by other agents
that are managing the infrastructure.
Today, I don't use uh
to go to GCP directly or through the C-
CLI. Agents are doing this on my behalf.
And then two more thing that I'm I'm I'm
doing more and more is once the code is
deployed, there are agents that are
going there and collecting all the
trajectories, all the logs.
And And the final thing is people are
talking about this a lot. I'm seeing a
significant a huge gain adopting this
this patterns is once the code is
deployed, the logs are are collected,
the trajectories are collected, there is
a final agent which is sitting on the
top of everything,
kind of super architect, and this final
agent is is closing the loop. So,
running the experiments, run learning
from these experiments, and then going
back and causing changes into the
original uh spec. So, there is this
continuous loop where you start from the
spec, you add test, you add
documentation, you deploy, you have uh
logs that are and trajectories that are
collected, and then the final step is
the self-improving loop.
And on the top of this, I
our role is to collect all this
information and validate this
information. So there are multiple
checkpoints where humans can can have a
can have a role. This the initial spec,
the
the test verification, the overall log,
and then the final loop where things are
improving in a self self self-improving
loop.
Cool.
Amazing. Thanks for your answers
everyone and all the guests. Thanks for
joining. That was our final question, I
believe, right, Smitha?
Yeah, thank you to all the guest
speakers for coming on here to answer
these questions. And a huge thank you to
everyone who's been following along so
far. So the depth of these questions has
been genuinely the best part of these
live streams. With that said, now let's
actually move on to the code labs with
Lovie.
Thanks Smitha. Thanks everyone. This was
a great
discussion and question and answer
session. So let me quickly
go to the code labs. Um
All right, let's start. Okay, so the two
code labs that you'll do today, again
it's in continuation of everything that
you have done so far.
This is one of the first code lab which
where we'll sort of talk about how to
deploy an ADK agent using agent CLI.
And this is written by one of my
colleague. Amazing job with both the
code labs that you'll see today.
So what we're going to do in this is
we're going to create an agent and that
agent essentially will create something
called as ambient expense agent. This is
just a sample that we're trying to do.
The two main objective of the code labs
today is one to create an agent and
deploy that, and the other is to sort of
add a layer of UI over it so that you
can also sort of use that in a general
environment. So you don't have to always
talk to a chatbot and there's a UI layer
to it. So the first code lab sort of
focuses on building that agent and then
sort of deploying it to agent runtime.
So the first thing that we'll do is
again you've done this multiple times.
I'm not going to bore you with this, but
you have to make sure that you set up
your cloud environment and agency line
anti-gravity can sort of do this
things with you and this is where it
it'll also sort of use lot of G cloud
automatically. So, make sure that you do
this. There are some issues that happens
with project ID. So, make sure that if
you're running into those issues, you go
to all the things that you've done
before the same thing will apply here.
Again, the third step is very simple
which is you set up the agency line
which is one of the most important
things. Do remember that agency line has
two sort of major components inside it
which is one it has all these
scaffolding that you need to work with
the whole agent platform
which is to say if you want to sort of
deploy certain things, if you want to do
any interaction with the GCP
from an agent platform perspective,
agency line is really good to do that
from an automation perspective and then
one of the amazing thing it also has is
the ADK skills which means that it has
all these skills which are required for
any coding agent be it anti-gravity
cloud code or code X or anything that
you want to use.
It'll it'll use those skills to build
those agents, deploy those agents and do
a lot of hand-holding.
It'll take lot of hand-holding from you
back to the agent. So, the second step
is for you to sort of deploy
second step for you is to sort of set
set up the agency line ADK skills.
Once that is done, what you're going to
go through is you're going to set up
your first
sort of scaffolding of the agent. This
is where the prompt if you read the
prompt, what it specifically says is
that you know, build me a
ambient expense agent
that sort of streamlines the expense
reporting by
you know, by employees and one of the
things that you're doing here is we're
explicitly asking it to
align with ADK 2.0 where you have the
graph workflows. And this use case
specifically, if you go through it,
you'll realize that it's a very
graph-based use case because there is a
uh, there's a sense of workflow that
needs to be integrated. There's a there
are steps that has to be taken if you
take one path versus the other. So, for
example, if you notice, we are saying,
"Hey, do the auto approve if the
expenses are less than $100.
If not, if they are more than send it to
review agents so that the human in the
loop can sort of review and approve the
uh, expense. Again, very real world,
this is pretty much a lot of the expense
trackers and expense systems do. So,
that's what we're doing. We're
explicitly asking it to align and build
for ADK 2.0.
Uh, it doesn't mean that you always have
to do it, but some of the use cases are
good uh, if you sort of uh, do the graph
part.
Now, once that is done, uh, one of the
coolest thing which I this is this is my
favorite thing about Agent CLI which is
um,
you can sort of even before you do the
deployment, you can sort of start doing
the scaffolding of the deployment. And
as well as well as do the dry run. What
what it does is it saves you all the
effort of, you know, going anything that
goes wrong during that process. So, the
next step that you're going to do here
is you're going to scaffold the
production deployment files for Agent
Runtime which is to say that it'll
create the Agent Runtime app. It'll
create the deployment metadata. These
are important for it to
um, you know, at some point go to the
Agent Runtime and deploy itself. So,
that's the second that's the fifth step.
And what you do next is like I said, the
the most important and the very
interesting part of this is that you
would ask it to specifically do a dry
run before it sort of goes and does the
actual implementation uh, or the
deployment of it. Now, this is important
also because again, when you are using
coding agents, if there is something
that is wrong and if there's anything
that is sort of um, not aligning with
how the deployment should go, the coding
agent will try to fix this for you and
you don't have to sort of worry too much
about it. And this is where the agent
CLI skills are very helpful because it
has all of these important requirements
written in the skill. So, again, a very
important step, even if you do this
outside of the lab, whenever you're
doing anything related to agent
development, you would want to sort of
do dry run all the time to make sure
that you're not pushing anything which
is not supposed to be pushed.
The next step is as simple as you know,
just saying deploy this agent to agent
runtime. What it does in the back is
like I said, agent CLI has two
components, so it has the scaffolding of
deployment. So, it will actually
leverage the agent CLI deploy. It will
use their project ID and the regions.
Now, sometimes and and this happens
um
sometimes which is the fact that you may
have issues with cloud region.
Um so, try using some other regions if
you're facing any issue. Uh if not, you
can sort of just um you know, directly
let it take care of this. Uh the other
thing that you can do is you can also
run this async, which is to um
explicitly tell uh Antigravity that, you
know, use no wait uh because then you
don't have to wait for it to, you know,
keep going and it will keep a check on
when this is done.
Now, once the deployment happens and do
remember that deployment sort of takes a
bit of time. It's anywhere from three
four three to five minutes to sometimes
it may take 10 minutes for uh different
reasons, but you would have to wait for
this time before you can actually test
the agent.
Now, testing the agent is also very
interesting because there are multiple
ways you can test this agent and this is
what step eight will help you figure
out, which is A, you can actually write
a simple prompt to Antigravity and say,
"Hey, just test my agent uh test my
agent on the agent runtime." And you
give it a scenarios and different test
cases so that it can uh go through all
these things uh if required. Um the
other thing that you can do is it
actually has um a UI in the agent
runtime. So, if you go to the cloud
console and see the agent runtime, you
have a playground there. In that
playground, you can actually put all of
these
different inputs and we've given you
examples of these inputs. Feel free to
sort of change it
and do the stress testing of your agent
just to make sure that everything is
good or not, right?
And at this point, if there are
something that fails and if there are
anything that is not working as you're
expecting,
you should
you know, let the Antigravity know and
it'll sort of fix that for you. So, you
have both of these options.
Now, a couple of additional things that
you can do with this lab. So, by this
point, you are sort of main objective of
building an agent and deploying is
almost done. But, there are a couple of
things that you can additionally do,
which is you can see the cloud trace and
the cloud log of an agent. This is very
important from monitor and the
observability part of agents. You can
also sort of write different SQL queries
and see what your traces of the agent
execution is. And this is very important
when you're sort of scaling it because
sometimes things go wrong and you want
to know why they went wrong, right? So,
there's another very good thing that you
can do with agent CLI in that
perspective.
The last thing, which is agent registry,
again, this is very interesting because
now, imagine a scenario where you're a
big company, you have like thousands of
agents, you would want your agent to be
somewhere centralized
and available as a registry so that you
can have a look at it and you can
use it if if your workflow sort of
requires it. So, you can test whether
your agents are automatically added in
the agent registry. By default, when you
do deploy, it will do that on behalf of
you. So, you can just test that part.
So, that's lab one and this is where
you'll see, like I said, starting an
agent, building the scaffolding,
deploying it,
testing it, and then monitoring and
observability part of it. The next one,
let me switch to the next lab.
is the fun part. So, the first part was
sort of slightly cumbersome because you
were doing a lot of serious things. But,
the next part, again, this is a lab by
Sita, amazing lab that she has written,
is is the fact that now you'll put a UI
in front of it. And this is like the fun
part because you don't have to always
sort of do the chat thing. You can
actually just create a UI, and that UI
actually coordinates everything with the
agent. Now, one thing which is very
interesting that you should spend time
is understanding why we are doing pub
sub here because we want these streams
of input to come to your agent runtime.
And then it should do all the remaining
logic at as it should. So, the way it
starts by is again, you should reconnect
with Antigravity and confirm deployment.
Again, very important because your UI
sort of relies on that deployment needs
to be done in the previous lab. So, make
sure that you're testing this. Once that
is done, it's as simple as just writing
a simple vibe code command. Now, one
thing to be very important one thing to
keep in mind here is that while you can
define however you want the UI to look
like, that's again your personal choice
and how you want to. But very important
is the instructions that the Antigravity
should remember in terms of how it
should connect to the
to the back end of the agent runtime.
And this is where you'll see that we're
doing fast API server. We're explicitly
mentioning that, "Hey, these are the
things that you need to keep in mind.
Connect to the session service." The
session service is where all the
connections and all the integrations
that will happen. And whatever you do on
the UI, it'll actually do all of those
things in the back. So, this sort of
prompt gives you a lot of
specific details so that you know your
Antigravity doesn't go in different
directions, and it sort of does this in
one go. Once that is done,
again, like I said, this is as simple as
this. Once this is done, you basically
do another deployment. So, this is not
the same deployment as the previous
case, which was the agent. This is where
you're deploying your actual UI. So,
again, we explicitly ask it to do a
Cloud Run deployment so that this is
separate from your agent runtime. And
then you define these two sort of
properties, which is the cloud project
and the runtime ID. Both are important
for it to function. And then you just
let it take care of that. Um the part
where I want all of you to focus on
today is the Pub/Sub, which is A, you
need to create that Pub/Sub, and the And
again, this is one of those architecture
choices. You don't have to do it. You
can do it in a very simple way as well.
But because we're trying to build it
from a scale perspective, the Pub/Sub
here sort of takes care of all the
incoming messages. You can imagine like
thousands of people in your company will
actually use this expense report. So,
you want something which can take care
of the streaming. You also need to have
something where if for some reason the
message fails to
go in the stream, there has to be
something where it can still catch it.
So, the five and the six is essentially
where you're building the Pub/Sub, again
just by prompting it. Once that is done,
you'll be able to sort of create and
connect this to the agent runtime as we
walked through in the first one.
Now, just to give you again a high-level
view of what this architecture looks.
So, there's a event ingestion part of
it. There's a management front-end part
of it, which is what we built. And then
there's a agent runtime part of it,
which is where all the agent things are
happening. So, these are three different
components, and we're trying to sort of
club them together by the push
subscription, which is going to push all
the events from the Pub/Sub to the
agent. And then there's a post API
section, which is coordinating with the
UI. So, this is very important. This
sort of gives you a high-level view of
what we're trying to do. And that's it.
That's That's what these two labs will
help you do, and I'm sure that you'll
have a lot of fun. So, with that, I'm
going to be back with Smita and Anant.
Awesome. Thank you so much, Ravi, for
going through both of those code labs.
That was great. Now, let's move on to
the pop quiz.
Amazing. Hope you're all ready for the
final pop quiz.
And first question for your pop quiz is
why is code considered disposable in
spec
driven development?
Is it because A,
AI generated code is always poor quality
and must be rewritten. B, because having
a rock solid version control spec
allows the entire code base to be easily
regenerated or translated. Or C, because
modern applications do not require a
persistent code base. Or D, because
database structures are decoupled from
the application logic. Think back to the
white paper and at least discussion
earlier. And the correct answer will be
shown in 3, 2, 1. And it's B.
Uh having a rock solid version control
spec allows the code base to be
generated as you wish and there's no
emotional attachment to it as Lee
pointed out.
All right, moving on to the second
question. Which structured natural
language syntax is recommended for
behavior driven specifications to keep
LLMs focused on state, action, outcome?
Is it A, YAML schemas? Or B, JSON schema
templates? Or C, Gherkin? Or D, UML
state diagrams?
Think about from our discussion earlier
and the white paper. And the correct
answer will be shown in 3,
2, 1.
And it's C. We discussed about Gherkin
for the spec specs earlier and that's
what it is.
Moving on to our third question.
Uh
what is the purpose of the cross tool
configuration file agents.markdown or
.md?
Is it
A, to define the specific API keys for
Google Cloud? Or B, to serve as a shared
cross tool foundation to prevent
instructional fragmentation?
Or C, to override product level sandbox
settings? Or D, to list the installed
third-party libraries?
Think about that And what could are my
Usually it's for especially agent to MD.
Correct answer is
Uh it should be shown in three, two,
one.
And it's B. It says that the shared
cross tool foundation um to prevent
structural fragmentation, especially
when there are multiple agents
operating.
All right, moving on to our fourth
question.
On very large legacy code bases, how do
diet three custom review runtimes
achieve deep structural code
understanding? Is it
um through option A, by loading the
entire code base into a single context
window, or B, by building a knowledge
craft by combining graph QL and vector
search and full text search, or C, by
ignoring the past and only can focusing
on the current pull requests, or D, by
running simple regex pattern scanners?
Think about that, and the correct answer
will be shown in three, two, one. And
the correct answer is B, by building
knowledge craft and vector search and
full text search. We're using by using
all of them. Yeah.
All right, then our last question. How
does a hybrid policy server evaluate
tool calling actions before execution?
Is it
A?
A It runs unit tests on the tools code
base.
B It prompts the developer for
multi-factor tool multi-factor
authentication.
Or C, by doing deterministic
structural gating via configurations and
specialized semantic LLMs. Or D, it's
best completely blocks all the dynamic
tool calls.
Think about it, and the correct answer
will be shown in three, two, one. And
it's C.
All right, that brings us to the end of
our pop quiz and we have a special
guest.
Hello folks. Uh it is a pleasure to see
you. Um I'm Renda from the Kaggle team.
If you've been on the Discord, we've
probably chatted because I've been
helping to moderate that and I'm here to
talk about the thing that everyone is so
excited about, which is the Capstone
project. The Capstone project is your
path towards getting to the badge and
certificate that people are really
pumped about with the um with the
course.
Uh so let's take everything that you've
learned over these last 5 days and put
it into practice. In this final project,
you're going to be called upon to build
an AI agent using all the tools and
concepts you've been learning. We've got
four categories that you can submit in.
The first one is agents for good. Uh
we'll be looking at submissions that
help solve problems for humanity, from
optimizing agriculture to managing
public health, advancing education,
supporting arts and literature. This is
your track for helping people.
The second one is agents for business.
Enterprises are using AI agents to solve
critical problems, from managing expense
submissions, my personal like favorite
thing to have AI do instead of me, to
creating pipeline actions, driving
insights, creating new products. In this
track, you'll create an agent designed
to solve compelling business problems
with cost or revenue on the line.
The third one is a concierge agent. The
opportunity for personal aid AI agents
to streamline and simplify people's
lives is incredible. From managing the
invite list for a party to planning a
garden or helping manage complicated
medications, safe and secure agents can
free time for things that really matter.
In this track, you'll solve individual,
family, or social challenges in a way
that keeps personal information private
and secure.
The last one is open-ended. It's a
freestyle track where you can exercise
your creativity. Are you tracking
satellite launches? Are you part of a
fandom that has problem? Uh whatever
your creativity can bring to the table
with the freestyle track, uh we'd love
to see.
We've covered a lot of concepts and
technologies through the white papers,
code labs, and QA sessions in the live
stream. We're really looking to see how
you incorporate these ideas in your
final project.
So, a few things to keep in mind about
the capstone. Each participant or team
can submit only one track, so choose
carefully. Although, we reserve the
right to adjust projects after we
evaluate them and be like, "Wait, this
this belongs to the other track."
You may work individually or in a team
of up to four people. It's completely up
to you.
The deadline for submitting your
projects is July 6th at midnight in
Pacific time, but we really strongly
recommend that you submit early in case
there are any issues or technical
challenges that come up. Once the
capstone is closed, it's closed, and
your chance to get a badge or
certificate is over.
Finally, the top three winning teams in
each category will receive Kaggle swag
and recognition on our social media
channels.
Every participant in the capstone will
get a Kaggle badge and certificate. And
this is only available to people who
participate in the capstone now.
Capstone link is on its way shortly. So,
check your email, keep your eyes on the
Discord server. We'll also be sharing a
link to our post-course survey. We'd
love to hear from you about what you
learned and what your experience was
with this course. Let us know your
thoughts.
Thank you so much. Back to Smitha and
Anant.
Thank you so much, Brenda. This has been
amazing. Do check out the capstone
project that Brenda was mentioning. And
that's a wrap on the five days of AI
agents live stream. This has been
amazing. Thank you to everyone who has
been sticking with us for the entire
five days. The engagement this week in
the chat, in Discord, and all the
questions has been incredible. Anant and
I have genuinely loved doing this with
all of you.
Yes, and thanks a lot to everyone. And
as Brenda mentioned,
the capstone and there'll be a little
bit after the capstone, a small thing as
well, which uh so to to keep uh you
engaged and experiencing your using your
skills in real life.
All right. Thanks a lot, everyone, for
joining.
Thank you.