This is how most AI agents get built
today. One giant framework, a sprawling
graph of moving parts. It runs until one
node deep inside throws an error and
nobody can tell where it actually broke.
And all you wanted was one call to a
model. What you got was an agent inside
an executor inside a chain. Your real
prompt buried eight layers down under
code you never wrote. So developers are
walking away. Their words, "Unstable,
the abstractions are overcomplicated,
the docs never match." Even LangChain's
own team now says it. For agents, don't
use LangChain.
Here's the secret no framework wants to
admit. Strip everything away and an
agent is just three things: a model,
some tools, and a loop. That is the
entire idea. So instead of one massive
agent that does everything, people build
the opposite. Many tiny agents, one job
each, two or three tools, a tight loop.
Micro-agents, think microservices for
intelligence. Two names keep coming up.
First, Pydantic AI from the team behind
the validation library half of Python
already runs on. Type-safe agents and 15
million downloads. Second, small agents
from Hugging Face. The whole agent
engine is about a thousand lines of code
and its agents write their actions as
plain Python instead of clunky JSON. And
that is the real unlock. A micro-agent
is small enough to read end-to-end every
line, every tool, the exact prompt. Try
doing that inside a framework. For two
years the answer to everything was, "Add
more framework." Now the pendulum is
swinging hard the other way. Less
framework, more engineering. So here is
the rule this whole video comes back to.
Use the smallest thing that works. Let's
break down why the giants got so big and
what the lightweight crowd does instead.
Start with the cost you can actually
measure. Every layer of abstraction is
code that runs on every single call.
Same task benchmarked, a direct API call
against the framework stacked on top.
The heavier the framework, the more time
and tokens you burn before the model
even thinks. But the bigger cost is the
one you can't see. The framework quietly
assembles your prompt, injects its own
instructions, and decides when to call a
tool, all behind the curtain. So when
the model does something dumb, you can't
even tell what it was actually asked,
which makes debugging miserable. The bug
isn't in your 50 lines, it's somewhere
in the framework's 50,000. So you end up
stepping through abstractions you didn't
write in a stack trace that's mostly
somebody else's code, and the ground
keeps moving under you. LangChain
rocketed past 100,000 GitHub stars and
became just as famous for breaking
changes between versions. The interface
you learned last quarter is deprecated
this one. Popularity was never the same
thing as stability. Now to be fair, the
giants grew up, the field consolidated.
LangChain pushed agents onto LangGraph.
Microsoft fused AutoGen and Semantic
Kernel into one agent framework. Crew AI
raised 18 million and now runs inside
most of the Fortune 500, and sometimes
you genuinely need that machinery.
Long-running workflows that must survive
a crash,
deep observability, state that persists
for hours, strict rules and audit
trails. If that is you, a heavy
framework earns its weight. The catch is
most projects simply aren't that. So
let's rebuild an agent from nothing.
Start with the model, give it tools it
can call, search, code, a database.
Add a little memory of what has happened
so far. That is the whole engine.
Anthropic calls it the augmented LLM,
and there's a fork most people skip
right past. A workflow runs your model
through steps that you wrote,
predictable and easy to trace. An agent
lets the model choose its own next move,
flexible but harder to control. Both are
valid. Most tasks only need the first.
This is the advice straight from
Anthropic's own guide on building
agents. Find the simplest thing that
solves the problem. Only add complexity
when it clearly pays for itself, and
don't reach for a framework when a plain
API call would do. Because most of what
people call agents are really just a
handful of simple patterns. Chain a few
prompts, route to the right one, run
some in parallel, have one coordinate
the workers, have one check another's
work, compose those, no giant graph
required.
And when you do need a true agent, the
loop is almost boring. Send the model
the goal and its tools, it picks one,
you run it, you hand back the result, it
looks again and decides keep going or
done. Round and round until the job is
finished. Now the architecture itself.
The monolith hands one model 30 tools
and a giant prompt and hopes for the
best. The micro agent approach splits
that into a team, each agent with a
narrow job, a couple of tools and a
prompt you can actually read. And the
real magic is in the constraint. Give a
model five tools and it picks the right
one almost every time. Give it 50 and it
gets confused. A small single-purpose
agent is more reliable precisely because
it is allowed to do less. So how do they
work together? A thin router reads the
request and hands it to the right
specialist. Or you wrap a whole agent as
a tool and let another agent call it.
There's no central graph to maintain,
just small pieces that plug into each
other. It keeps the context window lean,
too. A focused agent loads only what its
one job needs, not every tool, document,
and instruction at once. Less in the
window means a sharper model, lower
cost, and far fewer ways to quietly go
wrong. There's even a manifesto for
this, 12 factor agents. The themes keep
repeating. Own your prompts, own your
context window. Keep each agent small
and focused. Treat them like real
software you control, not magic you hope
behaves. So, what do you actually build
with? Start with Pydantic AI. It feels
like Fast API, but for agents. Model
agnostic, so you can swap GPT for Claude
or Gemini in a single line. Real
dependency injection and observability
through open standards, not a walled
garden. But, the headline feature is
types. You describe the shape of the
answer as a Pydantic model and the
framework forces the model's output to
match it or retries until it does. No
more parsing fragile text. You get a
clean validated object back every single
time. Then there's Small Agents,
minimalism taken to the extreme. It's
big idea, instead of emitting JSON to
call a tool, the agent writes actual
Python code and runs it. One block can
call three tools, loop, and branch where
JSON would need three slow round trips,
and the whole thing stays tiny. A
working agent is a handful of lines.
Pick a model, hand it a few tools, give
it the task. The code runs inside a
sandbox, so it stays safe. No executor,
no ceremony, just an agent that thinks
in code. And this is a whole movement
now, not two libraries. OpenAI shipped a
deliberately lightweight agents SDK.
There's Atomic Agents, Google's ADK,
Monster for TypeScript, DSPy for tuning
the prompts themselves.
Different flavors, the same instinct,
keep it thin. Look closely and they all
rhyme. Thin over thick, model agnostic,
so you're never locked in. They speak
MCP to reach their tools, they emit open
telemetry you can actually inspect, and
above all else, you own the loop, not
the framework. Now, the honest catch.
Thin does not mean free. The glue the
big frameworks handed you, retries,
state, coordination, you now write
yourself. And the moment you have many
small agents all talking, you've traded
one hard problem for a fresh one,
orchestration. And no, the giants are
not dead. For a genuinely complex,
stateful, regulated system, LangGraph
and the Microsoft Agent Framework are
doing serious work in production. This
was never framework's bad. It's match
the tool to the actual job, which brings
us right back to the rule. Start with
one agent. Reach for a framework only
when the pain is real and specific, not
because a tutorial told you to. Most of
the time, the smallest thing genuinely
does work, and the wind is at the back
of small. Standards like MCP now handle
the messy plumbing the frameworks used
to own. When the tools and the telemetry
are standardized, the framework can stay
thin and the engineering stays yours.
So, here's the whole thing in one
breath. An agent is a model, some tools,
and a loop. Keep each one small and
single-purpose. Own your prompts and
your context. And always, always use the
smallest thing that works. If this
finally made agents click, subscribe.
Cloud Code's takes apart one system like
this every single week. Build, solve,
deploy, and I'll see you in the next
one.