AI-Lab v2.5 Active: Head-to-Head Multi-File Production Benchmarking & Telemetry Receipts
Portfolio Exhibition • Multi-Stage Telemetry
•
26 Recorded Runs Across 8 Tasks
Empirical comparison of unconstrained single-thread agents against governed hierarchical sub-agents (@backend-core [Max], @frontend-ui [Medium], @qa-playwright [Low]).
Featured Benchmark
EXP-008: ResumeForge Interview Intelligence System
Multi-turn STAR question synthesis, 4-axis rubric evaluation, and evidence claim grounding.
PLAYWRIGHT 100% PASS
Cost Reduction: -64.54%
ARM A: Vanilla Monolith
$0.10955 • 10 turns • 64.2s
ARM B: Governed Sub-Agents (ICM)
$0.03884 (-64.5%) • 4 turns • 25.8s
Cost / Run
$0.03884
-64.54% vs Monolith
Duration
25.80s
-59.81% latency
Prompt Cache Read
88,400 tokens
10.8x higher reuse
Thinking Rationing
1,820 tokens
-1,030 tokens saved
Chaos & Telemetry
EXP-007: Dynamic API Gateway Telemetry Simulator
-63.05% Cost Reduction
Landing Page
EXP-005: SaaS Marketing Landing Page
-77.9% Cost Reduction