# Design Document: AI-Orchestrated Mobile Core Network Control Plane

Status: Draft v0.4
Audience: Core network architecture review board
Scope: Introduction of LLM-based agent orchestration (MCP, A2A) into the 5G mobile core control plane

## Executive Summary

This document proposes an architecture in which a chain of LLM-based agents assists human operators in reasoning about, evaluating, and optimizing the mobile core network's control-plane configuration. Rather than a single monolithic model attempting end-to-end reasoning over the entire network state, the proposal splits the work across three cooperating agents of decreasing size and increasing deployment locality: a large "Architect" agent that ingests raw design documents and structures them, a mid-sized "Protocol Engineer" agent that evaluates inter-agent communication protocol choices, and a small "Edge Analyst" agent that produces deployment-ready latency guidance for edge sites. The three agents never share a single model instance or a KV cache; they share only serialized intermediate artifacts, which is the central engineering question this document works through.

## Mobile Core Network Overview

The 5G mobile core network's control plane is composed of a set of network functions -- AMF (Access and Mobility Management Function), SMF (Session Management Function), UPF (User Plane Function), and PCF (Policy Control Function) among others -- that exchange signaling messages over the service-based interface (SBI). Historically, operators have relied on static rule engines and manually authored policy tables to decide how these functions route and prioritize traffic. As network slices multiply and edge deployments proliferate, the combinatorial space of routing and signaling decisions has outgrown what a static rule engine can reasonably encode, motivating the introduction of LLM-based reasoning agents that can be pointed at a natural-language description of a proposed configuration and asked to evaluate it before it is rolled out.

### Network Function Routing and Signaling Touchpoints

The following touchpoints enumerate, per network function, where the proposed agent orchestration layer would need to observe or influence existing routing and signaling behavior. Each touchpoint is written as a standalone item so the Protocol Engineer agent can evaluate them independently rather than needing to hold the entire network function inventory in context at once.

N-1. AMF registration routing: the AMF selects a serving AMF instance during initial registration based on Requested NSSAI and TAI; any agent-mediated advisory must not alter this selection path, only annotate it for later review.
N-2. AMF mobility signaling: N2 handover signaling between source and target AMF must complete within existing 3GPP timers regardless of whether an agent is concurrently evaluating the handover's routing metadata.
N-3. AMF-to-SMF signaling: the Nsmf_PDUSession_CreateSMContext service operation is a routing decision point the agent orchestration layer must treat as read-only telemetry, never as an interception point.
N-4. SMF session routing: SMF selects a UPF based on DNN, S-NSSAI, and UE location; the Protocol Engineer agent's evaluation role is to check that any proposed agent-visible routing metadata does not leak session-identifying information into a lower-trust log tier.
N-5. SMF PDU session signaling: N4 signaling between SMF and UPF establishes the actual forwarding rules; this is precisely the signaling class the orchestration layer's A2A hand-offs must never be mistaken for on a shared message bus.
N-6. UPF traffic routing: UPF forwarding decisions happen in the user plane and are explicitly out of scope for agent-mediated evaluation, since none of the three agents in this pipeline operate on user-plane packets.
N-7. UPF usage reporting signaling: N4 usage reports feed charging and policy decisions; if an agent orchestration layer is ever extended to reason about these, the volume of signaling traffic this generates at scale must be modeled first.
N-8. PCF policy routing: PCF's policy decisions determine QoS routing for a given session; any LLM-based advisory role here needs its recommendations logged with a routing-decision provenance trail distinct from the automated PCF decision path.
N-9. PCF-AMF signaling: the Namf_EventExposure interface is a signaling channel the agent orchestration layer could plausibly subscribe to for read-only awareness, but the design explicitly excludes this from the current phase.
N-10. NRF-mediated service discovery: every SBI producer/consumer pair discovers each other through the NRF; the orchestration layer's own agent-to-agent discovery mechanism is deliberately kept separate from NRF to avoid conflating agent hand-offs with genuine 3GPP service discovery.
N-11. Cross-slice routing isolation: network slicing requires strict routing isolation between slices; an agent evaluating one slice's configuration must never be handed token arrays derived from another slice's raw configuration text.
N-12. Roaming signaling: inter-PLMN signaling via the SEPP introduces a trust boundary the orchestration layer does not cross in this design -- all three agents run entirely within the home PLMN's compute environment.
N-13. Emergency session routing: emergency PDU sessions bypass normal policy routing entirely; the Protocol Engineer agent's evaluation logic must explicitly exclude emergency-tagged sections of any design document from routing-optimization recommendations.
N-14. Lawful intercept signaling: LI-related signaling paths are excluded from this entire agent orchestration design by policy, not merely by omission, and any future extension proposal must re-affirm that exclusion explicitly.
N-15. Charging signaling: Nchf_ConvergedCharging signaling volume is high enough that any agent subscribing to it for evaluation purposes would need its own rate-limiting, independent of the rate-limiting already applied to the charging function itself.
N-16. AMF-initiated deregistration routing: network-initiated deregistration must route through the same AMF instance that owns the UE's current registration context, and any agent-visible telemetry about this path must be tagged as terminal, since no further routing decision follows it.
N-17. SMF session release signaling: the N4 session release procedure must complete before SMF reports session termination upstream; an agent evaluating post-hoc session logs must not assume termination is instantaneous with the release request.
N-18. UPF anchor relocation routing: when a UE moves far enough that its anchor UPF changes, the resulting routing update is one of the more disruptive events in the control plane and warrants its own dedicated evaluation category distinct from ordinary mobility handover.
N-19. PCF policy update signaling: an Npcf_SMPolicyControl_UpdateNotify signaling exchange can be triggered independently of any UE mobility event, purely from a change in subscriber policy, and an agent's routing evaluation must not conflate the two triggers.
N-20. AMF-UDM signaling: subscriber data retrieval from UDM during registration is a signaling exchange the orchestration layer must never be positioned to delay, since UDM lookups sit on the critical path of every initial registration.
N-21. NEF-exposed signaling: the Network Exposure Function's northbound APIs represent a signaling surface already exposed to third parties; any agent-generated routing advisory referencing NEF-exposed data must be held to the same third-party data-handling standard as NEF itself.
N-22. AF-influenced routing: application function requests for traffic influence (via PCF) constitute a routing input that originates outside the core network entirely, and the orchestration layer's evaluation logic must clearly label such inputs as externally sourced.
N-23. SMF-UPF N4 heartbeat signaling: periodic N4 association heartbeats are low-value signaling from an evaluation standpoint and should be explicitly excluded from any agent-visible signaling volume metric to avoid diluting more meaningful signals.
N-24. AMF paging signaling: paging messages sent to a UE in idle mode are broadcast-like in character and must never be represented, in any agent-visible summary, as though they were a targeted routing decision.
N-25. Network slice admission control signaling: NSSF's slice selection signaling determines which AMF set serves a UE for a given requested slice, and this is a routing decision point the Protocol Engineer agent's evaluation logic must treat with the same care as UPF selection.
N-26. AMF-to-AMF signaling during inter-AMF mobility: when a UE moves between AMF sets, the source and target AMFs exchange UE context directly, and this exchange must be modeled as a distinct signaling category from the N2 handover signaling described in touchpoint N-2.
N-27. SMF-initiated network-triggered service request signaling: SMF-triggered service requests for a UE in idle mode traverse a different signaling path than UE-triggered requests, and the routing evaluation logic must not treat the two paths as interchangeable.
N-28. UPF idle-mode buffering behavior: while not itself a signaling exchange, UPF's decision to buffer or discard downlink packets during paging is a routing-adjacent behavior that any future extension of this evaluation pipeline into the user plane would need to account for explicitly.
N-29. PCF-initiated policy re-evaluation signaling: a change in subscriber location can trigger PCF to re-evaluate an active session's policy without any explicit UE-initiated signaling, and this asynchronous trigger must be represented distinctly in any routing evaluation timeline.
N-30. AMF congestion control signaling: NAS-level congestion control back-off timers are themselves a signaling mechanism the orchestration layer must never interfere with, since they exist specifically to protect the core network during overload.
N-31. UDM-initiated subscriber data update signaling: UDM can push subscriber data changes to AMF and SMF asynchronously, and an agent evaluating routing decisions against a snapshot of subscriber data must record the snapshot's staleness bound explicitly.
N-32. AMF-to-NEF signaling for monitoring events: event-monitoring subscriptions between AMF and NEF represent a signaling class that, unlike ordinary session signaling, is explicitly intended for third-party consumption, and must be governed by stricter data-minimization rules.
N-33. SMF-PCF session policy association signaling: the initial policy association at session establishment is itself a routing-relevant signaling exchange, since the policy it negotiates directly constrains the UPF selection made moments later in the same session establishment procedure.

## Agent Orchestration Layer: MCP and A2A Protocols

The orchestration layer connecting the three agents borrows two protocol concepts from the broader agentic-AI ecosystem: MCP (Model Context Protocol), used here to standardize how each agent receives structured context about the network state it is reasoning over, and A2A (Agent-to-Agent protocol), used to standardize how one agent hands off a completed unit of work to the next agent in the chain, effectively defining the routing path a unit of work takes through the pipeline. In this design, A2A hand-offs are not implemented as live RPC calls between resident agent processes; instead, each agent runs to completion, writes its output as a self-contained artifact, and exits. The next agent in the chain discovers that artifact, reads it, and proceeds. This "write, exit, discover" pattern -- and the routing and signaling assumptions it makes about how artifacts are discovered -- is what allows the three agents to run sequentially on a single GPU without ever needing to hold more than one model's weights in device memory at a time.

### MCP Context Object Field Considerations

The MCP context object is the structured container each agent receives describing the piece of network state (or, in this pipeline's concrete instantiation, the piece of design-document text) it is being asked to reason about. The following considerations were raised during architecture review and are recorded here for the Protocol Engineer agent to evaluate against the current field schema.

C-1. Every MCP context object must carry a stable identifier that survives serialization to and from shared memory, so a routing decision can always be traced back to the exact context that produced it.
C-2. The context object's provenance chain must record which agent produced it and at which pipeline stage, mirroring the block_id/source_agent/stage fields already present in this pipeline's OKF metadata schema.
C-3. Context objects must never embed raw credentials or session tokens belonging to a live network function, even for illustrative or test purposes, since a context object may be logged for audit.
C-4. A context object's routing metadata (which downstream agent should consume it) must be derivable from tags attached at creation time, not inferred later from the object's content, to keep the routing decision auditable.
C-5. Context objects that reference routing or signaling constraints specifically must be distinguishable, at the schema level, from context objects carrying only general architectural narrative, so a downstream Protocol Engineer agent can filter cheaply.
C-6. The size of a context object's payload should be bounded, since an unbounded MCP context object risks exceeding a downstream agent's effective context window before any agent-specific instruction is even appended.
C-7. Context objects must specify which tokenizer produced any pre-tokenized payload they carry, since this pipeline's entire optimization strategy depends on that field being both present and independently verifiable.
C-8. A context object that fails schema validation on read should cause the consuming agent to fail loudly rather than silently proceed with partial or default field values.
C-9. Context object timestamps should be recorded in a single canonical timezone across every agent, to keep multi-run debugging tractable when comparing artifacts produced hours apart.
C-10. Context objects produced by a crashed agent mid-write should never be mistaken for a complete, valid hand-off by the next agent in the chain.
C-11. The relationship between one context object and the object it was derived from (its "source block") should be explicit and queryable, not reconstructable only by string-matching filenames.
C-12. Context objects intended purely for human review, versus those intended for direct machine consumption by the next agent, should be visually and structurally distinguishable in the workspace directory.
C-13. A context object's tag vocabulary should be a closed, documented set rather than an open-ended free-text field, so that filtering logic in a downstream agent cannot silently miss a semantically-equivalent but differently-spelled tag.
C-14. Context objects should degrade gracefully if an optional field is missing, but must fail hard if a load-bearing field -- such as the pointer to a pre-tokenized payload -- is missing.
C-15. The MCP context schema should be versioned explicitly, so a future schema change does not silently break an agent that was written against an earlier version of the schema.

### A2A Hand-off Semantics Considerations

The A2A protocol's hand-off semantics govern exactly how one agent's completed work becomes the next agent's input. The following considerations were raised specifically about the "write, exit, discover" pattern used in this pipeline.

H-1. A hand-off artifact must be written completely, in one atomic operation from the perspective of any reader, so a downstream agent scanning the workspace directory can never observe a half-written file.
H-2. The discovery mechanism a downstream agent uses to find upstream hand-offs must be deterministic across runs, so re-running the same pipeline against the same input produces the same processing order every time.
H-3. A hand-off artifact's tags are the only mechanism a downstream agent should use to decide relevance; content-sniffing the artifact's body to guess relevance is explicitly discouraged, since it reintroduces the CPU cost this whole design exists to avoid.
H-4. Hand-off artifacts should never be mutated in place after being written; if a correction is needed, a new artifact with a new identifier should be written instead, preserving the original for audit.
H-5. A downstream agent must never assume it is the only consumer of a given hand-off artifact; the same artifact could, in principle, be relevant to more than one downstream agent's filter.
H-6. The hand-off mechanism must remain functionally correct even if the upstream and downstream agents are versions of the pipeline built weeks apart, provided the underlying OKF schema version is compatible.
H-7. A hand-off artifact's absence should be distinguishable, in the downstream agent's error reporting, from a hand-off artifact that exists but fails validation -- these are different failure modes with different remediation paths.
H-8. The A2A hand-off should carry enough provenance that an operator auditing a final report can walk the entire chain backward to the original raw input text without consulting any external log.
H-9. Hand-off artifacts should be cleaned up according to an explicit, documented retention policy, rather than accumulating indefinitely in the workspace directory across repeated pipeline runs.
H-10. The hand-off protocol must not assume the downstream agent starts immediately after the upstream agent finishes; an arbitrary delay between hand-off and pickup must not corrupt or expire the artifact.
H-11. A2A hand-offs in this design intentionally avoid live RPC specifically so that no agent's failure can block another agent that has already received its hand-off and is proceeding independently.
H-12. The hand-off artifact format should be human-readable without any special tooling, so an operator can inspect a stuck pipeline's state with a plain text editor during an incident.
H-13. Every hand-off should be idempotent to re-processing: running the downstream agent twice against the same upstream artifact should not silently double-count or duplicate work.
H-14. The hand-off mechanism's dependency on shared memory rather than persistent disk must be documented clearly enough that an operator understands hand-off artifacts do not survive a host reboot.
H-15. A hand-off artifact's declared tokenizer identity must be independently verifiable by the downstream agent before that agent trusts any pre-tokenized payload the artifact references.

### Deployment Topology Considerations for the Orchestration Layer

Because the three agents never run concurrently within this pipeline, the deployment topology question is less about load-balancing across replicas and more about sequencing, isolation, and resource hand-back between one agent's process lifetime and the next.

D-1. Each agent must run as its own operating-system process, not as an in-process function call within a shared Python interpreter, so that the CUDA context and every megabyte of VRAM it holds is released deterministically when that process exits.
D-2. The orchestrator responsible for sequencing the three agent processes must treat a non-zero exit code from any agent as a hard failure, halting the pipeline rather than attempting to proceed with a partial hand-off chain.
D-3. Shared-memory cleanup must happen in a `finally`-equivalent block at the orchestrator level, so a mid-pipeline crash never leaves stale token arrays occupying tmpfs capacity indefinitely.
D-4. The orchestrator must verify, before launching any agent process, that the tmpfs-backed shared-memory region has enough free capacity for the largest token array the upcoming stage is expected to produce.
D-5. Agent process launch order must be fixed and documented, since the entire hand-off chain depends on each agent finding its upstream artifacts already written when it starts scanning the workspace directory.
D-6. The orchestrator should log wall-clock duration per agent stage separately, since a regression in one agent's latency should not be masked by averaging it against the other two stages' durations.
D-7. A deployment topology that colocates all three agents on a single GPU is the baseline assumption of this design; any topology that instead spreads agents across multiple hosts must re-derive the shared-memory hand-off mechanism entirely, since tmpfs does not span hosts.
D-8. The orchestrator must not assume any particular agent's model weights are already resident in host page cache; cold-start disk I/O time for model loading should be accounted for separately from the token-injection latency this pipeline is optimizing.
D-9. Resource limits (VRAM, tmpfs capacity, process count) should be enforced at the orchestrator level, not left to each agent to discover only after it has already failed partway through its work.
D-10. The orchestrator's precondition checks -- including the tokenizer-equivalence verification described elsewhere in this design -- must run once, before any agent process is launched, rather than being re-checked redundantly inside every agent.
D-11. A deployment topology intended for production should treat the workspace directory (where OKF hand-off artifacts are written) as ephemeral scratch space, not as a durable store of record for any decision the pipeline makes.
D-12. The orchestrator must be able to re-run a single failed stage without re-running the stages that already completed successfully, provided their hand-off artifacts are still present and valid.
D-13. Any monitoring layered on top of this deployment topology should treat "agent process running" and "agent process making progress" as distinct health signals, since a hung agent still counts as a running process.
D-14. The orchestrator's own memory footprint must remain small relative to any of the three agents' model weights, so that orchestration logic never becomes the reason a deployment fails to fit its VRAM budget.
D-15. A deployment topology that eventually moves the smallest agent to a far-edge site must still satisfy the same tokenizer-equivalence guarantee this design enforces for the single-host case, or the token-injection optimization silently stops being valid.

## Routing and Signaling Integration Points

The Protocol Engineer agent's primary responsibility is evaluating routing and signaling integration points between the orchestration layer and the existing SBI-based signaling fabric. Specifically, it must assess: (1) whether a proposed A2A hand-off could be mistaken by the SMF's session routing logic for a legitimate N4 signaling message if the two are ever bridged onto the same message bus, (2) whether the routing metadata attached to each MCP context object is sufficient for an operator to trace a decision back to the specific signaling exchange that produced it, and (3) whether introducing agent-mediated routing decisions into the control plane creates a new signaling amplification risk during a mass registration event, of the kind operators have historically had to guard against with mobility-management signaling storms. These are precisely the routing and signaling concerns that must be routed to the Protocol Engineer agent rather than handled by the Architect or Edge Analyst agents, since they require protocol-level judgment the other two agents are not sized or prompted for.

### Detailed Routing Risk Register

R-1. A proposed routing advisory that references a UPF selection must never be interpreted as an instruction to actually change UPF selection, only as a recommendation surfaced to a human operator.
R-2. Routing metadata generated by an agent must carry a confidence indicator distinct from the deterministic routing decisions made by SMF's own selection logic, so the two are never conflated in an audit log.
R-3. Any agent-generated routing recommendation touching cross-slice traffic must be flagged for mandatory human review before it influences any configuration change, regardless of the agent's confidence indicator.
R-4. The routing evaluation pipeline must not introduce a new single point of failure into a path that is currently resilient to individual network function restarts.
R-5. Routing recommendations must be timestamped with both generation time and the network state snapshot time they were evaluated against, since network state can change between the two.
R-6. A routing recommendation that contradicts an existing, already-deployed policy must be surfaced as a conflict requiring explicit resolution, never silently discarded or silently applied.
R-7. The routing evaluation logic must treat roaming subscribers' routing paths as a distinct category from home-network subscribers, since the two have materially different signaling paths through the SEPP.
R-8. Any routing advisory touching emergency-session paths must be excluded entirely from the agent evaluation pipeline, as noted in the Mobile Core Network Overview's touchpoint N-13.
R-9. Routing recommendations must not reference internal IP addressing of network functions directly, since that addressing is considered sensitive topology information under the operator's existing data classification policy.
R-10. A routing evaluation that spans multiple network slices must document, per slice, why cross-slice comparison was necessary, since slice isolation is a hard requirement elsewhere in this design.
R-11. The routing risk register itself must be versioned alongside the design document it accompanies, so a stale risk register is never mistaken for a current one during review.
R-12. Routing recommendations generated during a maintenance window must be tagged distinctly from those generated during normal operation, since baseline traffic patterns differ significantly between the two.
R-13. A routing recommendation's rationale must be expressed in terms an operator without LLM-specific expertise can evaluate, not merely as an opaque confidence score.
R-14. Routing evaluation must account for the fact that UPF selection can change mid-session during a service and session continuity mode transition, and a stale routing recommendation must be invalidated when that happens.
R-15. Any routing recommendation that would increase inter-datacenter backhaul traffic must be flagged explicitly, since backhaul capacity is a distinct constraint from radio or compute capacity.
R-16. The routing evaluation pipeline must degrade gracefully -- producing no recommendation rather than a low-confidence one -- when the network state snapshot it was given is incomplete.
R-17. Agent-generated routing recommendations must never be automatically merged into a live configuration management system without a distinct, logged human approval step.
R-18. A routing recommendation touching a network slice reserved for a specific enterprise customer must carry that customer context explicitly, never inferred implicitly from slice identifiers alone.
R-19. The routing evaluation logic must be re-run, not silently reused, whenever the underlying design document it was evaluating is edited, since a stale evaluation against an edited document is worse than no evaluation.
R-20. Routing recommendations must document their assumed traffic model explicitly, since a recommendation valid under steady-state traffic may be actively harmful during a flash-crowd event.
R-21. Any routing advisory referencing a specific AMF or SMF instance by name must be reviewed for whether that reference will still be valid after the next planned network function scale-out.
R-22. The routing evaluation pipeline's own compute footprint must not compete for GPU resources with any latency-critical inference workload already running on the same physical node.
R-23. A routing recommendation must specify whether it applies to uplink traffic, downlink traffic, or both, since asymmetric routing considerations are common in mobile core deployments.
R-24. Routing evaluations that reference third-party interconnect partners must be reviewed against the operator's existing interconnect agreements before being surfaced to an external-facing dashboard.
R-25. The routing risk register must include an explicit "not applicable" disposition for touchpoints the current design phase does not address, rather than leaving them unaddressed by omission.
R-26. A routing recommendation must never be generated from a design-document block whose declared tokenizer identity could not be verified against the rest of the pipeline, since an unverified block could silently corrupt the very token IDs the recommendation was reasoned over.
R-27. The routing evaluation pipeline must record, per recommendation, which of the three agents produced it, so a reviewer can weigh a small edge-deployed agent's recommendation differently from a large centrally-hosted agent's recommendation if warranted.
R-28. Routing recommendations touching network functions scheduled for imminent decommissioning must be suppressed automatically, rather than relying on a human reviewer to remember which functions are already on a retirement timeline.
R-29. A routing recommendation must never be phrased as an imperative instruction ("change UPF selection to X"); it must always be phrased as an observation with a suggested action, preserving the human operator's authority over the actual change.
R-30. The routing evaluation pipeline's output format must remain stable across agent model upgrades, so that downstream tooling consuming these recommendations does not need to change every time a larger or newer model is substituted into the pipeline.
R-31. Routing recommendations must be excluded from any dataset used to further train or fine-tune a future version of the agents in this pipeline, to avoid a feedback loop where the model's own prior recommendations bias its future ones.
R-32. A routing recommendation generated against a design document later found to contain factual errors must be retroactively marked invalid, not silently left in place alongside recommendations generated against corrected versions.
R-33. The routing evaluation pipeline must expose, to a human reviewer, the exact block of source text a given recommendation was reasoned over, not merely a summary or paraphrase of that text.
R-34. Routing recommendations that would only make sense under a future, not-yet-deployed core network release must be labeled with that dependency explicitly, so they are not mistaken for actionable guidance against the current deployment.
R-35. The routing risk register's disposition for a given touchpoint must be revisited whenever the touchpoint's underlying 3GPP specification is updated to a new release, since a routing assumption valid under one release may not hold under the next.
R-36. A routing recommendation must never be silently re-ordered ahead of an already-queued human-authored change in the same configuration management pipeline, since the ordering of changes can itself affect the outcome of both.
R-37. The routing evaluation pipeline must expose a dry-run mode that produces recommendations without writing any hand-off artifact, for use during initial rollout when reviewers want to sample the pipeline's output before trusting its artifacts downstream.
R-38. Routing recommendations generated against a design document written in a language other than the operator's primary operations language must be flagged for translation review before being surfaced to on-call staff.
R-39. The routing evaluation pipeline must not be granted write access to any live network function's configuration store under any circumstance in this design phase; its output is advisory text only.
R-40. A routing recommendation's underlying reasoning, not merely its conclusion, should be retrievable on demand, since a reviewer rejecting a recommendation needs to know which specific input consideration the agent weighed most heavily.

### Detailed Signaling Risk Register

S-1. Signaling messages generated by the agent orchestration layer must be namespaced distinctly from 3GPP-defined signaling messages at the transport layer, precisely to prevent the message-bus confusion named in this section's opening paragraph.
S-2. A signaling amplification risk assessment must model the worst case of every agent in the pipeline being triggered simultaneously by a single mass registration event, not merely the average case.
S-3. Signaling exchange metadata must record which specific MCP context object triggered it, satisfying the traceability requirement raised in touchpoint (2) of this section's opening paragraph.
S-4. Any new signaling path introduced by the orchestration layer must be load-tested against the operator's documented signaling storm scenarios before being enabled in a production-adjacent environment.
S-5. Signaling messages related to agent hand-offs must never traverse the same physical message bus as N4, N2, or N1 signaling without an explicit, reviewed bridging design, per this section's primary concern.
S-6. A signaling exchange's retry policy must be bounded and must not itself become a source of amplification during a partial outage of a downstream agent.
S-7. Signaling volume generated per evaluated design-document block must scale sub-linearly with document size, or the orchestration layer risks becoming the amplification source it was meant to help prevent.
S-8. Every signaling exchange in the agent orchestration layer must have an explicit timeout distinct from any 3GPP-mandated timer, so a stuck agent cannot indefinitely hold up a hand-off.
S-9. Signaling metadata must never be used as a substitute for the actual 3GPP signaling audit trail already required by regulation; it is supplementary, not a replacement.
S-10. A signaling exchange that fails validation must be logged with enough context to reproduce the failure offline, without requiring access to the live network state at the time of failure.
S-11. The signaling risk register must explicitly address what happens if the tmpfs-backed shared memory region used for hand-offs is exhausted mid-exchange, rather than assuming it always has capacity.
S-12. Signaling exchanges that reference a specific subscriber's session must be redacted before being included in any design document intended for broad architecture-review distribution.
S-13. A signaling amplification test must include the scenario where all three agents in the pipeline are restarted simultaneously after a host-level failure, not only the steady-state running scenario.
S-14. Signaling exchange rate limits must be configurable per deployment, since a far-edge site's acceptable signaling budget differs materially from a centralized data center's.
S-15. The signaling risk register must be reviewed any time the underlying agent orchestration protocol version changes, since a protocol change can silently alter signaling volume or pattern.
S-16. Signaling exchanges must carry a schema version field so a future protocol change can be rolled out without breaking agents still running an older version during a phased upgrade.
S-17. A signaling exchange that spans the boundary between two network slices must be treated as a cross-slice event requiring the same isolation review as any other cross-slice routing decision.
S-18. Signaling metadata volume must be bounded independently of the size of the design document being evaluated, to prevent a single unusually large document from degrading signaling performance network-wide.
S-19. The signaling risk register must document the expected steady-state signaling rate per agent, so a deviation from that baseline can be detected and investigated promptly.
S-20. Signaling exchanges related to agent hand-offs must be excluded from any charging-relevant signaling counters, since they do not represent subscriber-billable network activity.
S-21. A signaling storm caused by a misconfigured agent retry loop must be distinguishable, in monitoring dashboards, from a genuine mass registration event affecting real subscribers.
S-22. The signaling risk register must specify who is paged when signaling amplification thresholds are crossed, since an unowned alert is equivalent to no alert.
S-23. Signaling exchange logs must be retained long enough to support a post-incident review, but no longer than the operator's data retention policy permits.
S-24. Any signaling path shared between the agent orchestration layer and an existing OAM (Operations, Administration, and Maintenance) system must be reviewed for whether it could introduce a new denial-of-service surface.
S-25. The signaling risk register's acceptance criteria must be re-verified after every material change to the pipeline's token-serialization mechanism, since that mechanism is what this entire signaling risk assessment is ultimately protecting.
S-26. A signaling exchange that references a network slice reserved for critical communications (such as public safety) must be subject to a stricter validation path than an exchange referencing a best-effort consumer slice.
S-27. The signaling risk register must document whether agent-orchestration signaling is included or excluded from the operator's existing network-wide signaling capacity planning models, since an undocumented omission there would understate true signaling load.
S-28. Signaling exchanges that fail their schema-version check must be quarantined for manual inspection rather than being silently dropped, since a dropped exchange could mask a real upstream protocol mismatch.
S-29. A signaling amplification incident caused by the orchestration layer must trigger the same incident-response runbook as a signaling amplification incident caused by any other network function, rather than a separate, less-rehearsed procedure.
S-30. The signaling risk register must specify a maximum acceptable signaling exchange rate per agent per second, calibrated against the smallest deployment target (the far-edge site), not against the largest.
S-31. Signaling exchange payloads must never include the full text of a design document; only the minimal routing/signaling-relevant metadata needed for traceability should be carried in the signaling layer itself.
S-32. A signaling path shared between two agents running on different physical hosts, should that topology ever be adopted, must be authenticated end-to-end, unlike the current single-host design where process boundaries provide sufficient isolation.
S-33. The signaling risk register must be reviewed by both the core network security team and the platform team responsible for the orchestration layer's compute infrastructure, since neither team alone owns the full risk surface.
S-34. Signaling exchange counters used for this risk register must be reset on a documented schedule, so that a long-running deployment's counters do not silently overflow or lose precision.
S-35. Any change to the signaling risk register itself must be reviewed with the same rigor as a change to the underlying routing risk register, since the two registers are two views of the same overall risk.
S-36. A signaling exchange between two agent stages that both happen to run on the same physical GPU must still be logged as a distinct event, even though no network transport is actually involved, to keep the audit trail consistent with a future multi-host deployment.
S-37. The signaling risk register must document what happens to in-flight signaling exchanges if the orchestrator itself is restarted mid-pipeline, since the orchestrator's own availability is a dependency the register has otherwise not addressed.
S-38. Signaling exchange rate metrics collected from this pipeline should feed the same observability stack the core network's own signaling metrics already feed, rather than a separate, agent-specific dashboard that reviewers must remember to check independently.
S-39. A false positive in the semantic-fidelity check that gates whether a generated evaluation is trusted must be treated as a signaling-integrity concern, not merely a model-quality concern, since a corrupted hand-off is indistinguishable from a coherent one at the transport layer.
S-40. The signaling risk register must be re-baselined whenever the document corpus being evaluated changes significantly in size or structure, since both routing and signaling overhead in this design scale with the size of the design-document blocks being processed.

## Edge Deployment and Latency Considerations

Once a protocol design has been evaluated, the remaining open question is deployment: which of the three agents, if any, can realistically run at a far-edge site alongside the UPF, where compute is constrained and the round-trip budget to a centralized inference cluster may not fit within the control loop's latency target. The Edge Analyst agent is deliberately the smallest of the three so that it is the one candidate for far-edge co-location, and its report should quantify, for a representative far-edge hardware profile, how much of the end-to-end decision latency is attributable to model inference versus to the token-serialization hand-off mechanism itself -- the same tokenization-bypass mechanism this whole pipeline exists to validate.

## Security and Compliance Notes

Any mechanism that shares raw token identifiers between processes, even on the same physical host, must be evaluated against the operator's data-handling policy before deployment, since a token-ID array is a lossless encoding of whatever text produced it. The current design confines this exchange to a tmpfs-backed shared-memory region that is never written to persistent disk and is explicitly cleared at the end of every pipeline run, which keeps the exposure window bounded to the lifetime of a single evaluation run and avoids leaving evaluation artifacts, which may reference real proposed network configurations, resident on disk after the run completes.

## Appendix: Glossary

AMF: Access and Mobility Management Function. SMF: Session Management Function. UPF: User Plane Function. PCF: Policy Control Function. SBI: Service-Based Interface. NRF: Network Repository Function, used for service discovery between network functions. SEPP: Security Edge Protection Proxy, the trust boundary at the inter-PLMN roaming interconnect. NAS: Non-Access Stratum protocol layer between the UE and the core network. NGAP: NG Application Protocol, the N2 interface protocol between RAN and AMF. MCP: Model Context Protocol, used here for structured context hand-off between agents. A2A: Agent-to-Agent protocol, used here for artifact-based hand-off between sequential pipeline stages.
