NetOvi: Putting Guardrails Around an AI Network Engineering Assistant
What started as ShowPilot - AI-assisted, read-only network troubleshooting - became NetOvi once a guardrailed write path was worth building: a NetBox-driven MPLS pipeline, and an AI interface that can propose config changes but cannot apply one without a literal YES.
The earlier MCP work in this journal - NetBox as inventory, Netmiko for read-only show commands - was called ShowPilot, and the name was accurate: it could see the fabric, never touch it.
This is a second, different lab, MPLS instead of VXLAN, and it’s where that boundary actually got tested and then deliberately moved.
ShowPilot became NetOvi once “read-only” stopped being the whole story: the same NetBox/Netmiko tools, plus NetBox as the source of truth for a real render/diff pipeline, plus an AI interface that can propose a config change and, after an explicit human YES, invoke the code that applies it.
That distinction matters. The AI is not the safety boundary. The tools are.
NetBox (config_context, native IPAM)
Live device (NAPALM / Netmiko)
Claude + the MCP tools
device allowlist, merge-only, one writable field (interface description), single-use token, no rollback tool
No “automate low-risk fixes” tier here, on purpose - every write, however small, goes through the same propose/confirm/apply gate. Nothing skips the human step because a field looked safe enough.
The pipeline: NetBox in, Cisco config out
NetBox (native IPAM + config_context) → GraphQL → Jinja2 → rendered config → NAPALM diff
render_config.py pulls everything a device needs - interfaces, VRFs, IP addresses, BGP/OSPF intent - in one GraphQL query, and renders it through role-specific Jinja2 templates (p.j2, pe.j2, ce.j2, internet.j2).
napalm_diff_push.py takes that candidate, loads it as a NAPALM merge candidate against the live device, and shows the diff. It never applies anything without an explicit prompt, and there is no --commit flag that bypasses that step.
The important part is that the AI sits on top of this pipeline rather than replacing it. It can ask the tools to inspect state, generate a candidate, and present the proposed change; the actual device operation remains deterministic code with its own checks.
Does the rendered config actually match the live devices?
First real NAPALM diff run found a genuine disagreement, not a rendering artifact:
+vrf definition INET
+ route-target import 65000:999
The rendered config was right. NetBox’s INET VRF object had quietly drifted from its own declared design - import_targets had become ['65000:999'] instead of the correct ['65000:100', '65000:200'].
Fixed at the source, re-verified with the same NAPALM diff: zero difference against all 11 live devices.
The lesson here is deliberately boring: don’t ask the AI to decide whether a configuration is correct when you can calculate the candidate and diff it against the device.
“All 36 devices are Active” - except one of them wasn’t
Asked the AI whether any device in the fabric was down. It checked NetBox’s status field across all 36 devices and reported back: every one Active, none Offline.
That was wrong. ARUBA-SW-LEY-01 was genuinely set to Offline in NetBox - checked directly, not inferred.
The tool behind that check, netbox_list_devices(), was returning the full raw NetBox device object for every device - not a trimmed summary. For 36 devices, each carrying its complete config_context inline, that’s a 133KB response. And that config_context wasn’t harmless: it included an enable_secret hash, a RADIUS shared key, and a user password hash, sent straight to the AI on every single device listing.
Pagination looked like the obvious suspect first - NetBox’s REST API paginates by default. It wasn’t that: count: 36, next: null, all 36 devices genuinely came back in one response. The actual problem was scale and shape, not truncation - a 133KB blob of deeply nested, mostly-irrelevant fields is exactly the kind of payload a model skims rather than reads field-by-field, and the one row that mattered got missed.
The fix was to return less, not more.
def _summarize_device(d: dict) -> dict:
return {
"id": d.get("id"),
"name": d.get("name"),
"status": (d.get("status") or {}).get("label"),
"site": (d.get("site") or {}).get("name"),
"role": (d.get("role") or {}).get("name"),
"primary_ip": (d.get("primary_ip") or {}).get("address"),
}
133KB down to 4.4KB. Re-ran the same check: ARUBA-SW-LEY-01 correctly came back Offline. Fixing the secret leak and fixing the reliability problem turned out to be the same fix, not two separate ones.
Deciding to let it write
The read-only boundary wasn’t a placeholder - it was a stated design decision, and reversing it wasn’t something to slip in quietly.
The actual bar was simple: it’s a lab, and with enough guardrails the risk is worth it for what it teaches.
The design, agreed on before writing a line of it:
- Two-step, never one-shot.
propose_config_change()only computes a NAPALM diff and returns a short-lived, single-use token. It never touches the device. A separateapply_config_change()call, with that exact token, re-verifies that the diff hasn’t gone stale and only then commits. - Merge only, never replace. Same discipline as the manual script.
- A device allowlist, checked before any device is contacted.
- No rollback tool. NAPALM’s
rollback()only works on the same connectioncommit_config()ran on - a stateless MCP call can’t hold that open. Recovery after a bad commit is a corrective propose/apply through NetBox, same as any other drift. - Everything logged outside the MCP process - a plain file on disk (
config_change_log.jsonl/netbox_write_log.jsonl), not server memory or the AI’s own account of what it did, so a restart doesn’t erase the record and nothing has to be taken on trust. Given there’s no rollback, this is what “safety net” actually means here.
Two things worth being precise about, since “two-step” undersells what’s actually happening on each side of it.
propose_config_change() stays provably read-only even when there’s a real diff to show. It renders a fresh candidate from NetBox, loads it onto the device with NAPALM’s load_merge_candidate() - which stages it in a separate candidate buffer, not the running config - asks the device itself to diff candidate against running with compare_config(), then calls discard_config() and closes the connection. Every single time, whether the diff is empty or not. The device never sees anything but a load-and-discard cycle from this call.
apply_config_change() doesn’t just trust the diff it was handed. It opens a fresh connection, loads the same candidate again, and re-runs compare_config() right before committing - re-checking reality, not memory. If that fresh diff doesn’t match what was actually shown to the human earlier, it refuses to commit and asks for a new propose_config_change() instead. That closes the gap where the device could’ve drifted in the time between “here’s the diff” and “apply it,” which would otherwise mean approving one change and silently getting a different one. The token connecting the two calls is just a correlation ID for that - a random hash, single-use, expired after five minutes - not a credential; it authorizes nothing beyond the exact diff it was issued for.
The log paid off almost immediately: one real session left a propose entry for Ethernet0/1 with no matching apply_committed after it - a proposal abandoned mid-conversation, token quietly expired. Not a bug, but the only reason that’s knowable at all instead of a guess from memory is that it’s sitting in the file:
{"event": "propose", "device": "MPLS-PE-01", "interface": "Ethernet0/1", "field": "description", "new_value": "test description", "token": "476e8...", "ts": 1788079677.2}
No apply_committed line anywhere after it. That’s the whole check - reading the file, not a special tool.
Then the same pattern got applied to NetBox itself: netbox_write_mcp.py, scoped down hard to a single low-risk field, interface description, nothing that touches routing, addressing, VRF membership, or admin state.
Both write paths, chained, on the same interface. Starting state - MPLS-PE-01’s Ethernet0/1, description MPLS-PE-01 -> MPLS-P-01:
Asked to change it via netbox_write_mcp.py, it didn’t just execute the request - it flagged what would be lost first, unprompted:
Confirmed with YES, NetBox updates - but the live device doesn’t, by design. That needs the separate NAPALM path:
Both sides now agree - NetBox:
and the device itself, sh int description run directly on MPLS-PE-01:
Reverted the same way afterward - both tools, same two-step flow, back to the real value:
The confirmation gate almost wasn’t real
The first version of the guardrail was a sentence in the server’s instructions= text: never apply without a literal YES.
Live-tested, and it failed in exactly the way you’d worry about: a bare y got the NetBox write applied anyway, while the same wording correctly blocked a bare y on the device-config path.
Same instructions, inconsistent behavior.
The fix moved the authorization check out of prose and into the function signature:
def apply_interface_change(token: str, user_confirmation: str) -> dict:
if user_confirmation.strip().lower() != "yes":
return {"error": f"Rejected: user_confirmation must be exactly 'YES', got {user_confirmation!r}."}
...
Rejected in code now, regardless of whether the model interprets y, yes, “go ahead”, or anything else as agreement.
That’s the more important architectural point: the model can request an operation, but it does not get to define what authorization means.
A related, smaller version of the same lesson: a 🔴 warning written into instructions= never actually showed up in the AI’s response. Moving it into the tool’s returned data - a field the client exposes directly - fixed that too.
Anything that must be enforced should live in code. Anything that must be communicated reliably to the model should live in structured tool output rather than relying exclusively on prompt prose.
What’s next
Everything here is MCP, Netmiko, and NAPALM. Nornir is the one major piece of this stack still untouched - a natural next lab, not an oversight in this one.
The more interesting direction, though, is continuing to push on where the AI belongs in the architecture.
The useful role isn’t “AI controls the network.” It’s closer to: AI is the interface to a set of deterministic network-engineering operations, with explicit boundaries around what those operations are allowed to do.
That’s a much more interesting thing to test.
A portal layer is the natural next piece: wrapping the propose/apply flow behind a real UI instead of a chat client, with a proper identity provider for auth instead of trusting whoever has the MCP config, and scheduled runs instead of one-off prompts. Still mid-build and not yet tested end-to-end, so no code link here yet - that’ll follow once it’s actually proven out, the same way the rest of this project’s code went up only after it worked.
Resources
- NAPALM -
inline_transferfor SCP-less lab devices, merge vs. replace candidates - Jinja2 -
keep_trailing_newline, the one Environment flag that cost an afternoon - FastMCP - in-process
Clientfor testing tools without a full MCP session - Model Context Protocol
- NetBox GraphQL API
