Your AI Factory Is Paid in Tokens. Defend Them.
Overview
Busy isn't the same as productive.
Running an AI factory means juggling GPUs, fabric, storage, and power, all to hit one number: tokens delivered per dollar per watt. When something slows down, finding the cause today takes an expert, five dashboards, and an afternoon.
AI NOC spans the whole stack. Ask a question in plain language, and it finds the responsible layer, evidence attached. Built for NeoClouds, service providers, sovereign clouds, and large enterprises running AI factories and AI grids.
It runs on the Aviz Autonomous Agentic Platform, the same private, vendor-neutral foundation behind Network Copilot, now extended to the GPU. Early access today, in development with NVIDIA, demonstrated in the context of NVIDIA DSX Air.
One question. Every layer.
Decompose
Your question splits into per-domain sub-queries, inference, GPU, compute, orchestration, fabric.
Collect
Agents pull from vLLM/Dynamo, DCGM, Base Command Manager, Run:ai, and Aviz ONES, in parallel.
Correlate
Signals join on the node or tenant, layers get ruled out until the cause is named.
Answer
KPI cards, charts, evidence. Read-first, nothing autonomous, nothing unattended.

Run the factory by asking
- "Give me a summary of all warning and critical alerts in the past 24 hours."
- "Generate a health report for the last 24 hours."
- "Show me the GPU utilization for the last hour."
- "Create tenant Blue, allocate all GPUs from hgx-node-00, schedule project 'noc-training'."
- "Show me all workloads for Tenant Blue."
- "Which jobs failed in the last 24 hours?"
The span no one else covers
| Others | AI NOC |
|---|---|
| GPU dashboards or network monitoring, never correlated | One agentic interface across network, compute, storage, and power |
| "GPU is 90% busy" | MFU, KV-cache economics, and a bottleneck verdict, is it actually productive? |
| Vendor-locked to one silicon stack or one cloud | Vendor-neutral and private, on your infrastructure |
| Static tooling and fixed dashboards | Ask in plain language; extend it with your own agents via the Aviz Agent SDK |
Cross-stack correlation across inference, GPU, compute, and fabric is demonstrated today; storage and power breadth are on the roadmap, in development with NVIDIA.
Three jobs. One interface.
Demonstrated end-to-end in Aviz's AI Factory lab and in the context of NVIDIA DSX Air; offered through early access.
Tokenomics
Fleet-wide, per-server, per-model, the economics your business runs on: tokens per dollar per watt, not GPU percent-busy.
What it measures
- Latency: TTFT, inter-token latency, end-to-end p50/p90/p95/p99 vs. your SLA baselines.
- Throughput: generation and prompt tokens/sec, requests/sec, goodput.
- Efficiency: Model FLOPs Utilization (achieved vs. peak) with a bottleneck verdict, compute-bound, memory-bandwidth-bound, decode-limited, or idle.
- KV-cache economics: cache-hit rate, turnover, allocation vs. true memory waste.
- Scheduler analysis: max-token overshoot, preemptions, queue depth.
- SLA readiness: a priority-ordered optimization plan vs. production targets.
Fleet Telemetry & Root Cause
Live per-GPU and fleet-wide via NVIDIA DCGM. Flags idle and stranded GPUs, allocated capacity producing nothing, so you can reclaim before buying more.
What it covers
- Telemetry: utilization, VRAM used/free, power draw, temperature, XID errors, row-remap failures, PCIe replays.
- Cross-layer root cause: joins inference (vLLM), GPU (DCGM), compute (BCM), and fabric (ONES) on one node or endpoint, names the cause with evidence attached.
- Network observability: historical trends, NetFlow/sFlow/IPFIX flow analysis, syslog analysis, all conversational.
Tenant Ops & Dashboards
Create a tenant, allocate GPUs, launch the workload, tear it down, one conversation, persona-scoped access throughout.
What it covers
- Tenant-aware ops: list, allocate/de-allocate GPU hosts, assign VNI/VLAN segments, organization-scoped isolation.
- Self-service lifecycle: integrates with orchestration like NVIDIA Run:ai, projects, quotas, workload submission, platform health.
- Dashboards & background jobs: any check schedules as a recurring job, ask once, it runs every morning.
Private by design
The same platform controls that carry Network Copilot into regulated networks carry AI NOC into sovereign AI environments.
Your infrastructure, your model
Private LLM on your own GPUs, or the endpoint you already run. Telemetry never leaves, never trains someone else's model.
Tenant isolation
Organization-scoped separation across fabric, compute, and inference.
Persona-scoped access
Infra admin, tenant admin, user, each sees and does only what their role allows.
Read-first operations
Observes, correlates, recommends. No unattended changes; governed action gates on the roadmap.
Sovereign-ready
On-prem and air-gapped paths, RBAC, and audit visibility for your control frameworks.
In Development with NVIDIA
Demonstrated to NVIDIA against AI Factory operations use cases, exercised in the context of NVIDIA DSX Air, the flight simulator for AI Factory operations. Complements the NVIDIA AI-factory stack (Dynamo, DCGM, Base Command Manager, Run:ai) as the agentic layer that correlates across it.
Early access, priced on value
AI NOC isn't tiered or metered yet, it's available through an early-access program, with licensing scoped per engagement as the product moves toward general availability.
What early access includes
Direct access to the Aviz AI Factory team, your use cases shaping the roadmap, and licensing scoped per engagement, not metered per GPU, per token, per flow, or per query.
Common questions
The most-asked questions from early-access conversations with AI Factory and NeoCloud operators.
What does AI NOC span today?
Demonstrated coverage: the inference runtime (vLLM, NVIDIA Dynamo), GPU hardware (NVIDIA DCGM), compute health (Base Command Manager), workload orchestration (NVIDIA Run:ai), and the network fabric (Aviz ONES), correlated through one conversational interface. Storage platforms, fabric managers, and DPU-based wire observability are on the roadmap.
What is the relationship with NVIDIA?
AI NOC is in development with NVIDIA; its capabilities are demonstrated in the context of NVIDIA DSX Air, the flight-simulator environment for AI Factory operations. It complements the NVIDIA AI-factory software stack as the agentic layer that correlates across it.
Can AI NOC run fully private, sovereign cloud, air-gapped, regulated?
Yes. AI NOC runs on the Aviz Autonomous Agentic Platform, on your infrastructure, with a private LLM on your own GPUs or the endpoint you choose. Your telemetry stays in your environment. On-prem and air-gapped paths, tenant isolation, and role-based access are platform substrate, not add-ons.
When will AI NOC be generally available?
AI NOC is in early access; we don't publish release dates. Early-access participants get the current capability set, direct access to the product team, and a voice in what ships next.
How does AI NOC relate to Network Copilot?
Same platform, different job. Network Copilot is agentic AI NetOps for your multi-vendor network. AI NOC extends that foundation up the stack for AI factories, inference, GPU, compute, and orchestration correlated alongside the fabric. They share the platform's connectors, governance, and agent SDK, so skills built for one carry to the other.
Does AI NOC take action on my infrastructure by itself?
No. Its current phase is read-first: it observes, correlates, and recommends. Operator-directed steps, like tenant and workload lifecycle actions, execute only what you ask, under persona-scoped, role-based access. Broader governed action gates are on the platform roadmap.
Proven in the Lab. Validated with NVIDIA.
Full-stack correlation, inference, GPU, and orchestration, is demonstrated in Aviz's own AI Factory lab, against the questions operators actually ask. In the context of NVIDIA DSX Air, AI NOC's fabric-operations components run today; GPU and inference-layer metrics are not yet part of the DSX Air environment.
"Why is this inference endpoint slow?"
Answered with cross-layer correlation across vLLM, DCGM, BCM, and ONES, ending in a named layer and its evidence.
"Give me a full token-efficiency summary across all servers."
Answered with fleet-wide TTFT, throughput, MFU, KV-cache, and a per-server bottleneck verdict.
"Show me fabric health and flag any link or topology anomalies."
Answered via Aviz ONES fabric telemetry, the fabric-operations slice of AI NOC demonstrated in DSX Air today.


