Aviz Logo
Contact
Modern network stack built for the AI era.
Build AI Factories
Start here: Neocloud, sovereign AI, and private AI cloud users
Network Automation with ONESToken Economics with AI ParserAgentic Operation with AI-NOCPOC Validation with FTAS
Modernize your Network
Start here: Cisco, Arista, Juniper network users; Gigamon and Netscout observability users
Autonomous Agentic AI PlatformNetwork CopilotAviz Packet BrokerAviz Service NodeVirtual Aviz Service Node / vTAP
SONiC & AI Experience Hub
Validate open networking and AI-driven infrastructure with real-world, multi-vendor testing environments.
Partner with Aviz Networks
Join our ecosystem of channel and technology partners. Together we deliver open networking solutions that drive innovation and growth.
Make Networks for AI. Introduce AI in your Networks.
End-to-end solutions, any NOS, any switch, any ASIC, any LLM, any application , backed by partner best practices, proven tech, and SLAs.
Explore why Aviz is the best partner to modernize your network with.
Explore Case Studies, TCO and ROI Calculators, Certifications, Community and News room
Aviz Training and Certification
Learn and Certify in SONiC and AI
24/7 World-Class SONiC Support & Proven Services.
Our dedicated team delivers round-the-clock, world-class SONiC support with unmatched quality, scalability, and efficiency, keeping your network optimized, secure, and always running at its best.
Hamburger
Aviz Logo

Make every AI request measurable.

AI-Parser — AI Traffic Observability
AI-Parser turns AI inference traffic into per-request, per-tenant intelligence: tokens, latency, throughput, identity, prompt signals, and completion status — without instrumenting the application.
Skip the browsing. Ask AI.
Hello! How may I help you today?
The Problem

GPUs can be healthy while AI applications perform poorly.

You're billed for tokens and blamed for latency, but the traffic between the user and the model is still a black box. Traditional observability tells you traffic is flowing — AI-Parser tells you what that traffic means.

01

Know who is consuming AI

Separate traffic by user IP, tenant, application, or stable API-key identity, not only by model.

02

Measure real token usage

Capture input, output, and total token counts carried in the AI response, then roll them up by tenant.

03

See the experience users receive

Track time to first token, end-to-end response time, inter-token latency, and tokens per second.

04

Find waste and contention

Identify aborted streams, truncated answers, over-allocation, noisy neighbors, and struggling workloads.

AI-Parser observes the request and response as delivered on the wire, creating an independent transaction record for AI usage and experience.

How It Works

See the AI call, not just the GPU.

AI-Parser reconstructs OpenAI-compatible HTTP and JSON traffic, including streamed responses, to expose the information needed for cost, performance, governance, and operations.

AI-Parser passively taps mirrored traffic between applications, inference engines, and GPU infrastructure, then streams enriched per-call records into your existing observability stack

One pass over the wire — cost, performance, and governance all read from the same record.

Turn shared GPU infrastructure into accountable services.

When the orchestration platform maps each tenant or application to an IP and GPU allocation, AI-Parser records align naturally with that model.

Token usage and cost by tenant for chargeback or showback
TTFT, end-to-end latency, inter-token latency, and throughput by tenant
Workload profiles by model, prompt size, request rate, and call type
Noisy-neighbor, fairness, and under-utilization insight
Aborted streams and wasted token-budget signals by tenant
Features

What AI-Parser adds beyond model-server metrics

Engine metrics stay the best source for GPU internals. AI-Parser adds identity, content, and delivered experience.

AreaModel-server viewAI-Parser viewWhy it matters
IdentityAggregate, labelled by modelClient IP, tenant, stable API-key identityChargeback, SLA, and accountability per tenant
Prompt contentNever exposedParsed from the body — fingerprint or textGovernance, DLP, prompt-cache analysis
TokensSelf-reported by the engineCounted independently on the wireAn audit trail against mis-billing and drift
End-to-end latencyRequest entry to completionFirst request byte to last response byteThe latency the user actually experienced
Time to first tokenAdmit to first tokenRequest in to first byte outCatches host stack, API server, and queue delay
Aborts and wastePartial failure countersTCP FIN/RST with no completion markerSees the client hang up, and attributes it

Vendor-neutral: vLLM, TGI, Triton, TensorRT-LLM, Dynamo, and SGLang all speak the same OpenAI-compatible schema, so AI-Parser reads them the same way.

Bridging network, GPU, and application observability

Most AI infrastructure monitoring operates in separate domains:

GPU telemetryTells you whether accelerators are busy.
Network telemetryTells you whether packets and flows are healthy.
Application telemetryTells you whether the AI service is responding.

AI-Parser is the correlation layer between them — particularly important for AI Clouds and GPU-as-a-Service providers, who need to tie network flows to tenant identity, workload type, and real-time performance signals rather than treating traffic as generic flows.

Use Cases

A different use case for every team, from one tap.

The same passive read of AI traffic answers a different question depending on who's asking — cost, performance, governance, streaming, or agentic-flow visibility.

Cost

Usage and token economics

Model, call type, request parameters, prompt tokens, completion tokens, total tokens, and request volume.

Performance

Performance as delivered

Time to first token, wire end-to-end time, inter-token latency, throughput, and stream completion.

Governance

Identity and privacy controls

Per-IP and per-tenant records, stable identity fingerprints, prompt hashing by default, and optional raw text.

Streaming

Streaming-aware analysis

Reassemble token chunks, detect first and last response events, count SSE chunks, and recognize completion.

Agents

Agentic-flow visibility

Observe the multiple backend model calls created by one user question and regroup related calls through identity.

Accuracy

Honest, structured records

Read or compute fields from observed traffic. When a value is not present, leave it blank instead of guessing.

From “Is the GPU busy?” to “Is the GPU delivering AI value?”

End-to-End Visibility
Understand AI traffic from the application and agent through the network to the LLM and GPU.
Real User Experience
Track TTFT and inter-token latency, not just network latency or GPU utilization.
AI Consumption
Measure token usage by application, workload, or model.
Faster Root-Cause
Pinpoint whether poor performance is the app, network, model, or infrastructure.
GPU Efficiency
Spot expensive GPUs waiting on traffic, requests, or application dependencies.
Infrastructure-Independent
Deploy on x86 or DPUs, alongside your existing AI and GPU platforms.
Open Telemetry
Export enriched inference KPIs via Kafka into your existing stack.

Technical questions.

Does AI-Parser require instrumenting my application?

No. It observes traffic as delivered on the wire — there's nothing to instrument in the application or inference engine.

Where does it deploy?

As a passive tap on an x86 host or a BlueField-3 DPU, out of the data path — it reads mirrored traffic, not live production traffic.

Which inference engines does it support?

vLLM, TGI, Triton, TensorRT-LLM, Dynamo, and SGLang — all read the same way, since they share an OpenAI-compatible schema.

Does it store my users' prompts?

Prompts are converted to one-way fingerprints by default. Raw prompt text is retained only when explicitly enabled.

What happens to fields AI-Parser can't observe?

They're left blank — never guessed or fabricated.

How does this relate to my model server's own metrics?

They're complementary, not competing. Model-server metrics remain the best source for GPU internals; AI-Parser adds the identity, content, and delivered-experience context that's only visible on the wire.

How to Buy

Talk to us about licensing.

Pricing details coming soon

AI-Parser licensing terms are being finalized. Talk to an Aviz architect for current pricing and deployment options for your environment.