ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers
ONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
One portable artifact vs an LLM server
Prefer
ONNX + ORT for encoders, classic ML, and edge
Fixed-shape or modestly dynamic graphs. One file across CPU, CUDA, TensorRT EP, CoreML, DirectML, OpenVINO. Small runtime, no Python training stack in prod.
- skl2onnx / onnxmltools cover trees; torch.onnx covers tensors.
- Provider order assigns each node to the first EP that claims it.
- ORT_ENABLE_ALL can change numerics — compare to a golden output.
Alternative
Hand-roll a 70B generative LLM in ONNX
You will miss paged KV, continuous batching, and multi-GPU TP. That is what SGLang, TensorRT-LLM, and vLLM-class servers are for.
- ORT GenAI exists; specialized servers still win on large decode.
- Shipping a scikit-learn model inside a vLLM container is the inverse smell.
- Unsupported ops silently fall back to CPU — a PCIe trip per token.
Export, optimize, execute
Declare every dynamic dim. Put the fastest EP first. Validate fusion on a golden input.
- 1
Export
torch.onnx.export or skl2onnx. Set opset, names, and dynamic axes / shapes. - 2
Optimize
ORT graph optimizers: fold constants, fuse MatMul+Add, optional EP-specific partitions. - 3
Execute
Session with provider list. CPU is the fallback, not the default you wanted. - 4
Validate
Compare optimized vs unoptimized on a golden input before trusting a fused graph.
Overview
ONNX is the interchange format that lets a model trained in one framework run in another runtime, on a different OS, accelerator, or language. ONNX Runtime (ORT) executes those graphs and rewrites them for speed.
If you can export cleanly, pick the right opset, and select the right Execution Provider (EP), you get a portable artifact that serves classic ML and transformer encoders at low latency without a Python training stack in production. If you cannot, you get silent numerical drift, unsupported-op failures at load time, or a graph that runs on CPU when a GPU sits idle.
torch.onnx / exporters
torch.onnx.export is the legacy tracer-based path. torch.onnx.export(..., dynamo=True) and the torch.export → torch.onnx pipeline are the modern, graph-capture path that handles control flow and dynamic shapes more faithfully.
Key arguments: opset_version, input_names / output_names, dynamic_axes (legacy) or dynamic_shapes (dynamo), do_constant_folding, and training=TrainingMode.EVAL.
Classic ML models (scikit-learn, XGBoost, LightGBM) go through skl2onnx, onnxmltools, or hummingbird, producing operator sets like TreeEnsembleClassifier that ORT executes natively.
Transformers: export the encoder / decoder as a graph with input_ids, attention_mask, and position_ids. Keep the KV cache as explicit graph inputs/outputs if you want stateful decoding; otherwise re-feed the full sequence.
ORT graph optimizers
ORT applies levels via GraphOptimizationLevel:
- ORT_DISABLE_ALL — debugging only.
- ORT_ENABLE_BASIC — constant folding, redundant node elimination, identity removal.
- ORT_ENABLE_EXTENDED — operator fusion (
MatMul+Add→Gemm,Conv+BN→Conv, attention fusions), layout transforms. - ORT_ENABLE_ALL — extended plus EP-specific fusions (for example TensorRT subgraph capture).
Optimizations are EP-aware: the TensorRT EP partitions the graph into TRT subgraphs and leaves unsupported nodes on CUDA/CPU. Always compare optimized vs unoptimized outputs on a golden input before trusting a fused graph.
Execution Providers
- CPU EP — MLAS kernels, always available, good for small models and fallback nodes.
- CUDA EP — cuDNN / cuBLAS kernels, FP16 / INT8, the default GPU path.
- TensorRT EP — builds TRT engines from subgraphs; best throughput on NVIDIA, slow first-run build, cached in
trt_engine_cache. - CoreML EP — Apple Neural Engine / GPU on macOS and iOS; FP16 by default.
- Others: DirectML (Windows), OpenVINO (Intel), ROCm, QNN (Qualcomm), WebGPU / WASM for browsers.
Provider order matters: ORT assigns each node to the first provider that claims it. Put the fastest provider first and CPU last.
This TensorRT EP is subgraph capture inside ORT, not the full TensorRT-LLM compile pipeline.
Pitfalls: dynamic axes, opset, unsupported ops
- Dynamic axes must be declared for every varying dimension; a batch dim fixed at export time will fail on batch 2.
- Opset too new → runtime rejects; too old → no fused attention. Match opset to the ORT version you deploy.
- Unsupported ops silently fall back to CPU, creating a PCIe round-trip per token. Inspect with
session.get_providers()and ORT profiling. - Tracer-based export drops Python control flow; use dynamo export or rewrite with
torch.where. - Data-dependent shapes and
NonZero-style ops break TensorRT partitioning.
When ONNX is right vs LLM-specific servers
Choose ONNX / ORT when: the model is a fixed-shape or modestly dynamic encoder, a classic ML model, an embedding / reranker, a vision model, or you need one artifact across CPU / GPU / edge / mobile.
Choose vLLM, SGLang, or TensorRT-LLM when: you need paged KV cache, continuous batching, speculative decoding, or multi-GPU tensor parallelism for a generative LLM. ONNX Runtime GenAI exists for LLM decoding, but the specialized servers still win on throughput for large generative workloads. Those servers are siblings / comparators — not this page.
Comparative: portable vs specialized
| Option | Pros | Cons | Wrong stack when |
|---|---|---|---|
| ONNX + ORT | Portable, many EPs, small runtime | Export friction, op coverage gaps | You need paged-attention LLM serving |
| TensorRT-LLM | Peak NVIDIA throughput | NVIDIA-only, engine rebuild per shape | You must run on CPU or Apple silicon |
| SGLang / vLLM | Continuous batching, easy HF load | GPU-centric, heavier deps | You ship a 5 MB edge classifier |
| TorchScript / eager | No export step | Python runtime in prod, slow cold start | You need a C++ / mobile runtime |
Wrong-stack smell: exporting a 70B generative LLM to ONNX and hand-rolling batching, or shipping a scikit-learn model inside a vLLM container.
Flow
- 1
1 PyTorch or sklearn model
- next2 Export ONNX + opset
- 2
2 Export ONNX + opset
- next3 Declare dynamic axes
- 3
3 Declare dynamic axes
- next4 ORT graph optimizers
- 4
4 ORT graph optimizers
- next5 Fold constants and fuse
- 5
5 Fold constants and fuse
- next6 Pick execution provider
- 6
6 Pick execution provider
- next7 CPU CUDA TRT or CoreML
- 7
7 CPU CUDA TRT or CoreML
- next8 Run and validate
- 8
8 Run and validate
Lesson map
ONNX & ONNX Runtime — Export, Graph Optimizers & Execution Providers
ONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 PyTorch or sklearn model"] b["2 Export ONNX + opset"] c["3 Declare dynamic axes"] d["4 ORT graph optimizers"] a -->|1 PyTorch or sklearn model| b b -->|2 Export ONNX + opset| c c -->|3 Declare dynamic axes| d
Sandbox: mock export + session (Python)
No GPU. Tries real torch / ORT if present; otherwise returns the same shapes.
Press Run. Snippets must be self-contained — no network, files, or native modules.
TypeScript session shape
onnxruntime-node API shape. The playground mocks run.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · LLM servers — pointer only
Paged KV and continuous batching: SGLang. Compiled NVIDIA engines: TensorRT-LLM. vLLM is named as a generative-LLM alternative, not taught. Stack choice: playbook.
Pitfalls
The model ran on batch 1 and died on batch 2. Which export flag did you miss — and how would you prove a node fell back to CPU?
Interview Q&A
Why does my ONNX model fail on batch size 2 when it worked on batch 1?
Answer
The batch dimension was static at export. Re-export with a dynamic axis on dim 0, or use dynamic_shapes with the dynamo exporter.
What does ORT_ENABLE_ALL actually change?
Answer
It enables extended fusions plus EP-specific partitioning. It can alter numerics slightly, so validate against a golden output.
How do you know if a node fell back to CPU?
Answer
Enable ORT profiling or inspect the optimized graph. Provider assignment is per-node and the last provider (usually CPU) catches leftovers.
When would you not use ONNX?
Answer
Generative LLM serving with paged KV cache and continuous batching — use vLLM, SGLang, or TensorRT-LLM instead. Those pages own that depth; this page owns portable graphs.
Go Deeper
Public onnx / onnxruntime docs only: