Optimizing GNNs for the Wild: PyTorch-to-ONNX Acceleration of GarNet on CPUs and GPUs
Benchmarking GarNet in PyTorch versus ONNX Runtime on multi-core CPUs and NVIDIA A100 GPUs: up to 5× CPU and ~2× GPU speedup with FP32 parity at 10⁻⁷, targeting the LHCb GPU trigger (HLT1/Allen).
May 25, 2026