|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH v2 0/2] xennet/xenvif: fix TX doorbell coalescing and improve RX NBL batching
On 28/09/2026 12:42, david ambu wrote: > Changes in v2: > - Replace Co-Authored-By trailer with Assisted-by (no email address) > - xennet: skip EtherType matching for LLC/SNAP frames (TypeOrLength <= > ETHERNET_MTU), > zeroing EtherType so those frames are excluded from > NDIS_RECEIVE_FLAGS_SINGLE_ETHER_TYPE > batching, matching the behaviour of xenvif!__ParseEthernetHeader Both patches: Reviewed-by: Tu Dinh <ngoc-tu.dinh@xxxxxxxxxx> > > Background > ---------- > NDIS passes packets to the miniport as chains of Net Buffer Lists (NBLs), > each NBL containing one or more Net Buffers (NBs). The xennet transmitter > signals xenvif's TX ring via a "More" flag: More=TRUE defers the ring > doorbell (kick), More=FALSE fires it. > > The intent is to batch all NBs within one NBL with a single kick at the > NBL boundary. Each NBL carries a single hash, so all its NBs map to the > same TX ring - making per-NBL kicking both correct and sufficient. > > Patch 1 - xenvif (TX fix) > -------------------------- > TransmitterQueuePacket was overriding More=FALSE unconditionally in the > XENVIF_PACKET_HASH_ALGORITHM_NONE case, causing a ring kick on every > individual NB instead of once per NBL. Removing the override lets the > More flag from xennet flow through correctly, giving ALGORITHM_NONE the > same per-NBL kick behaviour as ALGORITHM_TOEPLITZ. > > Patch 2 - xennet (RX batching) > ------------------------------- > Three improvements to the RX receive-indicate path: > > - Raise NblBatchMax default to 32 (configurable 1-64 via registry) > so NDIS receives larger batches in a single > NdisMIndicateReceiveNetBufferLists call, reducing per-indication > overhead. > > - Split batches at EtherType boundaries and set > NDIS_RECEIVE_FLAGS_SINGLE_ETHER_TYPE per batch, enabling NDIS to > apply protocol-layer optimisations. > > - In __ReceiverPushPackets: sample Returned before the InterlockedAdd > so it is read prior to the implicit full memory barrier that > InterlockedAdd provides; use the return value directly as Indicated > and remove the now-redundant KeMemoryBarrier. > > Performance (inter-host, Windows desktop VM, 10G, 3 iterations x 60s) > ----------------------------------------------------------------------- > > Without patch With patch (32 NBL default) > Single-stream TCP 4525 Mbits/s 5550 Mbits/s (+23%) > Single-stream UDP 372 Mbits/s 441 Mbits/s (+19%) > 8-stream TCP 11931 Mbits/s 17017 Mbits/s (+43%) > 8-stream UDP 87 Mbits/s 120 Mbits/s (+38%) > TCP latency (psping) 0.8 ms 0.9 ms (no regression) > ICMP latency (ping) <1 ms <1 ms (no regression) > > Tested at NblBatchMax = 8, 16, 32, 64. Peak throughput at 8-16; > larger values show marginal returns. Default of 32 is a conservative > mid-point. Latency is unaffected across all batch sizes. > > david ambu (2): > xenvif: fix TX doorbell coalescing for ALGORITHM_NONE traffic > xennet: improve NBL batching receive path > > src/xennet.inf | 6 +++ > src/xennet/adapter.c | 13 ++++++ > src/xennet/adapter.h | 8 ++++ > src/xennet/receiver.c | 110 +++++++++++++++++++++++++++++++++++---- > src/xenvif/transmitter.c | 1 - > 5 files changed, 107 insertions(+), 31 deletions(-) > > -- > 2.51.0.windows.1 > -- Ngoc Tu Dinh | Vates XCP-ng Developer XCP-ng & Xen Orchestra - Vates solutions web: https://vates.tech
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |