|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH 0/2] xennet: improve RX NBL batching and fix TX doorbell coalescing
On 17/09/2026 17:43, david ambu wrote: > This series addresses two independent throughput issues in xennet. > > Patch 1 — NBL batching on the receive path: > > The receive path previously called NdisMIndicateReceiveNetBufferLists > once per NBL (batch size of 1), causing excessive NDIS stack overhead > on high-throughput workloads. This patch raises NblBatchMax to 32 > (configurable up to 64), groups consecutive NBLs with the same > EtherType into a single indicate call with > NDIS_RECEIVE_FLAGS_SINGLE_ETHER_TYPE, and fixes __ReceiverPushPackets > to move KeMemoryBarrier before the Returned/Indicated reads and return > early when the queue is empty. > > Patch 2 — TX doorbell coalescing across NBL boundaries: > > __TransmitterSendNetBufferList computed the More flag as > (NET_BUFFER_NEXT_NB(NetBuffer) != NULL), which evaluates to FALSE for > the last NB of every NBL even when the outer loop still has additional > NBLs to submit. This caused a premature doorbell ring between > back-to-back NBLs. Fixed by passing MoreNbls from the outer loop. > > Note: a companion fix in xenvif is also required — > TransmitterQueuePacket unconditionally reset More = FALSE for > non-RSS traffic, overriding the value passed in from xennet. > > Performance (iperf2, 60 s, 3 runs, Windows PV guest): > > Single-stream TCP: 4,606 -> 6,056 Mbps (+31.5%) > 8-stream TCP: 14,663 -> 18,679 Mbps (+27.4%) > 8-stream UDP: 1,117 -> 1,230 Mbps (+10.1%) > Single-stream UDP: 327 -> 329 Mbps (no change, expected) > > david ambu (2): > xennet: improve NBL batching receive path > xennet: fix TX doorbell coalescing — More flag dropped at NBL > boundaries > > src/xennet.inf | 6 +++ > src/xennet/adapter.c | 13 +++++ > src/xennet/adapter.h | 8 +++ > src/xennet/receiver.c | 108 +++++++++++++++++++++++++++++---------- > src/xennet/transmitter.c | 9 ++-- > 5 files changed, 114 insertions(+), 30 deletions(-) > Hello, I don't think this is the right method. XenVif's TransmitterQueuePacket is eventually called by XenNet's __TransmitterSendNetBufferList which is eventually called by MiniportSendNetBufferLists. MiniportSendNetBufferLists states that different NBLs can be split into multiple transmit queues if they belong to different connections. In this case, the More flag is needed to terminate each NBL because if subsequent NBLs are not assigned to the same TX queue, you will not drain the queue of the previous NBL. So you may end up with packets stuck in the TX queue without being drained to the backend. For the purpose of batching NBLs, XenNet will need to be aware of the target queue of each NBL and group together NBLs destined for the same queue. Then further optimizations can be done by letting XenVif accept an entire chain of NBLs instead. -- Ngoc Tu Dinh | Vates XCP-ng Developer XCP-ng & Xen Orchestra - Vates solutions web: https://vates.tech
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |