[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [PATCH 0/2] xennet: improve RX NBL batching and fix TX doorbell coalescing


  • To: david ambu <david.preetham@xxxxxxxxxx>, win-pv-devel@xxxxxxxxxxxxxxxxxxxx
  • From: Tu Dinh <ngoc-tu.dinh@xxxxxxxxxx>
  • Date: Mon, 21 Sep 2026 15:07:15 +0200
  • Authentication-results: eu.smtp.expurgate.cloud; dkim=pass header.s=selector1 header.d=vates.tech header.i="@vates.tech" header.h="From:Subject:Date:Message-ID:To:MIME-Version:Content-Type:In-Reply-To:References:Feedback-ID"
  • Delivery-date: Mon, 21 Sep 2026 13:07:24 +0000
  • Feedback-id: default:8631fc262581453bbf619ec5b2062170:Sweego
  • List-id: Developer list for the Windows PV Drivers subproject <win-pv-devel.lists.xenproject.org>

On 17/09/2026 17:43, david ambu wrote:
> This series addresses two independent throughput issues in xennet.
> 
> Patch 1 — NBL batching on the receive path:
> 
>    The receive path previously called NdisMIndicateReceiveNetBufferLists
>    once per NBL (batch size of 1), causing excessive NDIS stack overhead
>    on high-throughput workloads. This patch raises NblBatchMax to 32
>    (configurable up to 64), groups consecutive NBLs with the same
>    EtherType into a single indicate call with
>    NDIS_RECEIVE_FLAGS_SINGLE_ETHER_TYPE, and fixes __ReceiverPushPackets
>    to move KeMemoryBarrier before the Returned/Indicated reads and return
>    early when the queue is empty.
> 
> Patch 2 — TX doorbell coalescing across NBL boundaries:
> 
>    __TransmitterSendNetBufferList computed the More flag as
>    (NET_BUFFER_NEXT_NB(NetBuffer) != NULL), which evaluates to FALSE for
>    the last NB of every NBL even when the outer loop still has additional
>    NBLs to submit. This caused a premature doorbell ring between
>    back-to-back NBLs. Fixed by passing MoreNbls from the outer loop.
> 
>    Note: a companion fix in xenvif is also required —
>    TransmitterQueuePacket unconditionally reset More = FALSE for
>    non-RSS traffic, overriding the value passed in from xennet.
> 
> Performance (iperf2, 60 s, 3 runs, Windows PV guest):
> 
>    Single-stream TCP:  4,606 -> 6,056 Mbps  (+31.5%)
>    8-stream TCP:      14,663 -> 18,679 Mbps  (+27.4%)
>    8-stream UDP:       1,117 ->  1,230 Mbps  (+10.1%)
>    Single-stream UDP:    327 ->    329 Mbps  (no change, expected)
> 
> david ambu (2):
>    xennet: improve NBL batching receive path
>    xennet: fix TX doorbell coalescing — More flag dropped at NBL
>      boundaries
> 
>   src/xennet.inf           |   6 +++
>   src/xennet/adapter.c     |  13 +++++
>   src/xennet/adapter.h     |   8 +++
>   src/xennet/receiver.c    | 108 +++++++++++++++++++++++++++++----------
>   src/xennet/transmitter.c |   9 ++--
>   5 files changed, 114 insertions(+), 30 deletions(-)
> 

Hello,

I don't think this is the right method.

XenVif's TransmitterQueuePacket is eventually called by XenNet's 
__TransmitterSendNetBufferList which is eventually called by 
MiniportSendNetBufferLists. MiniportSendNetBufferLists states that 
different NBLs can be split into multiple transmit queues if they belong 
to different connections.

In this case, the More flag is needed to terminate each NBL because if 
subsequent NBLs are not assigned to the same TX queue, you will not 
drain the queue of the previous NBL. So you may end up with packets 
stuck in the TX queue without being drained to the backend.

For the purpose of batching NBLs, XenNet will need to be aware of the 
target queue of each NBL and group together NBLs destined for the same 
queue. Then further optimizations can be done by letting XenVif accept 
an entire chain of NBLs instead.


--
Ngoc Tu Dinh | Vates XCP-ng Developer

XCP-ng & Xen Orchestra - Vates solutions

web: https://vates.tech

 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.