|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] [PATCH 0/2] xennet/xenvif: fix TX doorbell coalescing and improve RX NBL batching
These two patches address TX doorbell coalescing correctness in xenvif
and RX receive-path NBL batching efficiency in xennet.
Background
----------
NDIS passes packets to the miniport as chains of Net Buffer Lists (NBLs),
each NBL containing one or more Net Buffers (NBs). The xennet transmitter
signals xenvif's TX ring via a "More" flag: More=TRUE defers the ring
doorbell (kick), More=FALSE fires it.
The intent is to batch all NBs within one NBL with a single kick at the
NBL boundary. Each NBL carries a single hash, so all its NBs map to the
same TX ring - making per-NBL kicking both correct and sufficient.
Patch 1 - xenvif (TX fix)
--------------------------
TransmitterQueuePacket unconditionally overrode More=FALSE in the
XENVIF_PACKET_HASH_ALGORITHM_NONE case, causing a ring kick on every
individual NB instead of once per NBL. Removing the override lets the
More flag from xennet flow through correctly, giving ALGORITHM_NONE the
same per-NBL kick behaviour as ALGORITHM_TOEPLITZ.
Patch 2 - xennet (RX batching)
-------------------------------
Three improvements to the RX receive-indicate path:
- Raise NblBatchMax default to 32 (configurable 1-64 via registry)
so NDIS receives larger batches in a single
NdisMIndicateReceiveNetBufferLists call, reducing per-indication
overhead.
- Split batches at EtherType boundaries and set
NDIS_RECEIVE_FLAGS_SINGLE_ETHER_TYPE per batch, enabling NDIS to
apply protocol-layer optimisations.
- In __ReceiverPushPackets: sample Returned before the InterlockedAdd
so it is read prior to the implicit full memory barrier that
InterlockedAdd provides; use the return value directly as Indicated
and remove the now-redundant KeMemoryBarrier.
Performance (inter-host, Windows desktop VM, 10G, 3 iterations x 60s)
-----------------------------------------------------------------------
Without patch With patch (32 NBL default)
Single-stream TCP 4525 Mbits/s 5550 Mbits/s (+23%)
Single-stream UDP 372 Mbits/s 441 Mbits/s (+19%)
8-stream TCP 11931 Mbits/s 17017 Mbits/s (+43%)
8-stream UDP 87 Mbits/s 120 Mbits/s (+38%)
TCP latency (psping) 0.8 ms 0.9 ms (no regression)
ICMP latency (ping) <1 ms <1 ms (no regression)
Tested at NblBatchMax = 8, 16, 32, 64. Peak throughput at 8-16;
larger values show marginal returns. Default of 32 is a conservative
mid-point. Latency is unaffected across all batch sizes.
david ambu (2):
xenvif: fix TX doorbell coalescing for ALGORITHM_NONE traffic
xennet: improve NBL batching receive path
src/xennet.inf | 6 +++
src/xennet/adapter.c | 13 ++++++
src/xennet/adapter.h | 8 ++++
src/xennet/receiver.c | 110 +++++++++++++++++++++++++++++++++++----
src/xenvif/transmitter.c | 1 -
5 files changed, 107 insertions(+), 31 deletions(-)
--
2.51.0.windows.1
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |