Windows network card driver performance tuning guide

Updated at:

This topic describes how to troubleshoot and tune network performance issues such as intermittent packet loss on Alibaba Cloud ECS Windows instances that use the virtio network driver (driver name: NetKVM). It provides tuning recommendations for key parameters such as RxCapacity (number of receive buffers) and *NumRssQueues (number of RSS queues), and summarizes the display names, default values, value ranges, and configuration methods of the advanced parameters of NetKVM NICs.

Scope and usage guidance

This topic applies to Alibaba Cloud ECS Windows instances that use the NetKVM virtio network driver. For the display names, default values, value ranges, and configuration methods of the parameters, see NetKVM driver parameter reference.

NetKVM is the virtio network driver used on Alibaba Cloud ECS Windows instances. After the driver is installed, it registers a set of configurable parameters in the system to control behaviors such as checksum offload, large send offload, multi-queue receiving (RSS), receive segment coalescing (RSC), VLAN, MTU, and logging.

The default values are suitable for most scenarios. Do not modify any parameter unless you have a clear troubleshooting conclusion.

Troubleshooting intermittent packet loss

Symptoms and causes

Typical symptoms: network monitoring occasionally shows packet loss, and the TCP retransmission rate rises briefly, but sustained stress tests may not reproduce the issue. Packet loss often coincides with traffic bursts or periods when the system is busy.

One common cause: the CPUs in the instance cannot process received data in time. When the instantaneous receive rate exceeds the processing speed of the driver, the receive buffers of the NIC gradually fill up. After the buffers are exhausted, newly arriving packets have nowhere to be stored and can only be dropped. After the burst passes, the CPUs catch up and the buffers are freed, and the packet loss disappears. This is why the issue is intermittent and hard to reproduce.

How to identify

Before tuning parameters, use the process of elimination to determine whether the packet loss may be caused by receive buffer exhaustion. This avoids blind parameter tuning when the cause of the packet loss is unclear:

Note

Packet loss caused by buffer exhaustion occurs before packets reach the driver and is not reflected in the NIC statistics counters of the operating system. Therefore, it cannot be confirmed by a single metric. You need to evaluate the following clues together.

  • Rule out other stages: First confirm that the packet loss does not occur at the network link layer or on the peer side (for example, drops by security groups or bandwidth throttling). If other stages can be ruled out, receive buffer exhaustion in the instance is one of the most likely causes;

  • Check time correlation: Compare the times of packet loss with the CPU utilization and network traffic curves of the instance to confirm whether the packet loss coincides with high CPU load or traffic bursts;

  • Check CPU distribution: Observe the utilization of each CPU core and the distribution of receive processing to confirm whether receive processing is concentrated on a few CPU cores when packet loss occurs;

  • Increase-and-verify (final confirmation): If the preceding clues point to buffer exhaustion, increase RxCapacity and observe whether the packet loss disappears. If it disappears, buffer exhaustion can be confirmed as the cause.

If the initial assessment points to delayed processing on the receive side, tune along the following two paths.

Tuning path 1: Increase RxCapacity (absorb bursts)

Principle: a larger receive buffer (for example, increased from the default 256 to 1024 or 2048) can cache more packets during transient bursts while waiting for CPU processing, which prevents drops caused by buffer exhaustion. In Device Manager, this parameter is displayed as Init.MaxRxBuffers.

Usage boundaries (important):

  • Effective only for transient bursts: In the case of sustained overload (the CPU processing capacity is persistently insufficient), increasing the buffer only delays packet loss rather than eliminating it. You need to combine this with tuning path 2 (engage more CPUs in receive processing) or upgrade the instance;

  • Confirm the effective upper limit first: Not all instances can use values greater than 256. This depends on the instance type and image version. For details, see Effective upper limit. Confirm that the instance meets the requirements before you increase this value;

  • Estimate the memory overhead beforehand: The memory consumed by the receive buffers is proportional to the value and the number of NIC queues (the number of NIC queues is determined by the instance type and cannot be configured by the user). When the value is 4096, multi-queue instances can consume GB-level memory. For quantitative estimation, see Memory overhead. For instances with small memory, estimate the overhead before tuning.

Tuning path 2: Increase the number of RSS queues (engage more CPUs in receive processing)

Principle: RSS distributes received traffic across multiple receive queues by network connection, and different CPUs process the queues in parallel. On multi-core instances, increasing *NumRssQueues (for example, from the default 8 to a value close to the number of vCPUs) prevents the receive load from being concentrated on a few cores and improves overall receive processing capacity. In Device Manager, this parameter is displayed as Maximum Number of RSS Queues.

Prerequisites and boundaries (important):

  • *RSS (displayed as Receive Side Scaling in Device Manager) must remain enabled (enabled by default);

  • The configured value may differ from the effective value: The effective number of queues does not exceed the number of vCPUs of the instance. After configuration, check the RSS status information of the operating system to confirm the effective number of queues;

  • Not effective for a single high-traffic connection: The packets of the same connection (single flow) are always processed by the same CPU. Adding queues cannot distribute the load of a single flow. This path applies to multi-flow concurrent scenarios. For single-flow packet loss, prefer tuning path 1 or split the connections at the application layer.

Decision recommendations

Characteristics of packet loss

Priority parameter

Notes

Transient burst packet loss that recovers automatically after the burst

Increase RxCapacity (for example, 1024/2048)

Confirm the effective upper limit (see Section 3.1) and memory overhead (see Section 3.2) first

Multi-flow concurrency with the receive load concentrated on a few CPU cores

Increase *NumRssQueues (close to the number of vCPUs)

Requires *RSS to be enabled. Not effective for single-flow packet loss. Verify the effective number of queues after configuration

Sustained overload (the CPU processing capacity is persistently insufficient)

Combine both

Parameter tuning only mitigates the issue. We recommend that you also consider upgrading the instance (more vCPUs or an instance type with higher network capability)

The preceding table also applies to preventive tuning for high-throughput/high-PPS scenarios: RxCapacity can be increased, and *NumRssQueues can be set to a value close to the number of vCPUs. Keep the other offload parameters (checksum/LSO/RSS/RSC) at their default values (all enabled by default).

Note

Changes to the preceding parameters require a NIC restart to take effect. Perform the operations during off-peak hours and continue monitoring packet loss metrics after tuning to verify the effect.

Detailed description of RxCapacity

Effective upper limit

On Alibaba Cloud ECS, whether the number of receive buffers can use values greater than 256 (up to 4096) depends on the instance type and image version. Support for 4096 is guaranteed when both of the following conditions are met:

Condition

Requirement

Instance type

Alibaba Cloud ECS generation 8 or later instance types (such as g8i)

Image

Alibaba Cloud Windows public images released in September 2023 or later, or custom images created from such public images

For instances that do not meet the preceding conditions, the effective upper limit may be 256. Even if Device Manager allows you to select a larger value (such as 1024 or 4096), values beyond the effective upper limit do not take effect.

Note
  • Memory consumption is based on the effective value. On an instance whose effective upper limit is 256, even if you set the value to 4096, only 256 buffers are actually used, and no additional memory is consumed;

  • The actual configurable range and effective values may vary with instance types and image versions. The display in Device Manager and actual test results shall prevail.

Memory overhead

Each receive buffer consumes about 72 KB of memory. This memory is fixed when the driver is loaded and does not vary with traffic. It is also unrelated to the MTU/*JumboPacket settings. Receive buffers are allocated per NIC queue (the number of NIC queues is determined by the instance type and cannot be configured by the user, and is unrelated to *NumRssQueues). Therefore:

Total memory of receive buffers ≈ 72 KB × effective number of buffers × number of NIC receive queues

Estimation examples (based on the effective number of buffers):

Effective number of buffers

Single queue

4 queues

8 queues

256 (default)

About 18 MB

About 72 MB

About 144 MB

1024

About 72 MB

About 288 MB

About 576 MB

4096

About 288 MB

About 1.1 GB

About 2.3 GB

This parameter has a significant impact on memory consumption. The default value of 256 is a balance between memory usage and burst absorption capability. If you increase RxCapacity on instances with small memory, pay attention to the memory overhead (on instances with multiple NIC queues, memory consumption multiplies with the number of queues). Estimate against the preceding table based on the memory specification of the instance before tuning.

Note

Memory estimates vary by driver version. Actual test results prevail. To reduce receive memory consumption, upgrade the virtio driver to version 58131 or later and enable the Mergeable Rx Buffers parameter in Device Manager (disabled by default): when enabled, each receive buffer consumes approximately 4 KB of memory, and each estimate in the preceding table shrinks to about 1/18 accordingly. For applicable limitations, see Receive buffer merging mode (Mergeable Rx Buffers). For more information, see Install the virtio driver.

NetKVM driver parameter reference

This chapter is for your reference when viewing and tuning advanced NIC parameters. The default values and value ranges of the parameters shall be subject to the values actually displayed in Device Manager on the instance (differences may exist between driver versions).

Note

If you encounter network performance issues such as intermittent packet loss, or need to evaluate tuning of the number of RxCapacity/RSS queues, see Troubleshooting intermittent packet loss.

How to configure parameters

You can configure the parameters in Device Manager. Procedure:

  1. Open Device Manager > Network adapters and find the virtio NIC;

  2. Right-click the NIC and choose Properties > the Advanced tab;

  3. In the Property list on the left, select a parameter (such as Init.MaxRxBuffers), change the value on the right, and click OK.

Note
  • The list on the left of the Advanced tab shows the display names of the parameters (such as Init.MaxRxBuffers). This chapter identifies parameters by their display names, which correspond one-to-one with the Device Manager interface.

  • After you click OK in Device Manager to save the changes, the system automatically restarts the NIC (the network is interrupted for a few seconds).

How changes take effect

The driver reads parameter settings only once during NIC initialization. Therefore, you must restart the NIC after modifying parameters. After you click OK in Device Manager to save the changes, the system automatically restarts the NIC, interrupting the network for a few seconds.

Warning

Restarting the NIC briefly interrupts all connections that use it and can affect network performance and stability. Perform the operation during off-peak hours, record the original parameter values, and prepare a rollback plan before making changes.

Parameter quick reference table

The NetKVM driver provides the following 24 configurable parameters:

Display name (advanced property)

Default value

Value range

Description

Assign MAC

(Empty)

12 hexadecimal characters

Overrides the MAC address of the NIC (optional parameter)

Init.Do802.1PQ

1

0 / 1

Enables or disables 802.1P/Q priority and VLAN tagging support (master switch)

Init.MaxRxBuffers

256

16/32/64/128/256/512/1024/2048/4096

Number of buffers per receive queue

Init.MaxTxBuffers

1024

16/32/64/128/256/512/1024

Number of buffers per transmit queue

IPv4 Checksum Offload

3

0–3

IPv4 header checksum offload

Jumbo Packet

1514

590–65500

Maximum frame size (including the 14-byte Ethernet header)

Large Send Offload V2 (IPv4)

1

0 / 1

IPv4 large send offload (LSO v2)

Large Send Offload V2 (IPv6)

1

0 / 1

IPv6 large send offload (LSO v2)

Logging.Enable

1

0 / 1

Driver logging switch

Logging.Level

0

0–8

Driver logging level

Mergeable Rx Buffers

0 (Disabled)

0 / 1

Receive buffer merging mode: when enabled, each receive buffer occupies about 4 KB, which significantly reduces receive memory usage (for details, see the Receive buffer merging mode section)

Maximum Number of RSS Queues

8

1–16

Maximum number of RSS queues

Offload.Rx.Checksum

31

0/1/3/27/31

Receive checksum offload (master switch)

Offload.Tx.Checksum

31

0/1/3/27/31

Transmit checksum offload (master switch)

Offload.Tx.LSO

2

0/1/2

Large send offload (master switch)

Priority and VLAN tagging

3

0–3

Priority and VLAN tagging mode

Receive Side Scaling

1

0 / 1

Receive-side multi-queue distribution (RSS) switch

Recv Segment Coalescing (IPv4)

1

0 / 1

IPv4 receive segment coalescing (RSC) switch

Recv Segment Coalescing (IPv6)

1

0 / 1

IPv6 receive segment coalescing (RSC) switch

TCP Checksum Offload (IPv4)

3

0–3

IPv4 TCP checksum offload

TCP Checksum Offload (IPv6)

3

0–3

IPv6 TCP checksum offload

UDP Checksum Offload (IPv4)

3

0–3

IPv4 UDP checksum offload

UDP Checksum Offload (IPv6)

3

0–3

IPv6 UDP checksum offload

VLan ID

0

0–4094

The VLAN ID that the NIC belongs to. 0 indicates that VLAN filtering is disabled

Parameters by category

Checksum offload

Checksum offload lets the NIC compute and verify the IP, TCP, and UDP checksums instead of the CPU, which reduces CPU usage in the instance. These parameters are capability requests to the device, and whether they take effect depends on whether the device supports them. There are many checksum-related parameters (seven in the following table). We recommend that you keep all of them at their default values.

Display name

Default value

Valid values

Offload.Tx.Checksum

31 (All)

31 = All; 27 = TCP/UDP(v4,v6); 3 = TCP/UDP(v4); 1 = TCP(v4); 0 = Disabled

Offload.Rx.Checksum

31 (All)

Same as above

IPv4 Checksum Offload

3

3 = Rx & Tx Enabled; 2 = Rx Enabled; 1 = Tx Enabled; 0 = Disabled

TCP Checksum Offload (IPv4)

3

Same as above

UDP Checksum Offload (IPv4)

3

Same as above

TCP Checksum Offload (IPv6)

3

Same as above

UDP Checksum Offload (IPv6)

3

Same as above

Recommendation: keep all values at their defaults. Disabling any item moves the checksum computation of the corresponding protocol to the instance CPUs, which increases CPU usage and reduces throughput under heavy traffic. Disabling TCP checksum offload also disables the corresponding large send offload (LSO). Temporarily adjusting these parameters is recommended only when you troubleshoot difficult network failures related to checksums.

Large send offload (LSO)

LSO offloads the segmentation of large TCP data blocks to the NIC, which significantly reduces CPU usage in the transmit direction. Similar to checksum offload, these parameters are capability requests to the device, and whether they take effect depends on whether the device supports them.

LSO is controlled by two layers of parameters, which have an AND relationship (both must be enabled to take effect):

Layer

Parameter

Function

Capability declaration (master switch)

Offload.Tx.LSO

Declares which versions of LSO the driver supports (0 = all disabled, 1 = IPv4, 2 = IPv4+IPv6)

Enable/disable control (sub-switch)

Large Send Offload V2 (IPv4/IPv6)

Enables or disables LSO individually by IP version within the declared capability

By default, both layers are fully enabled and no adjustment is required.

Display name

Offload.Tx.LSO

Default value

2 (Maximal)

Valid values

2 = Maximal (IPv4 + IPv6); 1 = IPv4 (IPv4 only); 0 = Disabled

Display name

Large Send Offload V2 (IPv4), Large Send Offload V2 (IPv6)

Default value

Both are 1 (Enabled)

Valid values

1 = Enabled; 0 = Disabled

Recommendation: keep all values at their defaults. Disabling either layer disables LSO for the corresponding version, which may reduce TCP transmit throughput and increase CPU usage. Temporarily adjusting these parameters is recommended only when you troubleshoot difficult issues in the transmit direction.

RSS (multi-queue receive distribution)

Receive Side Scaling

Display name

Receive Side Scaling

Default value

1 (Enabled)

Valid values

1 = Enabled; 0 = Disabled

Function: controls whether to enable flow-hash-based receive load balancing. When enabled, the driver computes a hash for each received packet and distributes the packet to a designated CPU through the indirection table, which ensures that packets of the same connection are always processed by the same CPU. When disabled, received packets are no longer balanced across multiple CPUs and are all processed by a single CPU.

Recommendation: keep the default value 1. Disabling RSS removes the controllability of load balancing and may cause some CPUs to be overloaded in multi-flow concurrent scenarios.

Maximum Number of RSS Queues

Display name

Maximum Number of RSS Queues

Default value

8

Valid values

An integer between 1 and 16 (step: 1)

On the Advanced tab in Device Manager, this parameter is displayed as Maximum Number of RSS Queues. Set it to the default value 8.

Function: the upper limit of the number of queues that can participate in receive processing at the same time. More queues allow more CPUs to participate in receive processing. The effective number of queues does not exceed the number of vCPUs of the instance.

Recommendation: keep the default value 8. For scenarios with many vCPUs (>8) and high network throughput requirements, you can evaluate increasing this value. For the method and notes, see Tuning path 2: Increase the number of RSS queues (engage more CPUs in receive processing).

RSC (receive segment coalescing)

Recv Segment Coalescing (IPv4/IPv6)

Display name

Recv Segment Coalescing (IPv4), Recv Segment Coalescing (IPv6)

Default value

Both are 1 (Enabled)

Valid values

1 = Enabled; 0 = Disabled

Function: when enabled, the driver declares to the device that it can receive coalesced large packets. The device then coalesces multiple small packets of the same TCP connection into large packets before passing them to the driver, which reduces the number of receive operations and lowers CPU usage. This parameter is a capability declaration (request), and whether it takes effect depends on whether the device supports it.

Recommendation: keep the default value 1. After it is disabled, CPU usage in the receive direction may increase.

Queue capacities

Init.MaxTxBuffers

Display name

Init.MaxTxBuffers

Default value

1024

Valid values

16, 32, 64, 128, 256, 512, 1024

Function: the number of buffers of each transmit queue. A larger value makes it less likely that transmission is paused due to buffer exhaustion during traffic bursts.

Recommendation: keep the default value 1024. A smaller value causes transmit queuing and throughput jitter under high-concurrency transmission and provides no actual benefit.

Init.MaxRxBuffers

Display name

Init.MaxRxBuffers

Default value

256

Valid values

16, 32, 64, 128, 256, 512, 1024, 2048, 4096

On the Advanced tab in Device Manager, this parameter is displayed as Init.MaxRxBuffers.

Function: the number of buffers of each receive queue. A larger value reduces packet loss under burst traffic, at the cost of more memory consumption.

Recommendation: keep the default value 256. If you need to evaluate increasing it: the actual configurable upper limit depends on the instance type and image, and the memory overhead must be estimated beforehand. For details, see Detailed description of RxCapacity. A smaller value increases the risk of packet loss under burst traffic. Setting a value lower than the default is generally not recommended.

Receive buffer merging mode (Mergeable Rx Buffers)

Display name

Mergeable Rx Buffers

Default value

0 (Disabled)

Valid values

1 = Enabled; 0 = Disabled

Function: controls how receive buffers are allocated. By default (disabled), each receive buffer is allocated at the full packet size (about 72 KB). When enabled, each buffer is a single memory page (about 4 KB), and large packets are received across multiple buffers and merged by the driver before being delivered to the operating system. Network functionality and packet content are not affected. This essentially provides a way to increase RxCapacity at a low memory cost: with the same RxCapacity, receive memory usage is reduced to about 1/18 of the default mode, and with the same memory budget, the buffer pool depth becomes about 18 times greater. Applicable to memory-sensitive scenarios that require a relatively deep receive buffer pool.

Note: If throughput decreases after you enable this parameter, increase RxCapacity and observe, or disable this parameter to revert.

VLAN and priority

Controls whether the driver inserts or strips 802.1Q VLAN tags in transmitted/received Ethernet frames. When enabled, the transmit direction inserts a 4-byte VLAN tag (containing the priority and VLAN ID) into the frame header, and the receive direction strips the VLAN tag from frames before reporting them to the operating system.

The three parameters form a hierarchical control relationship:

Layer

Parameter

Function

Master switch

Init.Do802.1PQ

When disabled, all lower-level parameters become ineffective

Feature selection

Priority and VLAN tagging

Uses bits to select enabling priority (bit0) and/or VLAN ID (bit1)

VLAN ID

VLan ID

Specifies a concrete VLAN. When enabled, only packets of that VLAN are sent and received

Init.Do802.1PQ

Display name

Init.Do802.1PQ

Default value

1 (Enabled)

Valid values

1 = Enabled; 0 = Disabled

Priority and VLAN tagging

Display name

Priority and VLAN tagging

Default value

3 (All)

Valid values

3 = All (both priority and VLAN enabled); 2 = VLan (VLAN only); 1 = Priority (priority only); 0 = Disabled

VLan ID

Display name

VLan ID

Default value

0

Valid values

0–4094. 0 indicates that VLAN filtering is disabled

Recommendation: keep the default values. Whether this feature has an actual effect depends on whether the platform network supports VLAN passthrough. Under standard VPC networks, configuration is usually not required.

MTU (Jumbo Packet)

Display name

Jumbo Packet

Default value

1514

Valid values

An integer between 590 and 65500 (step: 1)

Function: the maximum length of Ethernet frames sent and received by the NIC. This value includes the 14-byte Ethernet header, that is, IP-layer MTU = this parameter − 14. The default value 1514 corresponds to the standard MTU 1500.

Note

The actual effect of this parameter depends on whether the platform network layer delivers an MTU configuration. If the platform has delivered one, the driver directly uses the MTU value provided by the platform, and your modification in Device Manager does not take effect. If the platform has not delivered one, the value configured here is used as the actual MTU.

Recommendation: keep the default value.

MAC address

Assign MAC

Display name

Assign MAC

Default value

Empty (not set; uses the MAC assigned by the platform)

Valid values

A 12-character hexadecimal string (such as 021122334455). Optional parameter

Function: overrides the MAC address currently used by the NIC. After the setting takes effect, both the source MAC of transmitted packets and the receive filtering use the new address. The address must be a unicast locally administered address (the lowest 2 bits of the first byte = 10, that is, the first byte is 02/06/0A/0E/...). If the configured address is invalid, the setting does not take effect, and the NIC continues to use the MAC address assigned by the platform.

Recommendation: keep the default value.

Logging and debugging

Display name

Logging.Enable, Logging.Level

Default value

Logging.Enable = 1; Logging.Level = 0

Used for internal driver diagnostics. Keep the default values.