Interface packet Drops on an octacore xeon with intel nics

Hi,

I cant understand why we’re getting rx drops. its an HP ML 350 Server with only intel pcixpress gigabit ethernet nics. low traffic and we’re getting rx drops

i tried to enable RPS on the interfaces that are getting drops, fixed one cpu to each card that having drops and still getting drops :frowning:


any idea how to fix it?
packetloss.JPG

just to keep you informed , i found out that the problem is caused by OSPF BUG in rc01 - rc02 prerelease. When i disabled ospf no more RX drops!

Hi,

Can you post PCI tab with IRQs?

What is the model of motherboard?

it is a HP Proliant ML 350

DEVICE VENDOR NAME IRQ

0 01:04.6 Hewlett-Packard Company Proliant iLO2 virtual UART (... 11
1 01:04.4 Hewlett-Packard Company Proliant iLO2 virtual USB co... 11
2 01:04.2 Compaq Computer Corporation Integrated Lights Out Proce... 11
3 01:04.0 Compaq Computer Corporation Integrated Lights Out Contro... 11
4 01:03.0 ATI Technologies Inc ES1000 (rev: 2) 3
5 03:00.0 Broadcom Corporation NetXtreme II BCM5708 Gigabit... 7
6 02:00.0 Broadcom EPB PCI-Express to PCI-X Bri... 0
7 14:01.1 Intel Corporation 82546GB Gigabit Ethernet Con... 7
8 14:01.0 Intel Corporation 82546GB Gigabit Ethernet Con... 4
9 13:08.0 Hewlett-Packard Company Smart Array E200i (SAS Contr... 4
10 13:04.0 Broadcom BCM5785 [HT1000] PCI/PCI-X B... 0
11 12:00.0 Broadcom EPB PCI-Express to PCI-X Bri... 0
12 0f:00.1 Intel Corporation 82571EB Gigabit Ethernet Con... 4
13 0f:00.0 Intel Corporation 82571EB Gigabit Ethernet Con... 10
14 09:00.0 Intel Corporation 82572EI Gigabit Ethernet Con... 5
15 06:00.0 Intel Corporation 82572EI Gigabit Ethernet Con... 7
16 05:01.0 Intel Corporation 6311ESB/6321ESB PCI Express ... 255
17 05:00.0 Intel Corporation 6311ESB/6321ESB PCI Express ... 255
18 04:00.3 Intel Corporation 6311ESB/6321ESB PCI Express ... 0
19 04:00.0 Intel Corporation 6311ESB/6321ESB PCI Express ... 0
20 00:1f.2 Intel Corporation 631xESB/632xESB/3100 Chipset... 5
21 00:1f.0 Intel Corporation 631xESB/632xESB/3100 Chipset... 0
22 00:1e.0 Intel Corporation 82801 PCI Bridge (rev: 217) 0
23 00:1d.7 Intel Corporation 631xESB/632xESB/3100 Chipset... 7
24 00:1d.3 Intel Corporation 631xESB/632xESB/3100 Chipset... 4
25 00:1d.2 Intel Corporation 631xESB/632xESB/3100 Chipset... 10
26 00:1d.1 Intel Corporation 631xESB/632xESB/3100 Chipset... 5
27 00:1d.0 Intel Corporation 631xESB/632xESB/3100 Chipset... 7
28 00:1c.0 Intel Corporation 631xESB/632xESB/3100 Chipset... 255
29 00:16.0 Intel Corporation 5000 Series Chipset FBD Regi... 0
30 00:15.0 Intel Corporation 5000 Series Chipset FBD Regi... 0
31 00:13.0 Intel Corporation 5000 Series Chipset Reserved... 0
32 00:11.0 Intel Corporation 5000 Series Chipset Reserved... 0
33 00:10.2 Intel Corporation 5000 Series Chipset FSB Regi... 0
34 00:10.1 Intel Corporation 5000 Series Chipset FSB Regi... 0
35 00:10.0 Intel Corporation 5000 Series Chipset FSB Regi... 0
36 00:05.0 Intel Corporation 5000 Series Chipset PCI Expr... 0
37 00:04.0 Intel Corporation 5000 Series Chipset PCI Expr... 0
38 00:03.0 Intel Corporation 5000 Series Chipset PCI Expr... 0
39 00:02.0 Intel Corporation 5000 Series Chipset PCI Expr... 0
40 00:00.0 Intel Corporation 5000Z Chipset Memory Control... 0

[admin@mk_brd01] /system resource pci>

Thank you, you are fast.

Maybe it is good idea to try to fix this 7 and 4 irq conflicts to avoid rx drops and other problems ?

I have a similar problem http://forum.mikrotik.com/t/cpu-load/44622/1

I’m sorry, but your “RX-Drop-Phobia” is just funny at best.

As far as i know there are 2 possible reasons for “RX Drops”:

  1. Driver drops packet cause it is convinced that packet is unnecessary/unusable (usually management frames like “pause” frames, multicast frames, frames that arrived 2nd time, frames with wrong header info etc) - basically driver is doing good thing by keeping out the bad stuff from your software

  2. Drivers rx buffer is full, it can’t receive more frames, cause there are not enough resources to process already received frames and empty the buffer.

Usually 2nd option are indicated by 6 7 8 digit numbers in RX drops over 24h

Thanks for the info. Maybe you can post ‘TX Drop’ reasons? on VLAN interface, I have about 100 dropped TX packets per second at ~130 Mbps TX rate. Ethernet drop/error counters are at zero

UPD: seems like problem was solved by increasing ‘ethernet-default’ queue length from 50 to 200 packets %)

UPD2: from 200 to 400 :smiley:

I have this problem too. x86, intel 82572EI. increazing ethernet default to 400 do not help

anybody have a solution . Using 82576 too many 200000 packets in and out . Dual quad core (8cores) cant handle it.

You need to make irqs and cores balance and then you will not have this problem in v5

If you have more lanes (pci-e) per one ethernet that is the better and you can setup it so packet drops will never appear.

I have an 8 core system
4 x 82576 gige Ports - shows 7rx and 7tx each ethernet auto (like eth2-tx-rx-0 there are 7 of these , similarly there are 14 for each 82576 eth )
2x onboard broadcom 5708 (this is pci-x bridge but im not using this anymore)

How do I do it in v5 . Have v5.4 right now installed. I want it to scale to 500k or more pps how can this be done ?
Where do I go what do I do to put in more lanes
Would getting an opteron 12 + 12 core = 24 core system help me process more even though they have a lower clock speed?

You should have 8 total.

Go in resources>irq and put different core for every lane (eth2-tx-rx-0 … eth2-tx-rx-1 .. eth2-tx-rx-2 etc). Put auto for some and for some put fixed core like for eth2-tx-rx-XX (XX is number, 0-3 fixed and 4-7 put auto).

Do that for every ethernet and you should be fine.

Done it let me watch tomm. no significant change so far. Right now the heavy packet senders are on a separate router but strangely on that I have 4 cores and it shows 8 streams per ethernet (same 82576 Intel) . It shows 4 streams of rx (eth0-rx-0 to 3) and 4 streams of tx ( eth0-tx-0 to 3) which means I have 4 ports (eth0,1,2,3) = 32 streams for all 4 eth

While on the 8 core system i have (eth0-txrx-0 to 7) 8 streams . Still total works out to 32 streams (4 eth * :sunglasses:. But the naming of ethX-txrx-X vs ethX-rx-X + ethX-tx-X why is there a difference ?

Does changing the no. of packets (pfifo) in ethernet-default (in queuing) also affect anything? It was 100 changed to 100000

Hasnt made much of a difference my cpu loads are still the same.

Did you turn off RPS and maybe try the multi-queue ?

Maybe you try to install linux ?? gentoo or ubuntu… On 8 core Xeon on mikrotik (2 Gb/s traffic, 600000 pps) without firewall and queue 15-20% cpu load. On same system and same traffic but with linux cpu load ~ 1%

With firewall (8000 rules) and queue (queue tree 8000 rules) on mikrotik cpu load - 50-70% and TX/RX drops on interface.
With queue (12000 rules) on linux cpu load ~ 2-4% and no TX/RX drops.

Question - maybe mikrotik need to change something ??

Martini, that kind of traffic could not be utilising the CPU only 1% ? Maybe Linux need to change something.

1500 users going through this router, why this traffic cant utilising CPU for 1% ??

@shaper172:~# uptime
15:15:24 up 14 days, 14:16, 1 user, load average: 0.00, 0.01, 0.05


bwm-ng v0.6 (probing every 0.500s), press 'h' for help
input: /proc/net/dev type: rate
| iface Rx Tx Total

eth1: 247.39 Mb/s 923.58 Mb/s 1.14 Gb/s
eth2: 923.44 Mb/s 246.43 Mb/s 1.14 Gb/s

total: 1.14 Gb/s 1.14 Gb/s 2.29 Gb/s


top - 16:09:59 up 14 days, 15:11, 1 user, load average: 0.04, 0.03, 0.05
Tasks: 84 total, 1 running, 83 sleeping, 0 stopped, 0 zombie
Cpu0 : 0.0%us, 0.0%sy, 0.0%ni, 99.6%id, 0.0%wa, 0.0%hi, 0.4%si, 0.0%st
Cpu1 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Cpu2 : 0.0%us, 0.0%sy, 0.0%ni, 99.7%id, 0.0%wa, 0.0%hi, 0.3%si, 0.0%st
Cpu3 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Cpu4 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Cpu5 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Cpu6 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Cpu7 : 0.0%us, 0.0%sy, 0.0%ni,100.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Mem: 4123928k total, 404336k used, 3719592k free, 90460k buffers
Swap: 4190204k total, 0k used, 4190204k free, 230644k cached

PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
1 root 20 0 3184 1800 1260 S 0 0.0 0:01.10 init
2 root 20 0 0 0 0 S 0 0.0 0:00.00 kthreadd


eth1 Link encap:Ethernet HWaddr 00:1b:21:92:01:d0
inet addr:10.10.7.7 Bcast:10.10.7.255 Mask:255.255.255.0
inet6 addr: fe80::21b:21ff:fe92:1d0/64 Scope:Link
UP BROADCAST RUNNING MULTICAST MTU:1500 Metric:1
RX packets:46726066747 errors:0 dropped:0 overruns:0 frame:0
TX packets:47601321122 errors:0 dropped:0 overruns:0 carrier:0
collisions:0 txqueuelen:1000
RX bytes:31720165396132 (31.7 TB) TX bytes:50510319588460 (50.5 TB)
Memory:d8800000-d8820000

eth2 Link encap:Ethernet HWaddr 00:1b:21:92:01:d1
inet addr:10.10.8.8 Bcast:10.10.8.255 Mask:255.255.255.0
inet6 addr: fe80::21b:21ff:fe92:1d1/64 Scope:Link
UP BROADCAST RUNNING MULTICAST MTU:1500 Metric:1
RX packets:47601307667 errors:0 dropped:0 overruns:0 frame:0
TX packets:46722252334 errors:0 dropped:0 overruns:0 carrier:0
collisions:0 txqueuelen:1000
RX bytes:50527044557868 (50.5 TB) TX bytes:31642723711897 (31.6 TB)
Memory:d8820000-d8840000




Any questions ???

When you tested the same with RouterOS, did you turn off RPS ? If not - please can you test again?

And make a screenshot of Tools->Profile. As well as a supout to eventually send to support.

Under Linux, can we see how many CPU are the ethernet controllers using, for interrupt requests etc ?

P.S. Please can you show the exact linux version you are using, exact Kernel and exact drivers with exact NIC models. Thank you.