ROS stop L3 forwarding for 2-3 minutes

Who has seen this?

ROS on different platforms stops L3 routing for 2-3 minutes. After this everything will return normaly. There are no log events, In the error time we could log in via mac telnet. There we could se that there is no high cpu and the connection count (DOS). We could ping from the router to different ip outside of routers subnet but no ping from outside (also from same subnet) is possible…

if you’re using bridges then you should configure admin-macs.
maybe it helps.

Thanks sup,

but we do not use bridges. As I said: L3 forwarding is gone. L2 still works…

No one else with this problem?

Post the routing table for the affected devices please

Is there a dynamic routing protocol running on this device?

Yes OSPF is in use.
Routes looks good.

For example:

Interface VLAN 11 is on 10.0.0.42/24
Other router with 10.0.0.30/24 can’t ping 10.0.0.42. Access with mac-telnet to the .42 router is possible
Ping from .42 (mac telnet) to .30 is working. Only inbound is not working
After 2 minutes the problem is gone and everything is working as normal

Are there any logs in the logging buffer when this occurs that show up?

Is there a pattern that can be repeated that you’ve seen that causes the layer 3 portion to stop forwarding?

Have you done any testing to repeat the issue?

There is nothing in the log. Absolute nothing.
The problem appears absolute at random times. RB1100, RB493AH and I386 are affected.
Some days silence and then it comes again.

Cause I don’t know what I should reproduce it is hard to reproduce :slight_smile:

How stable is the power to the box?

Not the issue:

3 Boxes at different locations (230V, POE 24V) …

Are you sure that the routes in the routing table are what you would expect during the event? I suggest enabling OSPF logging and having a look at the logs.

What could have a lower metric than an direct connected network? L3 is not working from same subnet 10.0.0.0/24…
I will activate OSPF logging and have a look…

It wasn’t clear to me that you were talking about directly connected networks. Which ROS version is this?

6.1

Maybe not impossible that you are seeing a 6.1 issue then…

If the problems tends to last 2 mins - 2 mins 30 secs and you have OSPF segments which have 30/120s Hello/Dead intervals then I would certainly give OSPF a check over.