Initially, I did not see any of the hAP ax S Wi-Fi issues that some people were reporting. I mostly chalked them up to configuration problems or issues involving features that we do not use.
Now that we have rolled out more of these devices, however, we have started receiving reports of problems. What we are seeing appears to paint a different picture from what has previously been reported, and I am beginning to think that at least some of the earlier issues may not have been correctly identified.
For context, we currently have a few hundred hAP ax S routers installed in production and have recently shipped several hundred more. We use completely standardized and automated configurations, and every device is running the same RouterOS version. Aside from the SSIDs and PPPoE credentials programmed into each device, the configurations are otherwise identical.
Using essentially the same deployment model, we also have more than 10,000 older hAP ac lite, ac², and ac³ series routers installed and in use today. At that volume, we have an opportunity to identify equipment patterns that may be difficult for individual users or smaller deployments to detect.
What we have now observed is that a subset of hAP ax S devices, all of which so far appear to have been shipped during April and May of this year, exhibit the same highly repeatable behavior. They may also all have serial numbers beginning with “HM,” although we only identified this pattern yesterday and are still verifying the affected range. UPDATE: We have now identified several devices with serial numbers starting with HJ that are also exhibiting the same problem, though they share the same production / shipping timeframe.
On the affected routers we found and have tested, all running 7.20.8LT but also confirmed the same behavior on 7.21.5LT, sustained Wi-Fi traffic above approximately 30 Mbps causes the router to reboot from the watchdog timer.
If Wi-Fi traffic remains below roughly 30 Mbps, the router can stay online for days or weeks. During that time, it can pass approximately 1 Gbps through the Ethernet ports without any apparent problem.
If the watchdog timer is disabled, the router no longer reboots at 30 Mbps of wireless traffic. It will instead pass up to approximately 150 Mbps over Wi-Fi before locking up completely and requiring a physical power cycle.
At sustained wireless loads somewhere between approximately 30 and ~100 Mbps, the wireless interfaces will instead disappear after several minutes and stop broadcasting, almost as though the radio had been physically removed from the router. This is again with the watchdog timer disabled.
The behavior changes significantly when the same affected router is upgraded to RouterOS 7.23.2. On 7.23.2, it remains stable and no longer crashes or triggers the watchdog timer, but Wi-Fi throughput appears to be capped at approximately the same 150 Mbps limit that crashed the non-watchdog enabled LT version regardless of the client or test conditions, though unlike on the Long Term builds, it is stable at that 150 Mbps speed.
If I replace that router with another hAP ax S running the exact same configuration and RouterOS version, but originating from a different shipment or manufacturing period, the replacement reliably passes approximately 600 Mbps over Wi-Fi in the same location, using the same client and test methodology.
In other words:
-
The failure follows the physical router.
-
The configuration and RouterOS version are identical.
-
Ethernet forwarding remains stable at approximately 1 Gbps.
-
Older RouterOS versions crash or lose the wireless subsystem under 30 Mbps load with watchdog timer enabled, and 150 Mbps with watchdog timer disabled.
-
RouterOS 7.23.2 prevents the crash, but the affected hardware remains limited to approximately 150 Mbps.
-
Other hAP ax S units manufactured at different times (locations?) pass approximately 600 Mbps under the same conditions.
My current theory is that particular production batches of hAP ax S routers manufactured or shipped (to us) around April and May may contain either a physical defect or a component variation within the wireless subsystem. This could involve the MediaTek radio, RF front end, antenna-chain circuitry, radio power delivery, calibration data, or a different component or silicon revision used during that production run.
One possibility is that one or more RF chains are defective or being disabled after the driver detects a fault. However, I do not think the throughput result alone is sufficient to prove that this is specifically a “single working chain” problem. A 150 Mbps ceiling could also result from a fallback operating mode, repeated transmission errors, a bus or DMA problem, thermal or power instability, or a driver workaround that deliberately limits the affected radio to avoid the complete lockups seen on earlier releases.
Whatever changed between the earlier RouterOS versions and 7.23.2 appears to have hardened the driver or operating system against the failure. It may prevent the watchdog reboot and complete lockup, but it does not appear to correct the underlying cause, because the same affected routers remain dramatically slower than otherwise identical units.
These findings are still preliminary. We identified the apparent shipment and serial-number pattern only yesterday, after spending some time trying to determine why these routers were repeatedly rebooting. We are now reviewing additional units and deployment records to establish whether all affected devices fall within the same production range.
I am posting this because the symptoms may have been interpreted as unrelated configuration or driver problems when they could instead point to a specific hardware batch or component revision. I would be interested to know whether MikroTik can identify any production, board, radio, or component changes associated with the affected serial numbers, and whether other users seeing similar problems can compare their serial-number prefixes and manufacturing dates.