AX S Manufacturing Issue?

Initially, I did not see any of the hAP ax S Wi-Fi issues that some people were reporting. I mostly chalked them up to configuration problems or issues involving features that we do not use.

Now that we have rolled out more of these devices, however, we have started receiving reports of problems. What we are seeing appears to paint a different picture from what has previously been reported, and I am beginning to think that at least some of the earlier issues may not have been correctly identified.

For context, we currently have a few hundred hAP ax S routers installed in production and have recently shipped several hundred more. We use completely standardized and automated configurations, and every device is running the same RouterOS version. Aside from the SSIDs and PPPoE credentials programmed into each device, the configurations are otherwise identical.

Using essentially the same deployment model, we also have more than 10,000 older hAP ac lite, ac², and ac³ series routers installed and in use today. At that volume, we have an opportunity to identify equipment patterns that may be difficult for individual users or smaller deployments to detect.

What we have now observed is that a subset of hAP ax S devices, all of which so far appear to have been shipped during April and May of this year, exhibit the same highly repeatable behavior. They may also all have serial numbers beginning with “HM,” although we only identified this pattern yesterday and are still verifying the affected range. UPDATE: We have now identified several devices with serial numbers starting with HJ that are also exhibiting the same problem, though they share the same production / shipping timeframe.

On the affected routers we found and have tested, all running 7.20.8LT but also confirmed the same behavior on 7.21.5LT, sustained Wi-Fi traffic above approximately 30 Mbps causes the router to reboot from the watchdog timer.

If Wi-Fi traffic remains below roughly 30 Mbps, the router can stay online for days or weeks. During that time, it can pass approximately 1 Gbps through the Ethernet ports without any apparent problem.

If the watchdog timer is disabled, the router no longer reboots at 30 Mbps of wireless traffic. It will instead pass up to approximately 150 Mbps over Wi-Fi before locking up completely and requiring a physical power cycle.

At sustained wireless loads somewhere between approximately 30 and ~100 Mbps, the wireless interfaces will instead disappear after several minutes and stop broadcasting, almost as though the radio had been physically removed from the router. This is again with the watchdog timer disabled.

The behavior changes significantly when the same affected router is upgraded to RouterOS 7.23.2. On 7.23.2, it remains stable and no longer crashes or triggers the watchdog timer, but Wi-Fi throughput appears to be capped at approximately the same 150 Mbps limit that crashed the non-watchdog enabled LT version regardless of the client or test conditions, though unlike on the Long Term builds, it is stable at that 150 Mbps speed.

If I replace that router with another hAP ax S running the exact same configuration and RouterOS version, but originating from a different shipment or manufacturing period, the replacement reliably passes approximately 600 Mbps over Wi-Fi in the same location, using the same client and test methodology.

In other words:

  • The failure follows the physical router.

  • The configuration and RouterOS version are identical.

  • Ethernet forwarding remains stable at approximately 1 Gbps.

  • Older RouterOS versions crash or lose the wireless subsystem under 30 Mbps load with watchdog timer enabled, and 150 Mbps with watchdog timer disabled.

  • RouterOS 7.23.2 prevents the crash, but the affected hardware remains limited to approximately 150 Mbps.

  • Other hAP ax S units manufactured at different times (locations?) pass approximately 600 Mbps under the same conditions.

My current theory is that particular production batches of hAP ax S routers manufactured or shipped (to us) around April and May may contain either a physical defect or a component variation within the wireless subsystem. This could involve the MediaTek radio, RF front end, antenna-chain circuitry, radio power delivery, calibration data, or a different component or silicon revision used during that production run.

One possibility is that one or more RF chains are defective or being disabled after the driver detects a fault. However, I do not think the throughput result alone is sufficient to prove that this is specifically a “single working chain” problem. A 150 Mbps ceiling could also result from a fallback operating mode, repeated transmission errors, a bus or DMA problem, thermal or power instability, or a driver workaround that deliberately limits the affected radio to avoid the complete lockups seen on earlier releases.

Whatever changed between the earlier RouterOS versions and 7.23.2 appears to have hardened the driver or operating system against the failure. It may prevent the watchdog reboot and complete lockup, but it does not appear to correct the underlying cause, because the same affected routers remain dramatically slower than otherwise identical units.

These findings are still preliminary. We identified the apparent shipment and serial-number pattern only yesterday, after spending some time trying to determine why these routers were repeatedly rebooting. We are now reviewing additional units and deployment records to establish whether all affected devices fall within the same production range.

I am posting this because the symptoms may have been interpreted as unrelated configuration or driver problems when they could instead point to a specific hardware batch or component revision. I would be interested to know whether MikroTik can identify any production, board, radio, or component changes associated with the affected serial numbers, and whether other users seeing similar problems can compare their serial-number prefixes and manufacturing dates.

When I read "we have hundreds devices deplyed" and "problems" the first thought is that you have to contact your distributor or/and Mikrotik directly as we, forum users, have no power to help.
Mikrotik follows more or less (!!!) news on the forum but the only official way is to raise a ticket in Mikrotiks's support system.

Try this: https://mikrotik.com/support

I think you're missing the point of Brian’s post.

He isn't really asking forum users to fix the problem, but rather trying to find out whether others can confirm the same symptoms

Of course, opening a support ticket with MikroTik makes sense, but his post and the user forum serve a different purpose here: gathering input from others who may have experienced similar issues.

Trust me I'm very familiar with opening support tickets, I've been doing this for more than 20 years (check my profile date :wink:). I posted this because other users have commented about AX S wifi problems, and I suspect that some of them may have not been config issues or driver issues like they thought, but instead I think we may have a bad batch of hardware, and the best way to confirm that is to widen the scope here to see if other users can contribute additional observations to help get a clearer picture. I didn't notice this pattern until yesterday, and that's with "hundreds of AX S devices deployed", if we can get information from other users that help add more detail to the picture or widen / bookend the scope of the problem, then it will help everyone including MikroTik to understand exactly what is and isn't affected.

Possible, but when I see:

it makes me think that the question is directed to the Mikrotik, however the rest of sentence:

gives users a chance to corelate their problems with AMPLIFIED scale of @BrianHiggins deployment, not asking for help.

It's obvious that forum is a place for comments/help/complains/sharing ideas etc. but the OP post is more a kind of complain to the Mikrotik than opening a discussion, so finally I decided to suggest opening a ticket. Especially that MT is quite resistant to react to complains raised on the forum as they expressed it lots of times.

You think that sounds like me complaining? Clearly you didn't read any of the Device-Mode hell saga thread, if you think that I'm complaining here... Your account is old enough that I'm going to give you the benefit of the doubt here and assume you really meant well.

However I do want to just put into perspective that I was the 13th person in the US certified on MikroTik, I was deploying MT gear for almost a full decade before you created your login here, and drinking beer and smoking cigars with Normis Janis and John T at USA MUM events (lets not talk about the Jager Bombs in Orlando) 7-8 years before you first logged into this forum. I have been deploying MikroTik gear since before this forum existed, was even one of the first people in the USA to purchase MikroTik manufactured hardware back when they only made 3 models of open circuit board and you had to buy the cases separately from 3rd party companies, and I've been deploying it ever since. I know I don't come to the forums as regularly as I used to, but I do appreciate the effort to help educate people and be helpful, but I also don't appreciate the condescending judgement about what you interpreted my intentions to be, @Larsa was exactly correct with their earlier response to you, and I followed that up with confirmation and additional context just in case that wasn't clear. I don't know what you think you're contributing at this point, but the fact is we have a collective interest in identifying a possible pattern of problem with this hardware, and the best way to accomplish that is getting more eyeballs on the problem and reports of shared experiences.

My initial interpretation of the problem might be incorrect, maybe it's more widespread than I suspect, maybe someone else can identify a pattern that I didn't, or potentially and best case, we can find out a way to definitively identify which routers are affected and which ones aren't. But what remains true regardless of the final answer, and even though I might have bought the majority of the AX S inventory that's been received into the US over the last couple months, I am still limited by the data I have available to me, and I didn't buy all the AX S inventory, so there's hardware sitting out there in people's hands that I can't see if it's affected or not.

Your explanation of there being some manufacturing difference actually makes a lot of sense. I'm mainly following the discussion around the wifi performance of the ax S out of curiosity, but the reports seem to be strangely bimodal, with 5x or so difference in performance.

If this was my product, I'd request that one of each device (one that performs well and one that's predictably problematic) be shipped back for analysis. I know MT has done similar things in the past.

From your description I'd assume a thermal problem right away. Cooling RF chips is usually hard from a mechanical engineering perspective. Maybe the difference is there? What really points to this is the delayed onset of the problem, as in "sustained throughput"...

Yeah, thermal issues definitely sound like a very plausible cause.

I don’t know what the hAP ax S looks like inside, but poorly mounted heatsinks, missing or incorrectly applied thermal paste, could potentially explain it.

At first sight the design seems prone to thermal paste issues:

I opened the case up on one here in the lab yesterday wondering if there could be a loose antenna wire or something, I can say that mine had much better (complete) contact of the thermal paste than the one in those photos had. There might be something to that idea.

I'll see if I can get someone onsite to the problems devices (I'm troubleshooting from 1.5k miles away) can open one and check the seating of that heatsink. Already know that'll have to wait until next week though, he left for the weekend already.

As you probably already saw in the post jaclaz linked to, it's one very large and substantial heatsink for the whole bottom of the router. Quite a nice design in that regard, but it does seem like it might be prone to poor contact as a result of the screws being under torqued resulting in the heatsink not being clamped to the bottom of the board. The board is sandwiched between the heatsink and the top of the case, it should be a pretty bulletproof design, but the one failure point I saw was insufficient torque could easily cause it to be left loose.

Once upon a time we were taught that a plane passes through three points not seven, so It could also be a planarity problem.

Though I personally hate thermal pads (that I also believe have worse thermal transmission) they would guarantee a more consistent thickness.

In assembly lines the thermal paste/silicone Is dispensed by high precision robots/dispensers, still ...

Thank you for providing this valuable statistics into the different hardware batches of hAP ax S. I have one and the serial number starts at HK. I never experienced any symptom as mentioned (router restarts on sustained wireless transfer around 30 Mbps). I'll be following this with interest and can provide some other observations as need arises.

/b

UPDATE!
I just got off the phone with the onsite tech who was doing the testing for me (as previously mentioned, I'm 1.5k miles away from where the hardware is installed), talking through everything again we realized that there was a possible fault in his testing after updating to v7.23.2 because he didn't confirm that he was in fact connected to the 5 GHz radio, and might have been connected to the 2.4 GHz network (and we do have DFS channels enabled, so there's a delay. He forgot to wait long enough after the reboot for 5 GHz to come online). He's going to go back onsite tomorrow morning to re-test the problem devices after validating that he is definitively connected to the 5 GHz radio and hope that we see different results.

I'm genuinely hoping that he was connected to the 2.4 GHz, because 150 Mbps is pretty much exactly what he should have seen there, and it opens the possibility that v7.23.2 actually does more than just address whatever manufacturing variance was causing some of these devices to crash / reset under load, and might have actually fully fixed the problem.

There 100% is some manufacturing variability between devices where certain devices have issues and others don't, but if v7.23.2 is capable of running stable across that variance, I'm happy, because that means no need to replace hardware, just update it.

Stay tuned, I'll add the results of the updated test as soon as I get them.

EDIT: revised test conducted and longer Note added below, but short version is that when connecting to the 5 GHz radio, it actually still crashes the router. It's the worst possible result....

I really like the one big heatsink design as well. Unfortunately these have a problem with maintaining thermal contact with many components at once, and that has to be solved in some way.

This is not primarily due to low torque or board flex. While jaclaz is totally right that a plane is described by three points, the PCB is actually quite flexible at this scale (tenths of mm over several cm), usually it's too high or uneven torque that causes a problem. The main reason why some sort of space-filing thermal solution has to be used is that during soldering, it's mainly the surface tension of the solder that pulls the components to their final position, and therefore this position is not exactly repeatable between assemblies - some sit higher, some lower, some with a bit of slope.

Mikrotik clearly uses some sort of specific goo to bridge this. I'm not familiar with this particular solution. You can also see what I was referring to as "especially difficult for RF chips" thing, in that they include an additional base elastic layer and a metal top for the shielding cages. Getting these arrangement to conduct heat reliably and pass emissions compliance tests is what nightmares are made of.

I was going to be very clever, and point out that the linked images show insufficient thermal contact, much like I would expect to see on the problem units - but as you have already pointed out, the person who made these photos is also experiencing problems.

If you look closely at the images, it can in fact be seen that the wifi chips that make poor thermal contact have a sort of ridge imprint on the thermal goo. This pattern does not correspond to anything on the mating side. Could this come from some sort of testing rig?

Finally: from your follow-on about possibly conducting the test on the 2.4GHz band, are you absolutely sure that on the problem devices, the 5GHz radio doesn't simply shut off, as in: even if the test is begun on 5GHz, is the device still on 5GHz at the end?

I concur. You can see how tiny (read: puny) the heatsink of the hex S 2025, which is an almost identical hardware compared to hAP ax S sans the wireless.

I was having so many hopes for this thread.

From the photos posted by Ca6ko on the referenced thread, the thickness of the "thermal goo" seems (to me) way too much, when compared on what is used on (say) PC CPU's or graphical cards CPU's.

Particularly this one:

We are talking of (seemingly) millimeter(s) vs. tenth(s) of millimeter (or even less).

The "disalignment" of the various components due to the surface tension you mention is anyway in the order of magnitude of tenth(s) of millimeter, not more.

But, given the position of the seven points of contact (5 main ones + 2 smaller ones) and of the five screws, I wouldn't be surprised if the board soft of pivots along a north/norteast to south/soutwest axis and the top left screw is the only thing that can force the contact between the heatsink plate and the two (I presume RF) chips, so an under torqued top left screw (or an over torqued central one) is IMHO likely to cause a less-than-optimal contact.

Getting the proper thermal exchange looks to me as an engineering nightmare.

The "thermal goo" seems a lot like silicone putty, like TG4040 pr TG6060:

https://www.tglobalcorp.com/products-detail/tg4040-putty/

or S-Putty/H-putty2:

https://lipoly.com/en/product/liquid-gap-fillers/thermal-putty/s-putty/

But this is only speculation, there might be other explanations, not connected to thermal issues.

Update, turns out it's worst possible scenario, he did in fact connect to the 2.4 GHz radio on the initial testing which was why he was limited to 150 Mbps, though stable. However when he connects to the 5 GHz radio it still crashes the router when running a speed test. :weary_face:

@lurker888 He went on site and did follow-up testing and confirmed that the 2 GHz side is actually fine it's the 5 GHz that is crashing even on v7.23.2

I'm going to see if he can open up the case on one of these problem routers to see if the heat sink is actually making contact