How To Troubleshoot Microsoft Teams with Call Quality Dashboard (CQD), Hands-on with Microsoft
Victor Guzman and Siunie Sutjahjo from Microsoft on troubleshooting Teams call quality with CQD: the seven-step Get to Green checklist, the QER 5.3 Power BI reports to start with, and how to read transport, VPN, weekly metrics and per-user data.
Prefer audio?
Who should listen: Teams service owners, M365 admins and network teams who own call quality, and anyone who has downloaded the CQD Power BI templates and not known where to start.
Guests: Victor Guzman — Senior Technical Program Manager, Microsoft · Siunie Sutjahjo — Principal Product Manager, Microsoft Teams Meeting, Microsoft
Microsoft's Get to Green programme was the team customers were escalated to when Teams call quality went wrong, and Victor Guzman spent five years running its quality reviews. In this episode he shares the seven-step checklist that resolved the vast majority of those escalations, then opens the CQD QER 5.3 Power BI templates on a real tenant and walks through the reports he actually uses: usage summary, media health, transport, estimated VPN, weekly metrics and the per-user search. Siunie Sutjahjo adds how customers shape the templates and how to turn the data into a conversation the network and security teams will act on.
Many thanks to Barco, sponsor of this episode.
Key insights
- Get to Green's quality review started with a seven-step checklist, and Guzman reckons around 99% of the escalations his team worked were resolved by following it. The rest of the session is about using CQD to check whether each step is actually true in your tenant. ▶ 4:52
- This is not a one-off exercise. Usage patterns have shifted from audio-only on managed office networks, to working from home with VPN problems, to a return to offices that were never sized for video and now give a worse experience than home. Review quality on a cadence and after every change in how people work. ▶ 6:20
- Step one: ports and protocols. Media services are now consolidated to three subnets, with UDP 3478 to 3481 and TCP 443. In many configurations all media types go to 3478 rather than mapping neatly to each port, so open the range to the published subnets rather than assuming per-modality ports. ▶ 7:30
- Step two: bypass proxies and deep packet inspection. Proxies are not sized for real-time media, they queue packets, and inspection cannot decrypt the media anyway, so all it does is add delay and saturate the proxy during video calls. ▶ 8:27
- Step three: split tunnel the media. Teams wants UDP, where a lost packet is simply lost and the codec copes. Forcing media into a VPN tunnel wraps UDP in TCP, adds retransmissions, latency and jitter, and gains nothing for security. Web traffic can stay on the VPN; media should not. ▶ 9:09
- Steps four and five: resolve DNS locally and take the shortest path to the internet. Microsoft uses geo-based DNS to pick the nearest media service, so users in Brazil or France resolving from the US end up in meetings hosted far away. Egress locally to a provider that peers with Microsoft, so media is on the Microsoft network within one or two hops. ▶ 10:20
- Step six: QoS only if you need it. It helps solely where there is congestion on your own network and does nothing for home users. Microsoft stopped recommending it for everyone after seeing badly configured policies that were fine until the reserved bandwidth was hit and then dropped packets even with capacity spare. ▶ 13:06
- Step seven: exclude the Teams processes from antivirus and DLP scanning. Scanning encrypted media adds load and quality loss for no security gain. Microsoft has worked with the antivirus vendors so it bites less than it did, and the Teams security guide sets out the trade-off if the security team wants the reasoning. ▶ 14:50
- Use the QER 5.3 templates or later, because the classifiers changed and older versions do not benefit. The pack has 45 report pages and is deliberately a template: Sutjahjo describes customers removing weekends from their quality percentage, or scoping the whole pack to the region or subnets an admin is responsible for. ▶ 16:40
- Start with the usage summary and note the audio, video and sharing mix, Wi-Fi versus wired versus mobile, and conference versus peer-to-peer (typically 80 to 85% conference). Then check meetings hosted by country: a US tenant whose meetings are hosted in France, Ireland and the Netherlands has a DNS or geolocation problem. ▶ 22:18
- The feedback rate is user perception rather than a measurement, and its usefulness varies: Japanese users are strict, other organisations barely answer the survey. Go to the feedback report, find the users giving the poorest ratings and contact them. Guzman recalls one VIP whose bad ratings came down to a damaged connector, and Sutjahjo notes the one persistent complainer is often speaking for a whole floor. ▶ 25:45
- The media health dashboard is the report to check daily or weekly. For a global company Microsoft's target is under 3% poor audio, video and sharing; a managed network in well-connected countries can aim for 2% or 1%. Setup failure, where the call never connects, should be under 1%. The demo tenant sat at 3.36% poor audio, 4.98% video and 3.58% setup failure, and the trend lines are worth reading for a step change that points to a network change. ▶ 28:40
- The transport report showed only 23% of the demo tenant's media on UDP, with most of it on Compound TCP, which means it is going through an HTTPS proxy. Those streams ran at over 7% poor, against about 6% overall, with jitter and packet loss peaks to match. Compound TCP is the worst case, then TURN TCP, then multi-TURN TCP; the only fix is opening UDP. ▶ 33:15
- Import your building data into CQD. It is tedious and coverage is never complete, but the reports become far more useful when they show office and network names instead of bare subnets that only the network team can read. ▶ 36:07
- Estimated VPN flags a session as probable VPN when the subnet mask is a single address. In the demo, users not on VPN ran about 6% poor and 0.31% setup failure; VPN users were 12% poor and 2.26% setup failure. Guzman has seen India-to-India calls hairpinned through US concentrators with two to three seconds of latency. The same applies to Zscaler, Cisco Umbrella and similar services: send web traffic there, not media. ▶ 39:50
- Weekly metrics is Guzman's most-used report because it exposes raw jitter, latency and packet loss over time, by country, ASN, public IP and subnet. Drill from a week to a day to an hour to the individual conference IDs. Patterns emerge: Mondays bad by hour, or US night-time problems that turn out to be India users routing via the US. ▶ 42:30
- The search experience takes a meeting ID, a UPN, a subnet or a public IP and drills through to the meeting health detail: who dropped, their packet loss, where they connected from. It settles whether the all-hands was bad or bad for one person, and the new dominant participant classifier reflects a presenter's problems onto everyone who heard them, so the room's complaints line up with the cause. ▶ 48:25
- Check client versions. Sutjahjo has dealt with large enterprises running clients two years old. Version used to be on the checklist and was removed when new Teams started blocking very old clients, but it is still worth a look. ▶ 56:05
- Take evidence, not opinion, to the network team. In Sutjahjo's words, if you are going to court you need the documents. Tom's version: it is never the network team until you can prove it. Inbound or outbound, which IPs, how many users affected, and a before-and-after from CQD's year of history is what gets a change made and a quick win recognised. ▶ 58:15
Insights summarised by AI from the episode transcript, reviewed by the Empowering.Cloud team.
Listen: Apple Podcasts · Spotify · Other platforms
Comments ()