AV, AI and Meeting Equity - Multi-Camera, Multi-Agent, Multi-AI with Bo Pintea
Bo Pintea of Huddly, previously a long-standing Microsoft veteran across Exchange, the RTC/OCS/Lync lineage and Skype, sets out the architectural case for IP-connected, edge-processed multi-camera rooms. He covers the human science behind shot selection and hyper-gaze fatigue, practical camera layou
Prefer audio?
Who should listen: AV designers, integrators and Microsoft 365 service owners specifying Teams Rooms and multi-camera or 360 camera deployments, plus anyone weighing edge versus codec processing, room audio quality for transcription, or the data governance implications of always-on meeting AI.
Guest: Bo Pintea — Industry and product strategy (ex-Microsoft, 25 years across Exchange, RTC/OCS/Lync and Skype), Huddly
Bo Pintea of Huddly, previously a long-standing Microsoft veteran across Exchange, the RTC/OCS/Lync lineage and Skype, sets out the architectural case for IP-connected, edge-processed multi-camera rooms. He covers the human science behind shot selection and hyper-gaze fatigue, practical camera layout guidance, and where AI agents, corporate knowledge and data ownership go next.
Many thanks to AVI-SPL, sponsor of this episode.
Key insights
- Three trends run in parallel, not just multi-camera: IP connectivity (most cameras still hang off USB to the codec), edge AI, and outside-in versus inside-out placement. Pintea expects the conversation to shift from 'multi-camera' to 'multi audio-video perceptual nodes' around the room — capturing and playing back audio on the perimeter — from roughly 2027-28 onwards. ▶ 4:41
- The prevailing multi-camera model sends every stream to the MTR compute for compositing; Huddly's position is that processing must happen on the camera itself, because you cannot push six to ten 6K/8K streams into any room compute without it drowning in data. Scalability to 5-10 cameras is the deciding factor. ▶ 6:12
- Three named in-room experience problems drive the design: mirror anxiety from the self-view screen, 'hyper-gaze' (a wall of faces staring back, when the brain can only process two or three streams at once), and feeling trapped in a single camera's field of view. The answer is a curated sequence of shots that preserves spatial relationships, not 20 simultaneous streams gridded out. ▶ 8:50
- A single 360 camera gives a cylindrical, effectively two-dimensional view — like turning your head with one eye. Adding a front-of-room bar only gets you to '2.5D' because both sit on the same optical axis, so remote viewers can't judge distance or line of sight. Overlapping fields of view from separate positions are what enable 3D reconstruction (SLAM), head pose estimation and 'who is looking at whom'. ▶ 11:55
- Practical five-camera layout guidance: one front-of-room camera above the display (with adaptive perspective correction so it can be mounted high, or dropped from a ceiling mount, without CCTV-style distortion) for the overview shot; cameras two and three roughly three metres in looking back into the room; cameras four and five roughly six metres in looking forwards — the goal being every person covered by at least two cameras. Pintea concedes the industry, including Huddly, has not educated the market well on this. ▶ 14:26
- Device silicon is becoming a real spec: Huddly has moved from an Intel NPU to a more powerful Ambarella part in the C1 to leave headroom for running more models at the edge. Pintea argues TOPS should be in buyers' evaluation matrices but currently isn't — and that you can't know what LLM capability you'll want to run on the device in two or three years. ▶ 16:57
- Audio is now processed in the visual domain — a spectrogram analysed as an evolving image — so advances in visual understanding transfer directly to audio, replacing hardcoded DSP pipelines with learned models. Combined with beam-forming mic arrays, the remote audio experience can be better than being in the room: someone ten feet away can sound like they're whispering in your ear. ▶ 18:29
- The AI agent is now a second customer for room audio alongside the remote human. Capture quality has to be good enough for accurate transcription and in-meeting contextual insight, not just human intelligibility — a shift in how systems should be specified and evaluated. ▶ 21:04
- Expect SMB and mid-market to leapfrog large enterprises on AI-first operating models. Following Christensen's 'plane of non-consumption', disruptive technology is adopted first by those doing nothing today — just as Lync landed with organisations that had no Nortel or Avaya PBX to protect. ▶ 26:10
- An unresolved governance question for service owners: if everything an employee says and does is recorded and modelled, does the resulting digital twin belong to the employer after they leave, can it keep impersonating them with their business network, or does the individual keep it? Pintea points to Inrupt-style personal data pods and federated learning as models where agents query data and reach consensus without disclosing or copying the underlying record. ▶ 28:45
Insights summarised by AI from the episode transcript, reviewed by the Empowering.Cloud team.
Listen: Apple Podcasts · Spotify · Other platforms
Comments ()