Sitting in remote meetings all day long I’m regularly wondering: Why is it so much harder separating voices from different participants if they cross talk than in a meeting room. I knew from university that we can identify where a sound comes from simply by utilizing delay and volume differences between the ears. Then I thought: That might make the difference and it sounds like something you can solve with software. I drafted a prototype with some AI tooling and was surprised how well it worked. I found out that Zoom has some kind of feature like this that they patented. It is in combination with the position of the video AFAIK. That’s where I stopped working on this. I’m not a patent expert, but I still think it is a fun demo and maybe something that is integrated into video call software everywhere.
What is this tool? #
When multiple people talk at the same time, it can be hard to follow any single conversation. This interactive tool uses interaural time differences (ITD) and interaural level differences (ILD) to spatially separate multiple voice recordings, making it easier to distinguish between different speakers even when they overlap.
The tool automatically loads three voice recordings and lets you position each one in stereo space by adjusting:
- ITD (Interaural Time Difference): A tiny delay between left and right ears (in microseconds), simulating sounds coming from different horizontal positions
- ILD (Interaural Level Difference): A volume difference between left and right channels (in decibels)
By giving each speaker a different spatial position, your brain can more easily separate and follow individual conversations, even when they’re playing simultaneously.
Multi-talkers Stereo ITD/ILD Tester
Try adjusting the ITD and ILD values for each track to position them in different parts of the stereo field. The preset buttons provide quick starting points for left, center, and right positioning.
Audio Credit: Includes excerpts from Passive_Houses_Presentation by MMCTV, licensed under CC BY 3.0.