Source
Room
Speakers
Listener
Paths
Top view — drag the speakers and the listener
Settings
Spectrum at the listening position
This is measured before the HRTF, and that is deliberate. Your two ears hear the same sound a fraction of a millisecond apart, so adding them together – which is what a single plot of the output would do – produces a comb filter that is an artefact of the summing rather than something either ear hears. One speaker aimed straight at you, no room, measured 15.5 dB of ripple from 1 to 10 kHz that way and 8.7 dB here.
What this models, and what it does not
Read the assumptions before you trust what you hear
- Modelled: direct sound and first order specular reflections from six surfaces, path length as delay and as 1/r attenuation, one reflection coefficient per bounce, a crude high frequency loss for wall and air, ideal frequency independent directivity (omni, cardioid, dipole — the dipole's rear lobe really is polarity inverted), HRTF panning from the image source position, and a synthetic exponential tail whose decay comes from Sabine on the current room.
- Below the Schroeder frequency a second model takes over, and the two are crossed over rather than mixed. One bounce per surface cannot produce a mode — a mode is what is left after infinitely many bounces — so the bottom end is computed analytically instead: every mode of an ideal rectangular room, f = (c/2)·√((nx/W)² + (ny/H)² + (nz/L)²), each one a bandpass filter whose level is the mode's pressure at the speaker multiplied by its pressure at your ears. Zero means you are sitting in a null, and that is the point.
- The crossover is the part that makes the nulls real, and it is worth understanding before you trust a null you hear. The two models describe the same sound field in two different ways, so running both across the whole band counts the same energy twice: the direct sound and twelve reflections keep playing at 56 Hz and fill in the null the mode bank just created. Measured here with a sine at a modal null, the mode bank produced 18.6 dB of difference between null and maximum while only 2.1 dB survived to the output. In a real room there is no separate direct sound at 56 Hz — the whole field is the modal field, and that is exactly why the null is deep. So the image source branch is high passed at the Schroeder frequency and the mode bank low passed at the same point, and the same measurement then gives 15.8 dB. Second order Butterworth, not Linkwitz-Riley: the two branches carry different signals — one phase comes from resonators, the other from delays — so they sum incoherently, and the criterion that matters is power, |LP|² + |HP|² = 1. Butterworth holds that exactly; LR would dig a 3 dB hole at the crossover. Turning the modes down opens the high pass again, so the anechoic reference you A/B against still has its bottom octaves.
- The mode bank's absolute level is an anchor, not a measurement. A mode's amplitude has no reference of its own the way direct sound has 1/r, so it is scaled to match what the high pass removes: averaged over five listening positions, the modal field is set to carry the same 30–80 Hz energy as the direct sound it replaces. Averaged, because any single position is either a null or a peak. The bank runs up to the Schroeder frequency, above which modes overlap into a statistical field; the absorption slider and the floor material move that boundary, and you can watch it move on the spectrum plot. Q comes from one Sabine number, so every mode decays at the same rate — real ones do not.
- The late tail has a unit now: 100% is the level a diffuse field of this room's absorption would have. In the model's own units that level is exactly 1/rc per channel – as loud as the direct sound at the critical distance, which is what the critical distance means – and the two channels are decorrelated, so two speakers give √2 of it on their own, as the theory wants. Before this the tail was a fixed coefficient: absorption and floor material changed how long it rang but not how loud it was, so a dead room got the same tail as a concrete stairwell. Two consequences worth knowing. The tail now gets quieter when you add absorption and when you pick a more directive speaker, because a directive speaker puts less power into the room to begin with. And in a very live hard-surfaced room the honest level would clip the output, so it is capped – the app says so beside the slider when that happens rather than pretending the room is calibrated.
- The measured ratio beside your seat is read from the running audio, and it is deliberately not the same number as the estimate. It taps every path before the HRTF, splits them into direct and room, and compares the two above the crossover – above it, because below the crossover there is no separate direct sound to compare against, only the modal field. It excludes the mode bank for the same reason. Expect it to read wetter than the estimate: the tail alone is calibrated to the estimate, and the six early reflections arrive on top of it, which is what a real room does near the speakers too. It also lags, since the ratio of two noise signals has to be averaged over about half a second before it is worth reading. If the two numbers disagree by more than a decibel or two after the tail is at 100%, the geometry is telling you something – a seat that close to a wall is not a diffuse field.
- The critical distance beside your listening distance is textbook theory, not a reading from this model. It is rc = √(Q·R/16π), where the room constant R = Sᾱ/(1−ᾱ) comes from the same Sabine numbers as the tail: the distance at which a diffuse field of this room's absorption would be as loud as the direct sound. The familiar 0.057·√(V/RT60) rule of thumb is that formula without the (1−ᾱ) term, so it reads about 20% lower in an absorptive listening room — worth knowing before you conclude one of the two is wrong. Q is the on-axis directivity factor: 1 for omni, exactly 3 for both the cardioid and the dipole, and for the measured speaker it is integrated from the polar data, which is why it arrives as a range instead of one number. How many speakers are playing cancels out — a second one adds direct and reverberant energy in the same proportion — so only the distance matters. Two things the figure does not claim: a small room has no diffuse field at all in its bottom octaves, and the direct-to-reverberant ratio of this page at your seat is a different number, because the late tail here is a slider and the reflections are first order only. Compare it with your own room rather than with what you hear here.
- An ideal shoebox is not your room. No furniture, no bass traps, no door openings, no non-parallel walls, nothing built in. The mode bank is capped at 60 modes and says so above when it runs out. It also arrives in mono to both ears, because below 200 Hz the wavelength is longer than 1.7 m and the difference between your ears is small — but not zero. And the surface mute buttons act on the reflections only: muting a wall does not remove the modes that wall is half of.
- Two symmetric speakers playing mono do not excite the odd width modes at all. Not "weakly" — exactly zero, because the mode's pressure is equal and opposite at the two speaker positions. So the 40.8 Hz width mode of the default room can stay silent no matter where you sit, until the two channels carry different bass. This is real, it is why a symmetric subwoofer pair is recommended against lateral modes, and it is worth knowing before you conclude that moving your seat does nothing. Walk along the room instead and listen to the length modes, or try the even width mode an octave up.
- Also not modelled: diffusion, scattering, furniture, real frequency dependent materials, second and higher order reflections, and therefore flutter echo, which cannot exist in a first order model.
- Measured directivity is an approximation of a measurement. The data is one real speaker exported from VituixCAD: horizontal and vertical polars, every 15°, normalised to the on-axis response — so it carries directivity only, never the speaker's own frequency response. Two planes are not a sphere, so directions in between are guessed as H(θ)·cos²ψ + V(θ)·sin²ψ. Each path then gets that curve fitted with one gain, two high shelves and one peaking filter, which is what lets it track a dragged speaker without clicking; the fit sits within about 3 dB of the measured curve on average and up to 8 dB at worst, and below 250 Hz it holds the 250 Hz value rather than returning to omni. The measurement is magnitude only, so unlike the ideal dipole here, no rear lobe is polarity inverted. The export does not state which way a positive vertical angle points, and the vertical polar is asymmetric — so that switch trades the floor reflection for the ceiling one, and up is the owner's answer rather than a documented convention.
- The HRTF is the browser's own generic set. It is not yours, there is no head tracking, and elevation is its weakest axis — which is unfortunate, because floor and ceiling are exactly the reflections people argue about. Judge side wall reflections first.
- Muting a reflection removes energy, so some of what you hear is simply a level change. That is honest here — with SBIR the level change is the phenomenon — but keep it in mind when comparing two positions.