Saturday, October 10, 2026

Exploring Linear Phase Speakers, Part II

After figuring out what a “linear phase speaker” means and checking some professional studio monitors in Part I, we can apply these ideas to the tuning process of my LXdesktop.

Original LXmini—Not Linear Phase?

Before throwing in the power of linear-phase DSP processing and paying the latency costs, let’s first check what the original LXmini implementation looks like in terms of phase rotation and group delay. For that, we first build a simple idealized approximation of it and compare it with the actual measurement.

There are three major factors that define the phase rotation in the original LXmini:

  • the sealed woofer pipe, which acts as a 2nd-order high-pass filter with a corner frequency of about 45 Hz;
  • the LR2 crossover at 700 Hz;
  • the high-frequency roll-off of the full-range driver—the actual corner frequency and the slope largely depend on the measuring conditions: the measuring microphone frequency range, and the sampling rate.

All these components are minimum-phase filters. Modeling the first two is rather trivial. Modeling the low-pass part cost me some time. I figured out that besides the phase rotation from the low-pass filter itself, the other contributing factor is the brick-wall filter employed by the DAC/ADC that are used for measurement.

I discovered this because a minimum-phase low-pass filter of the same magnitude shape did not exhibit the same phase behavior, so there had to be an all-pass component. Measuring the electrical chain alone, at higher sampling rates, confirmed that.

Thus, the LXmini model I’ve built here is not purely theoretical: I had to fit the properties of the low-pass filter to my measurement. However, the nice part is that even the impulse response of this model looks very close to the actual IR at the macro level:

Note that since in the LXmini the polarity of the full-range driver is inverted, the pulse actually starts on the negative side. However, since acoustical software typically looks at the largest positive peak, to avoid confusing it I have changed the polarity of the entire pulse—this does not affect the magnitude response, and allows Acourate to find the peak and reference the phase to it correctly. This is the graph of phase shifts of the actual speaker and its model:

So the total phase rotation (if we unwrap it) is approximately 520° of which 180° comes from the high-pass, another 180° from the crossover, and the rest is the low-pass filter plus the measurement artifact from the anti-aliasing filter of the audio interface.

The resulting group delay predicted by the model:

(Note that Acourate places the start of the impulse at sample position 6,000; thus, the baseline group delay is approximately 125 ms.)

We can see that the group delay starts rising below 2 kHz and reaches approximately 1.6 ms at 100 Hz. This is actually not that bad (and definitely better than the Genelec 8331A in the “low latency” mode!) because presumably the group delay stays within inaudible boundaries, according to the research considered in Part I of the post.

We can conclude that the original LXmini is definitely well designed from the psychoacoustic perspective. However, it is not a “linear phase speaker,” thus some improvement may still be possible.

Re-Tuning LXdesktop

LXdesktop differs from LXmini in two aspects: the pipe of the woofer is shorter, and I use a sealed subwoofer. Thus, it is a three-way system with two crossovers to define. Let’s start with the lower one.

Defining the Woofer/Subwoofer Crossover

Choosing the crossover frequency between the subwoofer and the woofer has turned out to be a challenging task. We don’t want to overload the woofer by forcing it to go too low: due to the smaller enclosure volume it can’t go down to 45 Hz like the original LXmini. Since I’m tuning this speaker setup for my room, the choice of the crossover is largely dictated by the interaction of the woofer and the subwoofer with the room modes. Unlike the Dutch & Dutch speaker, my subwoofer is not co-located with the woofer, thus they couple to the room modes differently—in fact, significantly differently. This dictated the choice of the crossover: find the point on the frequency scale where their phase behavior is more or less close, and narrow down the crossover overlap region to that frequency as much as possible.

However, since speaker drivers are naturally minimum-phase systems, any deviations in the magnitude response affect the phase. Thus, the drivers must be linearized before doing the comparisons. Another important thing is that measurements of both drivers must use the same time reference: preferably a high-frequency driver for producing a narrow synchronizing peak.

We can see that the phases of both drivers (red: subwoofer, green: woofer) have similar slopes in the region around 60 Hz, so I ended up choosing 60 Hz as the crossover frequency. Then I started looking for a brickwall-like crossover to minimize the overlap region. With linear-phase crossovers even steep slopes by definition don’t pose any phase alignment issues; however, they have a potential issue of pre-ringing. The pre-ringing definitely must lie beneath psychoacoustic pre-masking thresholds. I considered several types of steep crossovers: LR8, 2nd-order Neville-Thiele, and 2nd-order polynomial UB (Brüggemann). All of them offered negligible pre-ringing, and UB are the steepest ones, so I chose that type. Here is how its pre-ringing looks:

The pre-ringing above -60 dB from the peak lasts for approximately 5 ms.

Note that under ideal conditions: when the speaker drivers actually behave like the crossovers, and with the listening point at the tuning position, the pre- and post-ringing parts of the driver IRs cancel each other completely. However, as we will see this is hardly achievable in reality, thus the pre-ringing of the crossover must be kept to the minimum.

Defining the Woofer/Full-Range Crossover

For this problem, we have two constraints. First, since the woofer driver is oriented upward and is located approximately at ear level, we should limit its range to the region of its omnidirectional radiation pattern. According to the technical documentation on the SEAS L16RN-SL driver, it is omnidirectional up to 500 Hz. However, this is the area where the second constraint is not met—the full-range driver, used without a baffle, simply can’t go that low without risking over-excursion. Thus we need to pick a trade-off between operating the woofer driver mostly in its omnidirectional range while avoiding overloading the full-range driver. LXmini uses 700 Hz as the crossover point, and I decided to keep it.

Side note: previously I was trying to find a full-range driver that would exhibit the lowest distortion in baffleless operation. However, maybe a better criterion would be to pick the one which allows going as low as possible without going over its excursion limits.

Although I kept the crossover frequency, I also decided to narrow the crossover overlap region: the LR2 used by LXmini is too broad. I checked steeper crossovers first, and unfortunately they all had significant pre-ringing, which is a bigger concern here than in the bass region because human hearing is much more sensitive in the midrange. So I reduced the crossover order and ended up using 1st-order Neville-Thiele:

(Note that the time scale on this graph is different from the previous one.)

The pre-ringing above -60 dB from the peak lasts for approximately 1.2 ms. While temporal backward masking (pre-masking) is generally weak, it is highly effective within a brief 1–2 ms window immediately before a loud transient, so this pre-ringing should stay hidden.

The Ideal Speaker Model

With both crossovers defined, it’s time to take our mixed-phase ideal speaker model from Part I and split it into the bands for each driver. Due to the minimum-phase nature of our high-pass filter, it affects the phase rotation of every band. That means it must be convolved with each of the speaker bands. After that, we need to double-check that all the crossovers actually sum back into our ideal speaker magnitude response, and that the shape of the impulse response stays the same. This is indeed the case:

Note: if we look at the IRs of individual crossover bands, we see that their peaks do not align at the same position. This is totally correct because our model has a minimum-phase component:

Having this reference alignment will be very handy later for time-aligning the drivers. But before that, the drivers must be linearized in order to coerce them into the shape of the crossover bands we have defined. And before doing even that, let’s do a quick experiment.

The Ideal Speaker, Acourate’s Take

What happens if we apply Acourate’s room/speaker correction to the IR of our ideal speaker? Interestingly, it “disagrees” that it is ideal and corrects it! How exactly? Acourate actually brings it to a minimum-phase behavior:

(The red trace is the phase of the ideal speaker after Acourate’s correction; the magnitude response stays the same.)

Why? The philosophy of Acourate is that speakers are minimum-phase devices. Plus, as we discussed before, minimum-phase filters exhibit no pre-ringing at all, by definition. One of the goals of the acouStep algorithm is to avoid pre-ringing.

The process of the speaker correction builds a filter which removes the excess phase. Typically, this excess phase comes from analog crossovers (in passive speakers) and speaker driver imperfections (we will see that later, once we start to combine linearized drivers). However, the linear-phase behavior at the end of the speaker frequency range in our ideal model is “excess phase” too! Acourate does not distinguish between “good” excess phase (like in our model’s case: it has excess phase because the model is linear-phase, but this pre-ringing is imperceptible) and “bad” excess phase (from minimum-phase crossovers, potentially audible); it just removes all of it for the direct speaker sound.

Should we worry about this? I don’t think so. First, we actually still use the ideal speaker model for defining crossovers. Second, when the speaker magnitude response is mostly flat, there is no big difference in the phase behavior between the linear-phase and minimum-phase model. Third, as follows from my brief analysis of the audio interface behavior, the phase behavior at the end of the frequency range may be severely affected by its brick-wall filters, thus we risk over-correcting.

Driver Linearization with Sinc-Pulse

Now, back to the main path. With the linear-phase crossover components developed, we can use them directly with the sinc-pulse linearization function of Acourate. Applying it to each driver of this system has its own intricacies, though.

The most straightforward is the woofer because it uses a classical sealed enclosure arrangement. In order to simplify microphone placement, I used a version of LXdesktop stripped of the full-range driver:

Although the full-range driver as a physical object certainly creates some reflections, due to its size they mostly affect mid- and high-frequency ranges which we taper with the crossover.

The full-range driver works without a baffle (as a dipole) and thus has backward radiation; measuring it in the near-field does not capture the dipole cancellation and thus does not reveal its actual far-field behavior. However, we don’t want to measure it from too far (the main listening position, MLP) because that will incorporate unrelated reflections. I experimented with the distance to ensure that the microphone is not very far from the driver, while the impulse response still looks similar to what I’m seeing from the MLP. Because of these constraints, the IR of the full-range driver looks more jagged; however, that is fine because Acourate applies some smoothing in the sinc-pulse linearization procedure anyway.

I have to say that the sinc-pulse driver linearization procedure is much more straightforward than my previous experience with Acourate V2 where I was applying the speaker correction to each driver and also manually removing phase deviations by means of all-pass filters. One important thing to remember is that the process of probing the speaker driver with a sinc pulse is much more sensitive to background noise. In order to get a measurement with a high SNR, the procedure needs to be repeated about 200 times (though this is not a major issue, since it runs fully automatically).

And for the subwoofer, I actually didn’t use the sinc-pulse linearization at all because the near-field sinc-pulse measurement has shown that it does not have many irregularities to compensate. So I decided simply to apply my raw crossover to it, and then iron out the magnitude response shape during the room correction procedure. After all, the low-frequency output in a room is dominated by room modes, which induce far greater magnitude swings than any native irregularities of the driver itself.

Ideal vs. Real Speaker

As the result of linearization, we have the impulse response of each driver brought as close as physically possible to its crossover band. What we can do at this stage is to see what happens if we align them in the same way as the IRs of our model’s crossovers and then sum them. This is what happens:

You can see that although it looks close to the shape of the ideal IR, there are some notable differences. First is that it is contaminated with reflections—this is unavoidable because the linearization was done in a home room, and Acourate does not try to fill up every possible dent in the magnitude response in order to avoid over-correcting.

The second difference is more interesting: we can see that the real speaker’s impulse has some pre-ringing ripple. This is not the result of bad time alignment—the impulses of real drivers are aligned as close as possible to their ideal speaker counterparts. The ripple is caused by non-ideal summation of pre-ringing from the drivers’ crossovers. Drivers simply can’t behave as ideal crossover band filters. Due to physical constraints, they have what is called “passband ripple”—unavoidable deviations from the target response.

Why this pre-ringing is not a problem: first, thanks to the initial choice of the crossovers, it falls under pre-masking thresholds; second, after we assemble the real speaker’s response, we can optimize it with acouStep. In fact, we can preview the result by applying Acourate’s macros to our intermediate IR as if we had measured it. Below is the optimized IR produced by running the “test convolution” with Acourate’s correction filter:

We can see that acouStep is capable of shifting the energy from the pre-ringing to post-ringing thus totally sweeping it under the rug of the post-masking of human hearing which has an even longer span than pre-masking. Finally, below are group delays of the ideal speaker, the uncorrected “real speaker”, and its corrected version:

As a reminder: the red trace is the “real” speaker, the green trace is the “ideal” speaker, and the brown is the corrected “real” speaker. Mostly we see that the variation of the group delay is reduced across the frequency range, and the group delay below 20 Hz is reduced (although this is not that important). Let’s also look at the phase:

The phase graph above shows once again that Acourate corrects the speaker to minimum phase—the phase is not flat in the high frequency region where our ideal speaker (it was used as the target here) has the roll-off.

Time Alignment

With driver linearization complete, and having confirmed that our assembled speaker should mostly conform to the ideal speaker behavior, the next step is to “assemble” the output from the whole speaker by aligning individual driver outputs in time and adjusting their relative gains. I used to employ the sine-wave convolution approach for that; however, after considering the coherence (or, mostly, the lack of it) between the woofer and the subwoofer, I started to doubt that approach: aligning them at a single frequency point may be an over-optimization which makes summation in the adjacent regions not so ideal.

Learning from the experience of assembling the speaker response from the IRs of individual drivers, I realized that the same approach can be used for measurement of the drivers from the MLP. One difference is that we absolutely need to establish the timing reference. Typically, the driver at the highest frequency band is used for that because its impulse is the narrowest and “tallest” one. In order to split out the IRs of individual drivers within one single measurement, we can use the “impulse shift” method which is described both in M. Barnett’s book on Acourate and the free manual by Dr. Keith Wong available at the Acourate user forum.

The idea is that we delay the time anchor driver (in the case of the LXdesktop this is the full-range driver) by some known amount, for example 1000 samples or more. This makes the IR of the driver being aligned come before it. For example, for the woofer driver, the measured IR with the “impulse shift” looks like this:

The green trace is the IR of the crossover at its aligned position, shifted left by 1000 samples. This is our reference that we can align our real IR against. The strong peak to the right is the impulse of the full-range driver.

Because the woofer’s IR precedes the IR of the full-range driver, it is not obscured by the room reflections of the full-range driver’s pulse. Note that the IR of the woofer is relatively compact, thus the separation by 1000 samples is enough to see its main pulse part. The IR of the subwoofer is more smeared in time, and I had to use a 3000-sample offset instead.

You can see that the subwoofer’s impulse is so long that the delayed impulse of the full-range driver “rides” on it. Also note that there is a significant distance between where the subwoofer’s impulse needs to be and where it is before the alignment.

How do we perform the alignment against the reference? For the woofer’s IR this can be easily done “by eye” by looking at the IR peak (Acourate can also show its sample index, as the local maximum in the selected area). The subwoofer’s IR (especially the one from the actual physical subwoofer) may look more ambiguous. I ended up asking Claude to write me a MATLAB script for aligning IRs via cross-correlation. It also suggested other methods for subwoofer alignment such as GCC-PHAT, however their application actually resulted in poorer alignment than straightforward cross-correlation (which is not too surprising: the PHAT weighting whitens the spectrum, so for a band-limited signal like the subwoofer’s it gives equal weight to the out-of-band bins which contain only noise). But I assume this is where YMMV based on your actual room setup.

After time-aligning the drivers, I also performed level alignment by looking at magnitude responses with a frequency-dependent window (FDW) applied and adjusting the gain of filters.

The final check of the alignment is done by measuring the complete speaker. We can see that the IR of the actual speaker is indeed close to the ideal speaker, and exhibits the same pre-ringing issue that we already saw with our synthetic “real speaker”:

But we know that this can be corrected. In fact, because each driver would require a different passband ripple correction, it’s much easier to accept it first, and then remove it at the final stage.

Final Correction

As the last step, I performed the overall room/speaker correction in Acourate. I used the new “acouStep” option when making it, and the resulting step response looks rather nice, featuring a sharp onset and almost no pre-ringing:

(Note that here red and green are the left and the right speaker, respectively.) Did we achieve the desired linear-phase behavior? Yes—if we look at the phase and the group delay of the FDW-windowed response, we can see that they are flat, except for the low end, below 100 Hz where, as we know, group delay non-uniformity is imperceptible:

Recall that Acourate actually corrects toward minimum-phase behavior. This is why there is a phase roll-off at the high end. However, this is beyond my hearing range anyway. Also, as I mentioned in the beginning, if we try to correct that “imperfection,” we risk over-correcting.

As mentioned before, Acourate places the start of the impulse at sample position 6,000; thus, the baseline group delay is approximately 125 ms. The significant phase deviation near 100 Hz (and its corresponding group delay dip) is caused by a boundary reflection, as is the phase ripple near 300 Hz in the right channel. Trying to correct these boundary cancellations in DSP makes little sense.

Approach Recap

Just to outline our approach once again, step by step:

  1. We start by defining the ideal speaker behavior as a mixed-phase band-pass filter and ensure that the pre-ringing is within masking thresholds, and the group delay only deviates from flat below 100 Hz.
  2. Then we choose crossover points based on the capabilities of drivers and create linear-phase crossover bands (with the minimum-phase high-pass filter of the low-end cutoff applied to all the bands).
  3. We use Acourate’s sinc-pulse linearization to actually coerce the drivers into the crossover band-pass behavior.
  4. Then we time-align and level-align the drivers to achieve a close-to-ideal transition between them.
  5. The resulting impulse response is close to ideal, but exposes the imperfections of physical drivers, such as the passband ripple. These imperfections are swept under the rug of psychoacoustic masking after applying Acourate’s speaker/room correction to the whole speaker.

And as a result we end up with a “linear phase” speaker (by industry standards) which is actually mixed-phase, but has linear-phase behavior where it matters—in the passband.

Some Listening

Having the sound system set up, I did some listening to albums that feature enveloping, tonally rich compositions with a non-trivial spatial layout. I was listening using the psychoacoustic correction for the center channel and the diffuse field, which essentially adds virtual center and surround speakers (although this effect is limited to the MLP only).

These are the tracks that I used:

  • Too Much Rope / Amused to Death 1992 by Roger Waters; Stereo and Multi-Channel versions. This album was recorded using QSound technology, which could probably have developed into something similar to today’s Dolby Atmos, but unfortunately ended up being used on just a handful of albums. This track in particular features an audio rendering of a horse carriage moving from the far front left to the far back right. It was featured on Archimago’s blog for demonstrating the effect of crosstalk cancellation, but it also works correctly on my setup. An interesting observation is that the HALO Upmixer extracts some correlated components and sends them to the virtual center speaker, and this sounds more convincing than the multi-channel version, in which the center channel is mostly silent during this fragment.
  • Don’t Give Up (feat. Kate Bush) / So 1986 by Peter Gabriel. A high-quality, only slightly compressed recording featuring male and female vocals with synth ambience, bass guitar, and some other instruments. I think it demonstrates the importance of a coherent wavefront, which allows for easy pinpointing of instrument locations.
  • Heptapod B / Arrival OST 2016 by Jóhann Jóhannsson. This composition is even more interesting than the previous one—there is more ambience and more strange sounds coming from various locations and moving around the listener.
  • Track 18 from Ghosts I-IV (2008) by Nine Inch Nails. This album set is a dump of various previously unreleased compositions created by Trent Reznor. His attention to detail in music always amuses me, and you can hear it even in these simple but well-thought-out tracks. The samples used on this track are clean, and spatially arranged in a way that extends beyond the speakers. It’s a sort of “Chesky Records going industrial” thing. If this track feels too relaxed, try track 24 instead.
  • Rocksavage / The Nation’s Most Central Location 2023 by Warrington-Runcorn New Town Development Plan. Probably the longest band name I’ve ever seen. In the music, I can feel the influence of Kraftwerk and Vangelis. This composition features a very articulate, bouncing bass which creates a rhythmic pattern reminding me of the one used in “The Man-Machine,” but the samples used sound more “modern.” The ambience of the virtual space also feels more real than in Kraftwerk`s compositions.

Conclusions

I’m almost finished creating the new speaker setup; the only missing piece is the diffusers (arriving soon!), which should help remove some asymmetric reflections from the back wall. I think that this setup can match the symmetry and linearity of earspeakers, so that comparing them one-to-one becomes more straightforward.

Thursday, October 1, 2026

Exploring Linear Phase Speakers, Part I

I’m getting closer to experimenting with comparisons between earspeakers and real speakers. Since earspeakers (any headphones, really) are usually rather symmetric, I decided to physically change my desktop speaker setup to make it more symmetric, too. In the process, I also decided to retune my LXdesktop using the new tools added in Acourate V3 and V4: speaker linearization using a sinc pulse and the new acouStep tuning algorithm, which virtually eliminates pre-ringing in the tuned speaker response.

From my experience with the last tuning of LXdesktop, I realized that it’s much easier to tune speaker drivers by modeling them as components of a linear-phase crossover rather than a traditional minimum-phase crossover. Mathematically, modeling the drivers as minimum-phase crossover bands and then linearizing the phase afterward (as described in the Grimm Audio paper on the LS1 speaker) is equivalent to using linear-phase crossovers in the first place. In practice, however, speaker drivers that behave like linear-phase filters are much easier to time-align. Plus, the sinc-pulse linearization tool introduced in Acourate V3 makes it very straightforward to bring drivers to linear-phase crossover behavior.

The Idealized Speaker Model

Let’s state the goal first. Ideally, we want the entire speaker to behave as a linear-phase band-pass filter. Why “linear-phase”? Because linear phase guarantees constant group delay. That means that if the speaker plays a short wideband impulse, such as a click, all of its component frequencies reach the ear simultaneously. In other words, a linear-phase speaker preserves the waveform of transients in its direct sound (within its passband, of course), ensuring natural reproduction.

And why “band-pass”? Because the speaker is physically limited in the range of frequencies that it can reproduce. Even a big, high-quality speaker can’t cover the entire range from 0 Hz to infinity. First, this is not needed, because the range of human hearing is limited. Second, there are physical limits on how low and how high acoustic transducers can go. Going too low or too high requires too much energy and can easily overload the drivers. Thus, any speaker covers a limited frequency range, which technically makes it a band-pass filter.

Note that within the flat part of the passband, far away from the band edges, a minimum-phase filter also produces a similarly linear result. The defining property of a minimum-phase system is that its phase response is uniquely determined by its amplitude response (the two are related via the Hilbert transform), and so is its group delay. Where the frequency response is flat, a minimum-phase filter produces a flat (zero) phase response and a constant group delay. However, it introduces phase shifts as soon as the frequency response bends, and a band-pass filter, by definition, bends at both ends.

As a theoretical foundation for our experiment, we can use the open-access AES paper “Modeling and Delay-Equalizing Loudspeaker Responses” by A. Mäkivirta, J. Liski, and V. Välimäki, written as part of a collaboration between Genelec and the Acoustics Lab of Aalto University in Finland. The paper takes exactly this approach: it starts by modeling a speaker as an ideal band-pass filter.

This is fairly straightforward for a single-driver loudspeaker. In practice, however, most speakers use multiple drivers. Rare exceptions are speakers with a single full-range driver, and full-range electrostatic and magnetic planar speakers. All of my speakers are multi-driver, which means that their ideal counterpart, the band-pass filter, is in turn split into overlapping band-pass components. The splitting is done by crossovers. In my LXdesktop speakers, the crossovers are implemented in DSP, so they can have any desired behavior. Mäkivirta’s paper discusses the classic case of minimum-phase crossovers and describes the challenges of time-aligning them. One important property of minimum-phase crossovers is that, in a speaker with more than one crossover, they must be chained using a special topology:

This is because the phase shift created by the woofer-mid crossover must be applied to both the midrange and the tweeter drivers; otherwise, the parts of this composite crossover will not sum to a flat frequency response. Note that since the cutoff of the woofer-mid high-pass filter lies far below the tweeter’s band, this filter has no effect on the amplitude of the tweeter signal; it only affects its phase.

Linear-phase crossover filters do not have this issue because their effect on the phase is strictly linear: they only produce a pure time delay. As long as all the filters share the same delay (that is, their impulse responses are centered at the same sample position), they can be arranged in a simpler way:

The basic idea behind driver tuning stays the same, though. We do not consider the crossover filters and the speaker drivers separately. Instead, we try to coerce each speaker driver into becoming an ideal crossover component by developing a correction filter for it, based on the transfer function of the corresponding part of the ideal crossover.

Because of the extra complexity of minimum-phase crossovers, I use linear-phase crossovers in my project. In this case, each driver is coerced into its crossover component behavior using Acourate’s sinc-pulse linearization. This process takes two inputs: the actual response of the driver measured close to it, including the effects of diffraction at the baffle edges (note that a true near-field measurement, with the microphone right at the cone, would exclude them), and the desired crossover impulse response. The output is a filter that brings the driver as close as possible to the desired behavior.

The Catch at the Low End

There is one important caveat that Mäkivirta’s paper explains. Since neither the whole speaker nor its lowest-frequency component (in my case, the subwoofer) goes down to 0 Hz, there is a roll-off at low frequencies. For a subwoofer, or any large driver, the roll-off frequency can be quite low; let’s use 15 Hz as an example. The caveat is that if we represent this roll-off as a linear-phase filter, the filter has massive pre-ringing. If we represent it as a minimum-phase filter instead, group delay is inevitable because, as I noted earlier, in a minimum-phase system amplitude changes are unambiguously coupled with phase changes.

Let’s look at both alternatives. Below is the frequency response of a 4th-order Butterworth high-pass filter at 15 Hz:

And this is the RMS view of the impulse response (IR) of its linear-phase version:

We can see pre-ringing that lasts about 300 ms. Although its energy is concentrated around the cutoff frequency, where hearing is not very sensitive, with a capable subwoofer it is likely audible on bass-heavy transients, sounding like “breathing.” So, it is much more natural to use a minimum-phase filter for the high-pass part (after all, speaker drivers naturally tend to be minimum-phase devices). By definition, minimum-phase filters have no pre-ringing; however, the minimum-phase 4th-order Butterworth high-pass filter has a significant group delay, which reaches about 22 ms at 20 Hz:

Psychoacoustic research has shown that the human ear is relatively insensitive to the late arrival of low frequencies relative to the initial wavefront, especially with music and in typical rooms, where reverberation time at the low end can be quite long. For references, see another paper from the same Genelec–Aalto University collaboration, “Audibility of loudspeaker group-delay characteristics”, and Section 9.6.5 of the 4th edition of F. Toole’s “Sound Reproduction” book. The listening tests described in the paper found that group delay can exceed 10 ms below 200 Hz without the difference being audible, whereas in the 300 Hz–1 kHz range, differences became audible above just 1–2 ms.

Going back to the linear-phase high-pass filter: its 300 ms of pre-ringing far exceeds the span of pre-masking (backward masking) in human hearing, which, according to the studies reviewed in the same paper, is only about 5–20 ms. That explains why this pre-ringing is audible.

Practical Tradeoff: A Mixed-Phase Filter

Based on this, it is fine to keep the minimum-phase 4th-order Butterworth roll-off at low frequencies. This means that even our idealized, theoretical model of the speaker is not completely linear-phase. Instead, it is mixed-phase: it has minimum-phase behavior at the low end and linear-phase behavior above it. Mäkivirta’s paper comes to a similar conclusion: it is simply not practical to coerce the speaker into linear-phase behavior at low frequencies.

Here are the characteristics of our idealized speaker model. The frequency response is tapered at both the low and high ends:

The group delay is flat everywhere except at the low end, where it is determined by the Butterworth high-pass filter:

(Note that since this is a linear-phase filter for a 48 kHz sampling rate, centered at sample position 32,768, the group delay baseline is about 682 ms.) And this is the step response:

From this distance, it might look like a minimum-phase response. However, if we zoom into the onset, we can see some minimal pre-ringing from the linear-phase low-pass filter:

Its duration is well under a millisecond (a few periods of the cutoff frequency), and it is thus covered by the pre-masking mechanisms of human hearing.

Examining Linear-Phase Studio Monitors

As described in the Grimm Audio paper on the design of the LS1 speaker, its designers chose a minimum-phase LR4 crossover and then corrected its phase by applying an anti-causal all-pass filter (the sum of an LR4 crossover is a 2nd-order all-pass with a Q of 0.7, and its inverse is a non-causal filter, which is implemented as an FIR with added delay). Let’s see what other makers of linear-phase studio monitors do. I measured the behavior of two professional studio monitors: the Genelec 8331A and the Dutch & Dutch 8c.

The anti-causal all-pass filter introduces extra latency: from 4 ms to 30 ms, depending on the frequency of the crossover being compensated. Since this may be critical in live monitoring scenarios, manufacturers usually keep the natural minimum-phase behavior available, calling it the “low latency” mode.

Also note that this type of correction is only practical with DSP because it involves non-causal (look-ahead) filters, which can only be realized by adding delay. As a consequence, although both of these studio monitors accept a traditional analog input, it is there only for legacy compatibility. Their native input is digital, so the analog signal first goes through analog-to-digital conversion. Because of that, both monitors have a baseline latency of 3.2 ms, even in the “low latency” mode.

Genelec 8331A

This is a small desktop-size monitor with a bandwidth of 45 Hz to 37 kHz. The low end is extended by means of a port. Two notable design features of this speaker are the coaxial mid-tweeter driver and the two “racetrack”-shaped woofers at the top and bottom of the front panel, hidden under a cover.

This is the smallest speaker in Genelec’s “The Ones” series. The next model up (8341A) uses the same mid-tweeter driver, while the two larger models (8351B and 8361A) use a slightly bigger one. As the speakers get bigger, so do their woofers. This makes the largest model in the series look a bit unusual, earning it the nickname “Cyclops.” Even this big guy only extends down to 30 Hz, so covering the full audio range requires a subwoofer or the specially designed W371A, a stand with an integrated woofer. See the full specs here.

One caveat with these speakers is that to get full access to their features, including the ability to switch between the “low latency” and “linear phase” modes, the user needs to buy a proprietary “Loudspeaker Manager User Kit” (GLM), which alone costs as much as a mid-priced speaker:

To characterize the linear-phase correction filter, I measured the speaker in the same physical setup in the “low latency” and “linear phase” modes. Then I could simply divide the obtained transfer functions (complex division in the frequency domain, that is, deconvolution of the impulse responses) to derive the correction filter. Below are the phase and group delay differences between the two modes of the 8331A (the frequency response stays the same):

(Note that the group delay is contaminated by room interaction.) We can see that, similar to the Grimm Audio LS1 approach, the “linear phase” mode adds an anti-causal all-pass filter that corrects the phase deviation caused by the minimum-phase crossover. Genelec does not reveal which type of crossover the speaker uses. The manual only specifies the crossover frequencies: 500 Hz for woofer/midrange and 3 kHz for midrange/tweeter (both of the latter are part of the same coaxial driver). The all-pass filter applies up to 540° of phase rotation, so the crossover might be a combination of an LR4 (360° of phase rotation) and an LR2 (180°). I modeled an LR4 + LR2 crossover pair at the 8331A’s crossover frequencies. Here is the phase of the compensating all-pass filter for this model, overlaid with the actual filter of the 8331A:

The slope of the phase is very similar, but the details differ. This might be because Genelec uses different crossover types, and also because they employ a cleverly engineered approach with windowed IIR filters instead of a straightforward inverse FIR filter. Their design aims to achieve the lowest possible latency, even in the “linear phase” mode. According to the manual, this mode adds only 6.9 – 3.2 = 3.7 ms of extra latency. The RMS view of the all-pass filter’s impulse response shows that the pre-ringing is at -60 dB at approximately -3.7 ms relative to the peak:

This means that the extra delay is due to the filter itself. In particular, there is no evidence of delay caused by FFT buffering, which would be inevitable if the designers used a classic FIR filter applied via frequency-domain processing.

The choice of crossover frequencies also helped keep the filter short. However, I think that the 500 Hz woofer crossover frequency was actually dictated in large part by the physical arrangement. Since the woofers sit behind the front baffle, which acts as a waveguide for the coaxial driver (see this paper on the design of “The Ones”), they radiate mostly upward and downward, not toward the listener. Thus, the designers had to keep the woofers in the frequency region where their radiation stays omnidirectional.

We also need to confirm that the low end of the speaker stays minimum-phase in the “linear phase” mode. Indeed, it does. Since my group delay graph is contaminated by room interaction, here is a clean group delay graph of the 8331A in this mode from Genelec’s site:

The group delay starts to increase below 200 Hz, reaching 30 ms at 50 Hz, which is totally acceptable according to the research cited above.

Dutch & Dutch 8c

This is a big, full-bandwidth speaker (20 Hz to 20 kHz). Full-bandwidth output is achieved by integrating two rear-firing 8-inch subwoofer drivers in a sealed enclosure. They are designed to couple with the front wall (the wall behind the speakers), so the speaker should be installed close to it, about 20–80 cm away. This is an unusual recommendation because, normally, placing a speaker near a wall causes speaker-boundary interference response (SBIR) problems at low-mid frequencies. However, the clever design of this speaker mitigates this issue in two ways. First, the midrange woofer sits in its own sub-enclosure with acoustically resistive slots on both sides of the cabinet. The slots let out the delayed and attenuated rear radiation of the driver, which cancels the front radiation toward the back. This makes the woofer’s radiation pattern a cardioid over its entire working range (100 Hz and up). See the measurements in Erin’s Audio Corner for full details. Second, since the subwoofers themselves are so close to the wall, the first SBIR notch falls well above their 100 Hz cutoff, in the region covered by the cardioid midrange. Note that SBIR caused by large surfaces in front of the speaker (the floor!) is still present.

The acoustic design of this speaker is great, and so are its electronics. It features an Ethernet connection, which allows accessing its DSP controls from a smartphone app and manipulating the correction filters directly from Room EQ Wizard. Switching between the “low latency” and “linear phase” modes is done from the app. Let’s examine how the speaker’s response differs between the modes. As with the Genelec, the frequency response stays the same; only the phase, and consequently the group delay, changes:

(The ripple on the group delay graph is exactly the floor reflection I mentioned earlier.) Compared to Genelec’s all-pass filter, there is more phase rotation here: roughly two full turns. Unlike Genelec, Dutch & Dutch does not conceal the details of its crossover. The official documentation specifies that there are two LR4 crossovers: one between the tweeter and the woofer at 1250 Hz, and another between the woofer and the subwoofers at 100 Hz.

Just to confirm, I modeled the same configuration of minimum-phase crossovers and derived the compensating all-pass filter. It matches the measured all-pass filter of the 8c:

Due to the much lower crossover frequency, the all-pass filter’s impulse response is necessarily longer. Dutch & Dutch specifies that it adds 30 ms of delay (33.2 ms total versus 3.2 ms in the “low latency” mode). Interestingly, the RMS view of the filter’s impulse response suggests that it could probably be windowed to 8.5 ms:

(This graph overlays the measured response of the speaker’s all-pass filter, in green, with the theoretical one, in purple, which looks much cleaner.)

The designers may have chosen 30 ms either to achieve better frequency resolution of the filter, or simply because this is the granularity of their DSP processing (Martijn Mensink, the designer of the speaker, discusses the system’s latency here).

Preliminary Conclusions

Representing a loudspeaker as a mixed-phase band-pass filter makes it possible to achieve flat group delay over almost its entire working range, except for the bass region, where the ear is insensitive to group delay deviations.

Modern professional DSP-based studio monitors offer a tradeoff between the low latency of minimum-phase behavior and the better sonic precision of linear-phase behavior. Since linear-phase behavior inevitably introduces extra latency, it cannot be the default and must be enabled manually in the control software.

For LXdesktop, though, I don’t need to support a low-latency mode, so I can start with linear-phase crossovers, which make driver time alignment easier, especially without access to an anechoic chamber.

Sunday, July 5, 2026

Earspeakers Calibration and Acoustic Measurements

The next step in my ongoing study of the phantom center and diffuse sound colorations on stereo speakers (see Part I, Part II) will be about the “ideal” conditions for reproducing the phantom center that can be created in a normal, non-anechoic room by means of earspeakers. These listening devices are also known in research circles as “free-field” or “extraaural” headphones. Their main difference from regular headphones is that they create much less occlusion of the ears. This design factor allows them to be used in situations where we need to compare sound from headphones with external sounds directly. With earspeakers this is possible because the listener does not have to take them off in order to be able to hear the external sound practically unaffected by the headphones (although, as we will see, this is not entirely true).

And this property of earspeakers is really unique. Even open-back circumaural headphones of lightweight construction, for example electrostatic headphones, do attenuate high frequencies severely (for example, see the paper “Comparing the effect of different open headphone models on the perception of a real sound source…”) and also affect sound localization of external sources (see the paper “The Influence of Headphones on the Localization of External Loudspeaker Sources”).

I read about a study in which the researchers actually tried to compensate for the headphone-induced attenuation by applying a reverse filter to the speaker signals. I tried that myself with the Sennheiser HD800. However, I was not satisfied with the result for two reasons:

  1. Applying significant high-frequency amplification (an inverse filter for the occlusion from the headphones) to the sound from speakers evokes stronger room reflections and the overall result does not sound fully natural.

  2. Since having headphones on the head also affects other components of the HRTF such as ILD and ITD, external sounds filtered by open-back headphones still are not perceived the same as without them, even with the spectrum being compensated. In particular, this affects the phantom center because it relies on the symmetry of the stereo pair.

So, having your ears unoccluded while experimenting with the phantom center is actually a good idea. The problem is that all models of acoustic headphones that leave your ears unoccluded do look really strange, if not completely weird (see this open-access paper for photos). There are only a couple of commercial models, namely: AKG K1000 (discontinued), Sony PFR-V1 (discontinued), MySphere 3.2 (still in production; however, is rather expensive). These are very niche products, and the discontinued models are hard to find for a reasonable price.

However, AR/VR researchers do love earspeakers because of their property to allow hearing both natural and synthesized acoustic reality at the same time, creating “augmented reality.” And researchers come up with various designs and ideas for substitutes. One interesting example is a modification of AKG K702 headphones with custom earpads that are cut out in front and in the back (see “DIY Modifications for Acoustically Transparent Headphones”). I don’t have the K702, but I do have the K701 which is really the same model, just without the detachable cable. So along with the K1000 and the PFR-V1 I will try these as well.

I emphasized “acoustic headphones” in the paragraph above because another well-known way of having ears unoccluded is to use bone conducting headphones. I considered them initially, but then rejected them for a number of reasons:

  1. I don’t know how to calibrate them properly, as bone conduction works differently from acoustic transmission.

  2. They usually use wireless (Bluetooth) connection which adds a lot of latency.

  3. According to this research, “BC [bone conducting] earphone can’t provide enough interaural level difference (ILD)” and thus can affect perception of phantom sources, too.

But what about all these wireless earbuds with the “transparency” feature? Sure, they have lots of microphones inside and outside, and a DSP. In theory, they could implement an acoustic pass-through that is indistinguishable from wearing no earbuds at all. However, the question is why anyone would need that (apart from VR researchers). The goal of engineers working on wireless earbuds is to ensure that the user does not get hit by a car while listening to their favorite podcast, and that the user can communicate with a flight attendant while having airplane engine noise in the background—that’s it. For these scenarios, earbuds don’t actually need to replicate the true transfer function of an unoccluded ear. In fact, the DSP may artificially boost up certain frequency bands in order to improve the external voice clarity, which is the opposite of what I need for my research.

We see that achieving high-fidelity acoustic transparency is non-trivial. So the goal of the exploration described here is to measure the key aspects of earspeakers in order to ensure that I’m aware of their shortcomings. Also, since I bought the K1000 and PFR-V1 second-hand, they are a bit old (it’s practically impossible to buy them in new condition because they have been discontinued) and thus may have some aging-related problems. Another task is to check how the “transparentized” K701 compares to real earspeakers.

This is what I measured for all three headphones:

  • electrical impedance;
  • usable frequency range;
  • distortion;
  • cross-talk;
  • acoustical transparency.

Also, since I tried using miniDSP EARS (or is it HEARS?—that’s what is written on the label that I see on my device) for some of the measurements, there are some notes about this peculiar acoustic tool.

Electrical Impedance

This measurement is needed to ensure that the drivers are matched. For this, I used QuantAsylum QA490 unit. My AKG K701 shows good alignment:

However, in the PFR-V1 due to its age the left coil shows lower impedance than the right one:

Thankfully, the difference in the impedance seems to be constant across the frequency range which means we can compensate for that by increasing voltage (volume level) into the right speaker. I have found that I need to add +2 dB to the right channel.

The AKG K1000 is more interesting because it uses a 4-pin XLR plug. I don’t have an adapter from it to a TRS. I first tried to use QuantAsylum QA460 (a speaker amplifier), however I found that the nominal impedance of the K1000 is 120 Ohm, and that’s too high for the default value of the current sensing resistor in the QA460—0.01 Ohm. QuantAsylum’s Matt recommends that the resistor value should be 100 times less than the load.

I did not want to resolder the resistor just for this measurement, so instead of the QA460 I used Dayton Audio DATSv3. It worked fine and has shown that the K1000 also has an imbalance between its speakers. Unlike the PFR-V1, the imbalance of the K1000 is not uniform and is mostly pronounced at mid to low frequencies:

Sorry, DATS uses a bit less readable palette. So, for the K1000 I ended up adding a low shelf filter at 2.36 kHz, Q 0.8, gain +2 dB, and a peak at 3.3 kHz, Q 5.5, gain +3 dB. All this alignment was checked acoustically on miniDSP EARS and Neumann KU-100.

Note that although I read in many places that AKG K1000 are “hard to drive”, and sometimes people have to connect them to speaker amplifiers, I did not have any issues with driving them. I used two headphone amplifiers: the SMSL SP200 with “high gain” setting driven from the XLR input, and Monoprice Monolith (THX AAA 788) driven from its digital input. Neither of the amplifiers had any problems driving the K1000 to 100 dB SPL and even higher. That’s the level which is enough for my needs since that’s 100 dB near your ears, not at a speaker one meter away.

Usable Frequency Response

Since earspeakers operate in a free-field condition—they do not create a pressure chamber on the listener’s ear—all constraints of loudspeaker drivers apply to them. Due to the small surface area of the driver, they can’t be efficient at low frequencies. Typically, small (3"–4") loudspeaker drivers that must produce bass are designed to have wider excursion so that they can push and pull air more efficiently. However, it’s hard to design a miniature headphone driver this way.

In an attempt to overcome this limitation, Sony PVR-V1 has a “bass duct” which is intended to be placed near the ear canal entrance:

This creates challenges for measurement because artificial pinnae are simplified and are made of silicone thus their retention force is much lighter that of a real ear. So I tried measuring them both on the miniDSP EARS and the Neumann KU-100. The results are actually consistent and show severe bass limitation:

Note that for this measurement I used the miniDSP EARS with “raw” calibration file. I did not intend to produce a measurement comparable with data from other rigs, I just wanted to figure out where they start to roll off low frequencies. The result looks similar to free-field measurements performed long ago by Rin Choi, modulo the peak between 4–5 kHz which comes from EARS. Basically, the PFR-V1 earspeakers are only usable down to 650 Hz (-6 dB from the midrange) as trying to compensate for bass loss via equalization will just drive up distortion in the driver.

The driver of the AKG K1000 is much better, and its bass starts to roll off only below 90 Hz. Although Hans Zimmer would not endorse use of these headphones for listening to his music, this range is probably enough for classical music.

This is also obtained on EARS with “raw” calibration. Note that apart from the lack of standardized calibration, both miniDSP EARS and Neumann KU-100 are perfectly adequate for measuring earspeakers because their operation does not depend on replicating the true impedances of a human ear canal.

The AKG K701 with cut out earpads loses its sealing completely and thus starts to roll off the bass early:

For comparison, this is the FR of the left speaker for all three earspeakers:

All in all, I think the K1000 is the most adequate earphone in terms of the usable frequency bandwidth.

SPL Calibration of miniDSP EARS

Yet another myth I had read, apart from the K1000 being hard to drive, is that the miniDSP EARS is impossible to calibrate for loudness. This is how I performed their SPL calibration. I found that the silicone ears can be easily removed, opening access to the microphone capsule:

The diameter of its wrapping grommet matches the size of the Beyerdynamic MM-1 microphone so I used its adapter for coupling the acoustic calibrator:

I found that the capsules are well-matched. However, the sensitivity factor specified in the calibration files is wrong which makes the SPL level reported by REW to be off by about 20 dB. I have switched the EARS to 0 dB gain and edited the calibration file to specify correct sensitivity factor: 2.7 instead of 0.1 dB.

One interesting thing that I did not know before is that on Windows at least, REW monitors the software microphone gain and adjusts the reported SPL number accordingly. This is a good thing because it allows the user to adjust the gain so that the sound captured by the EARS at high SPL does not get clipped in the digital domain.

Distortion

I was mostly interested in whether my earphones have any significant distortion that can affect loudness perception. By “significant” I mean “higher than the room noise level.” Since earphones offer no isolation from the room noise, this criterion works well. I ensured that the room noise was 42 dB or below (C weighted) while I was measuring.

When measuring using the KU-100, the distortion products for all three headphones were below the noise floor in the mid-to-high frequency range. I couldn’t measure bass distortion adequately due to higher level of ambient noise in this region. Since with the open design the drivers have to work hard, I would expect that they distort. However, it is known that our hearing system is much less sensitive to low-frequency distortions, and they also do not play an important role in my experiments, so that’s not a problem.

However, with the miniDSP EARS there were peaks of the 2nd harmonic between 4–5 kHz. It was suspicious to me that this happens the same way for all three headphone models. I also tried placing Genelec 8331A speaker close to EARS and measuring it, and it has shown the same peak. So I think what happens here is that the construction of the silicone pinnae of EARS creates a resonance in that region and that overloads the microphone even when the resulting peak is at modest 100 dB SPL. That means, EARS are not suitable for measuring headphone distortion.

Cross-talk

That’s a practically interesting measurement. Since earphones are “open” designs, there is cross-talk by definition. Will this create a coloration of the phantom center? I measured it on the KU-100 in order to create a realistic head shadowing condition.

What I have found is that the level of cross-talk is actually quite low. If we consider the cross-talk from my speaker setup and hypothetically reduce it by what can usually be achieved with application of Cross-Talk Cancellation (CTC) in a non-anechoic room, the cross-talk from earspeakers is still below that level.

By the way, while exploring this, I have discovered a nice feature of REW that I did not know about for all these years of using it. On the SPL and phase graph it’s possible to select a frequency range by dragging the mouse cursor with Shift key pressed, and this shows min, max and average SPL for that region. So calculating the average cross-talk level consists of two steps:

  1. Take the difference in dB (A / B) between the contra-lateral and ipsi-lateral transfer functions (“Trace Arithmetic” in REW). Since there are HRTF-induced notches and peaks all over the frequency range, applying 1/3-octave smoothing makes sense.

  2. On the calculated SPL graph, select with Shift the frequency range of interest (I usually use 100–14000 Hz to avoid areas where SNR is lower) and read the average SPL.

If we look this way on the SPL graph of the difference between the ipsi- and contra-lateral sound for my desktop speakers (measured on the KU-100), it shows the effect of the head shadowing (the more shadowing—the better). The average level of shadowing is about 11 dB, and it is slanted towards high frequencies:

Below is a table comparing the average shadowing among earspeakers:

Earspeaker Avg. shadowing, dB
AKG K1000 19
AKG K701 DIY Transp. 48
Sony PFR-V1 36.5

If we consider the PFR-V1, their shadowing is 25 dB better than on external speakers. Considering that CTC algorithms for loudspeakers can at best remove about 20 dB of cross-talk, Sony already performs better.

The AKG K1000 is the worst performer, probably due to large size of speakers and completely open design. On the graph below we can see that at 3 kHz its shadowing drops and matches the shadowing from speakers:

(The numbers on the legend correspond to the 3 kHz point).

So, probably it is a good idea to apply CTC to the K1000. It appears especially effective because traditional nearfield CTC requires head tracking since the head is constantly moving relative to the speakers. However, since the K1000 always moves together with the head, the CTC filter remains constant.

However, so far I did not succeed with creating a CTC filter for the K1000. I tried using a method based on actual HRTF measurement. That’s because the proximity of the earspeaker to the head makes the transfer function of ipsi-lateral crosstalk highly dependent on the parameters of the head.

Also, unlike the louspeaker situation, the wavefront which reaches the opposite ear is not planar—it is spherical. Because of that, when measuring the actual transfer function at the opposite ear of the KU-100 I can see a lot of group delay deviations, and the resulting impulse response does not even have any distinctive peak—it looks more similar to ripples on a sea surface. Because of that, achieving correct time alignment is insurmountably challenging. I will try harder next time.

Acoustical Transparency

This is an important consideration since I plan to switch between playback in earspeakers and loudspeakers (that’s the whole point of using earspeakers in the first place!). As I mentioned in the beginning, even open-back headphones create significant alterations to the frequency response of external sources, as well as to their ITD and ILD. Earspeakers are a bit more transparent but not entirely.

A paper by C.  Porschmann “How Wearing Headgear Affects Measured Head-Related Transfer Functions” measured how wearing the AKG K1000 affects the HRTF of the KU-100. I also measured my earphones on the KU-100, using Genelec 8331A as an external sound source:

I compared the measurements from the paper with mine (using the SOFA files they have provided) and discovered that although the IRs for bare KU-100 look quite similar to my measurements—modulo the effect of a different distance from the speaker to the head—the measurements with AKG K1000 on the head look significantly different. The peaks and notches simply do not match. I started exploring this discrepancy and realized that the acoustical shadowing that the K1000 creates is highly dependent on the angle at which they are opened.

I think this leads to the important realization that measuring these occlusion transfer functions must always be done for the current setup and can’t be considered as “generic.” I suppose, the external speaker directivity also may heavily influence the result, thus occlusion transfer functions measured with a Genelec speaker will also differ from those measured using an LXmini. Because of that, let’s consider the compensating transfer functions for these three headphones just as a guideline.

Besides the data from the paper on headgear, another paper which had introduced the idea of the “transparent” K701 (“DIY Modifications…”) also has measurements of them and the K1000 on the KU-100. However, these measurements are from certain directions: front, back, and top only. Note that these two papers calculate the transparency in an opposite way. The headgear paper divides “reference” (bare head) HRTF by the HRTF of the head with the gear on, while the K701 paper carries the division the other way around. In my view, since we need to compensate for the occlusion by applying a reverse filter to the external speaker, we need to follow the headgear paper approach.

Below are occlusion graphs for each earspeaker, for frontal (0°), side (42°), and rear (138°) directions. For side directions, this is for the ipsilateral ear only. Also, for compatibility with the graphs from papers, my graphs are time windowed for 3 ms and 1/3 octave smoothing applied.

We can see that Sony PFR-V1 is the most unoccluding earspeaker, with AKG K1000 coming next, and the “DIY Transparent” K701 modification is actually not so transparent for rear sources.

Conclusions

Unfortunately, none of the earspeakers is an ideal one. Here is their comparison on the acoustic parameters:

Earspeaker LF cutoff, Hz Shadowing, dB Transparency
AKG K1000 ~60 19 Fair
AKG K701 DIY Tr. ~300 48 Poor
Sony PFR-V1 ~700 36.5 Good

Thus, for simulating anechoic listening on speakers, it makes sense to use the Sony PFR-V1 as much as possible, by limiting the sound sample choice according to the headphone bandwidth. One caveat here is that by cutting out this frequency range we sufficiently limit the bandwidth below 1.5 kHz where the auditory system uses ITD for sound source localization.

The AKG K1000 can be used for wider selection of samples and does not have any issues providing ITD cues; however, we need to take care about reducing its cross-talk, and compensating for partial loss of transparency when comparing its output with external loudspeakers.