I’m getting closer to experimenting with comparisons between earspeakers and real speakers. Since earspeakers (any headphones, really) are usually rather symmetric, I decided to physically change my desktop speaker setup to make it more symmetric, too. In the process, I also decided to retune my LXdesktop using the new tools added in Acourate V3 and V4: speaker linearization using a sinc pulse and the new acouStep tuning algorithm, which virtually eliminates pre-ringing in the tuned speaker response.
From my experience with the last tuning of LXdesktop, I realized that it’s much easier to tune speaker drivers by modeling them as components of a linear-phase crossover rather than a traditional minimum-phase crossover. Mathematically, modeling the drivers as minimum-phase crossover bands and then linearizing the phase afterward (as described in the Grimm Audio paper on the LS1 speaker) is equivalent to using linear-phase crossovers in the first place. In practice, however, speaker drivers that behave like linear-phase filters are much easier to time-align. Plus, the sinc-pulse linearization tool introduced in Acourate V3 makes it very straightforward to bring drivers to linear-phase crossover behavior.
The Idealized Speaker Model
Let’s state the goal first. Ideally, we want the entire speaker to behave as a linear-phase band-pass filter. Why “linear-phase”? Because linear phase guarantees constant group delay. That means that if the speaker plays a short wideband impulse, such as a click, all of its component frequencies reach the ear simultaneously. In other words, a linear-phase speaker preserves the waveform of transients in its direct sound (within its passband, of course), ensuring natural reproduction.
And why “band-pass”? Because the speaker is physically limited in the range of frequencies that it can reproduce. Even a big, high-quality speaker can’t cover the entire range from 0 Hz to infinity. First, this is not needed, because the range of human hearing is limited. Second, there are physical limits on how low and how high acoustic transducers can go. Going too low or too high requires too much energy and can easily overload the drivers. Thus, any speaker covers a limited frequency range, which technically makes it a band-pass filter.
Note that within the flat part of the passband, far away from the band edges, a minimum-phase filter also produces a similarly linear result. The defining property of a minimum-phase system is that its phase response is uniquely determined by its amplitude response (the two are related via the Hilbert transform), and so is its group delay. Where the frequency response is flat, a minimum-phase filter produces a flat (zero) phase response and a constant group delay. However, it introduces phase shifts as soon as the frequency response bends, and a band-pass filter, by definition, bends at both ends.
As a theoretical foundation for our experiment, we can use the open-access AES paper “Modeling and Delay-Equalizing Loudspeaker Responses” by A. Mäkivirta, J. Liski, and V. Välimäki, written as part of a collaboration between Genelec and the Acoustics Lab of Aalto University in Finland. The paper takes exactly this approach: it starts by modeling a speaker as an ideal band-pass filter.
This is fairly straightforward for a single-driver loudspeaker. In practice, however, most speakers use multiple drivers. Rare exceptions are speakers with a single full-range driver, and full-range electrostatic and magnetic planar speakers. All of my speakers are multi-driver, which means that their ideal counterpart, the band-pass filter, is in turn split into overlapping band-pass components. The splitting is done by crossovers. In my LXdesktop speakers, the crossovers are implemented in DSP, so they can have any desired behavior. Mäkivirta’s paper discusses the classic case of minimum-phase crossovers and describes the challenges of time-aligning them. One important property of minimum-phase crossovers is that, in a speaker with more than one crossover, they must be chained using a special topology:
This is because the phase shift created by the woofer-mid crossover must be applied to both the midrange and the tweeter drivers; otherwise, the parts of this composite crossover will not sum to a flat frequency response. Note that since the cutoff of the woofer-mid high-pass filter lies far below the tweeter’s band, this filter has no effect on the amplitude of the tweeter signal; it only affects its phase.
Linear-phase crossover filters do not have this issue because their effect on the phase is strictly linear: they only produce a pure time delay. As long as all the filters share the same delay (that is, their impulse responses are centered at the same sample position), they can be arranged in a simpler way:
The basic idea behind driver tuning stays the same, though. We do not consider the crossover filters and the speaker drivers separately. Instead, we try to coerce each speaker driver into becoming an ideal crossover component by developing a correction filter for it, based on the transfer function of the corresponding part of the ideal crossover.
Because of the extra complexity of minimum-phase crossovers, I use linear-phase crossovers in my project. In this case, each driver is coerced into its crossover component behavior using Acourate’s sinc-pulse linearization. This process takes two inputs: the actual response of the driver measured close to it, including the effects of diffraction at the baffle edges (note that a true near-field measurement, with the microphone right at the cone, would exclude them), and the desired crossover impulse response. The output is a filter that brings the driver as close as possible to the desired behavior.
The Catch at the Low End
There is one important caveat that Mäkivirta’s paper explains. Since neither the whole speaker nor its lowest-frequency component (in my case, the subwoofer) goes down to 0 Hz, there is a roll-off at low frequencies. For a subwoofer, or any large driver, the roll-off frequency can be quite low; let’s use 15 Hz as an example. The caveat is that if we represent this roll-off as a linear-phase filter, the filter has massive pre-ringing. If we represent it as a minimum-phase filter instead, group delay is inevitable because, as I noted earlier, in a minimum-phase system amplitude changes are unambiguously coupled with phase changes.
Let’s look at both alternatives. Below is the frequency response of a 4th-order Butterworth high-pass filter at 15 Hz:
And this is the RMS view of the impulse response (IR) of its linear-phase version:
We can see pre-ringing that lasts about 300 ms. Although its energy is concentrated around the cutoff frequency, where hearing is not very sensitive, with a capable subwoofer it is likely audible on bass-heavy transients, sounding like “breathing.” So, it is much more natural to use a minimum-phase filter for the high-pass part (after all, speaker drivers naturally tend to be minimum-phase devices). By definition, minimum-phase filters have no pre-ringing; however, the minimum-phase 4th-order Butterworth high-pass filter has a significant group delay, which reaches about 22 ms at 20 Hz:
Psychoacoustic research has shown that the human ear is relatively insensitive to the late arrival of low frequencies relative to the initial wavefront, especially with music and in typical rooms, where reverberation time at the low end can be quite long. For references, see another paper from the same Genelec–Aalto University collaboration, “Audibility of loudspeaker group-delay characteristics”, and Section 9.6.5 of the 4th edition of F. Toole’s “Sound Reproduction” book. The listening tests described in the paper found that group delay can exceed 10 ms below 200 Hz without the difference being audible, whereas in the 300 Hz–1 kHz range, differences became audible above just 1–2 ms.
Going back to the linear-phase high-pass filter: its 300 ms of pre-ringing far exceeds the span of pre-masking (backward masking) in human hearing, which, according to the studies reviewed in the same paper, is only about 5–20 ms. That explains why this pre-ringing is audible.
Practical Tradeoff: A Mixed-Phase Filter
Based on this, it is fine to keep the minimum-phase 4th-order Butterworth roll-off at low frequencies. This means that even our idealized, theoretical model of the speaker is not completely linear-phase. Instead, it is mixed-phase: it has minimum-phase behavior at the low end and linear-phase behavior above it. Mäkivirta’s paper comes to a similar conclusion: it is simply not practical to coerce the speaker into linear-phase behavior at low frequencies.
Here are the characteristics of our idealized speaker model. The frequency response is tapered at both the low and high ends:
The group delay is flat everywhere except at the low end, where it is determined by the Butterworth high-pass filter:
(Note that since this is a linear-phase filter for a 48 kHz sampling rate, centered at sample position 32,768, the group delay baseline is about 682 ms.) And this is the step response:
From this distance, it might look like a minimum-phase response. However, if we zoom into the onset, we can see some minimal pre-ringing from the linear-phase low-pass filter:
Its duration is well under a millisecond (a few periods of the cutoff frequency), and it is thus covered by the pre-masking mechanisms of human hearing.
Examining Linear-Phase Studio Monitors
As described in the Grimm Audio paper on the design of the LS1 speaker, its designers chose a minimum-phase LR4 crossover and then corrected its phase by applying an anti-causal all-pass filter (the sum of an LR4 crossover is a 2nd-order all-pass with a Q of 0.7, and its inverse is a non-causal filter, which is implemented as an FIR with added delay). Let’s see what other makers of linear-phase studio monitors do. I measured the behavior of two professional studio monitors: the Genelec 8331A and the Dutch & Dutch 8c.
The anti-causal all-pass filter introduces extra latency: from 4 ms to 30 ms, depending on the frequency of the crossover being compensated. Since this may be critical in live monitoring scenarios, manufacturers usually keep the natural minimum-phase behavior available, calling it the “low latency” mode.
Also note that this type of correction is only practical with DSP because it involves non-causal (look-ahead) filters, which can only be realized by adding delay. As a consequence, although both of these studio monitors accept a traditional analog input, it is there only for legacy compatibility. Their native input is digital, so the analog signal first goes through analog-to-digital conversion. Because of that, both monitors have a baseline latency of 3.2 ms, even in the “low latency” mode.
Genelec 8331A
This is a small desktop-size monitor with a bandwidth of 45 Hz to 37 kHz. The low end is extended by means of a port. Two notable design features of this speaker are the coaxial mid-tweeter driver and the two “racetrack”-shaped woofers at the top and bottom of the front panel, hidden under a cover.
This is the smallest speaker in Genelec’s “The Ones” series. The next model up (8341A) uses the same mid-tweeter driver, while the two larger models (8351B and 8361A) use a slightly bigger one. As the speakers get bigger, so do their woofers. This makes the largest model in the series look a bit unusual, earning it the nickname “Cyclops.” Even this big guy only extends down to 30 Hz, so covering the full audio range requires a subwoofer or the specially designed W371A, a stand with an integrated woofer. See the full specs here.
One caveat with these speakers is that to get full access to their features, including the ability to switch between the “low latency” and “linear phase” modes, the user needs to buy a proprietary “Loudspeaker Manager User Kit” (GLM), which alone costs as much as a mid-priced speaker:
To characterize the linear-phase correction filter, I measured the speaker in the same physical setup in the “low latency” and “linear phase” modes. Then I could simply divide the obtained transfer functions (complex division in the frequency domain, that is, deconvolution of the impulse responses) to derive the correction filter. Below are the phase and group delay differences between the two modes of the 8331A (the frequency response stays the same):
(Note that the group delay is contaminated by room interaction.) We can see that, similar to the Grimm Audio LS1 approach, the “linear phase” mode adds an anti-causal all-pass filter that corrects the phase deviation caused by the minimum-phase crossover. Genelec does not reveal which type of crossover the speaker uses. The manual only specifies the crossover frequencies: 500 Hz for woofer/midrange and 3 kHz for midrange/tweeter (both of the latter are part of the same coaxial driver). The all-pass filter applies up to 540° of phase rotation, so the crossover might be a combination of an LR4 (360° of phase rotation) and an LR2 (180°). I modeled an LR4 + LR2 crossover pair at the 8331A’s crossover frequencies. Here is the phase of the compensating all-pass filter for this model, overlaid with the actual filter of the 8331A:
The slope of the phase is very similar, but the details differ. This
might be because Genelec uses different crossover types, and also
because they employ a cleverly engineered approach with windowed IIR
filters instead of a straightforward inverse FIR filter. Their design
aims to achieve the lowest possible latency, even in the “linear phase”
mode. According to the manual, this mode adds only
6.9 – 3.2 = 3.7 ms of extra latency. The RMS view of the
all-pass filter’s impulse response shows that the pre-ringing is at
-60 dB at approximately -3.7 ms
relative to the peak:
This means that the extra delay is due to the filter itself. In particular, there is no evidence of delay caused by FFT buffering, which would be inevitable if the designers used a classic FIR filter applied via frequency-domain processing.
The choice of crossover frequencies also helped keep the filter short. However, I think that the 500 Hz woofer crossover frequency was actually dictated in large part by the physical arrangement. Since the woofers sit behind the front baffle, which acts as a waveguide for the coaxial driver (see this paper on the design of “The Ones”), they radiate mostly upward and downward, not toward the listener. Thus, the designers had to keep the woofers in the frequency region where their radiation stays omnidirectional.
We also need to confirm that the low end of the speaker stays minimum-phase in the “linear phase” mode. Indeed, it does. Since my group delay graph is contaminated by room interaction, here is a clean group delay graph of the 8331A in this mode from Genelec’s site:
The group delay starts to increase below 200 Hz, reaching 30 ms at 50 Hz, which is totally acceptable according to the research cited above.
Dutch & Dutch 8c
This is a big, full-bandwidth speaker (20 Hz to 20 kHz). Full-bandwidth output is achieved by integrating two rear-firing 8-inch subwoofer drivers in a sealed enclosure. They are designed to couple with the front wall (the wall behind the speakers), so the speaker should be installed close to it, about 20–80 cm away. This is an unusual recommendation because, normally, placing a speaker near a wall causes speaker-boundary interference response (SBIR) problems at low-mid frequencies. However, the clever design of this speaker mitigates this issue in two ways. First, the midrange woofer sits in its own sub-enclosure with acoustically resistive slots on both sides of the cabinet. The slots let out the delayed and attenuated rear radiation of the driver, which cancels the front radiation toward the back. This makes the woofer’s radiation pattern a cardioid over its entire working range (100 Hz and up). See the measurements in Erin’s Audio Corner for full details. Second, since the subwoofers themselves are so close to the wall, the first SBIR notch falls well above their 100 Hz cutoff, in the region covered by the cardioid midrange. Note that SBIR caused by large surfaces in front of the speaker (the floor!) is still present.
The acoustic design of this speaker is great, and so are its electronics. It features an Ethernet connection, which allows accessing its DSP controls from a smartphone app and manipulating the correction filters directly from Room EQ Wizard. Switching between the “low latency” and “linear phase” modes is done from the app. Let’s examine how the speaker’s response differs between the modes. As with the Genelec, the frequency response stays the same; only the phase, and consequently the group delay, changes:
(The ripple on the group delay graph is exactly the floor reflection I mentioned earlier.) Compared to Genelec’s all-pass filter, there is more phase rotation here: roughly two full turns. Unlike Genelec, Dutch & Dutch does not conceal the details of its crossover. The official documentation specifies that there are two LR4 crossovers: one between the tweeter and the woofer at 1250 Hz, and another between the woofer and the subwoofers at 100 Hz.
Just to confirm, I modeled the same configuration of minimum-phase crossovers and derived the compensating all-pass filter. It matches the measured all-pass filter of the 8c:
Due to the much lower crossover frequency, the all-pass filter’s impulse response is necessarily longer. Dutch & Dutch specifies that it adds 30 ms of delay (33.2 ms total versus 3.2 ms in the “low latency” mode). Interestingly, the RMS view of the filter’s impulse response suggests that it could probably be windowed to 8.5 ms:
(This graph overlays the measured response of the speaker’s all-pass filter, in green, with the theoretical one, in purple, which looks much cleaner.)
The designers may have chosen 30 ms either to achieve better frequency resolution of the filter, or simply because this is the granularity of their DSP processing (Martijn Mensink, the designer of the speaker, discusses the system’s latency here).
Preliminary Conclusions
Representing a loudspeaker as a mixed-phase band-pass filter makes it possible to achieve flat group delay over almost its entire working range, except for the bass region, where the ear is insensitive to group delay deviations.
Modern professional DSP-based studio monitors offer a tradeoff between the low latency of minimum-phase behavior and the better sonic precision of linear-phase behavior. Since linear-phase behavior inevitably introduces extra latency, it cannot be the default and must be enabled manually in the control software.
For LXdesktop, though, I don’t need to support a low-latency mode, so I can start with linear-phase crossovers, which make driver time alignment easier, especially without access to an anechoic chamber.





















No comments:
Post a Comment