Saturday, July 6, 2024

LXmini Desktop Version (LXdesktop)—Part II: DSP Tuning

This post continues my story about the desktop version of LXmini speakers that I have built and set up on my computer desk in a somewhat unusual way:

So, why are the speakers are "toed out"? The idea is that since the full range driver has a dipole dispersion pattern, if we turn it outwards, then the null of the dipole becomes directed towards the opposite (ipsilateral) ear, thus naturally contributing to the suppression of the acoustic cross-talk between speakers. This effect these days is usually achieved using DSP by injecting a suppressing signal into the opposite speaker (see a great post by Archimago and STC on this topic). However, it would be nice if the opposite ear would be just naturally blocked hearing the sound from the speaker.

I've estimated the angle between the full range driver and the opposite ear to be approximately 75°, thus the suppression is not maximal. However, it should still add extra -5 to -10 dB attenuation to head shadowing, depending on the frequency. I plan to measure the exact attenuation profile some time later. Another feature of setting the speakers this way is that the back of the speaker gets farther from the back wall, at about the recommended minimum of 1 meter.

Ideas for Tuning

Since the original LXmini tuning was aimed to achieve flat response on-axis (see the design notes), my unusual speaker arrangement required a dedicated tuning. I started looking around for ideas on to achieve close to ideal response in the time domain.

The author of Acourate Dr. Brüggemann holds a very strong position on using linear phase crossovers. Acourate can generate various kinds of crossovers, both in minimum phase and linear phase versions. Also, there are some tools (including a new one added in the recent Version 3) which are intended to bring each driver as close as possible to the corresponding band pass filter of the crossover, both for amplitude and the phase. Together with proper time alignment of the sound from each driver at the listening position, this allows to achieve "ideal" summing of the acoustic crossover components, yielding the perfect Dirac impulse response for the speaker as a whole.

Though, my initial concerns were about the pre- and post-ringing behavior of the linear phase filters. As we know, they are symmetric around the center, and the pre-ringing may potentially exceed the thresholds of masking. When the components of a linear phase crossover sum up as intended—with their peaks coinciding, the pre- and post-ringing components from each crossover band cancel each other. However, if there are time shifts—even as small as a fraction of a millisecond—this does not happen. The example below is for a two-band linear phase Neville Thiele crossover:

This is how the summed impulse response looks like on the logarithmic scale when the components are properly time aligned, and also for 0.23 ms and 0.5 ms time of arrival difference:

The red vertical line is the ideal IR which occurs in the ideal time alignment case, and on the right are the IRs when one of the crossover components is shifted. Recall that these delays correspond to a distance difference of just about 7.88 cm and 17 cm—that's comparable to the size of the human head.

When I started discussing this topic on the Acourate forum, one of the members has pointed me out to the white paper by B. Putzeys and E. Grimm on their ideas behind the DSP-based implementation of the professional Grimm Audio LS1 speaker (which costs quite a lot!). The authors used a minimum phase Linkwitz-Riley filter, but compensated for its phase deviations using an inverse all-pass filter. If we think about this approach, it effectively also yields a linear phase filter. In fact, when crossover components get time shifted, the combination of the crossover plus reverse all-pass filter also exhibits pre-ringing, although its level is a bit lower, and what's more important, the duration is shorter:

(Note that the red IR is not an ideal Dirac pulse because although the phase response of the all-pass filter I created is close to the phase response of LR4, it is not exactly the same). However, these improvements over the ringing of the Neville Thiele crossover are just due to the fact that the LR4 crossover has more relaxed slopes to start with:

Thus, instead of compensating for phase deviations of a minimum phase crossover, which can be quite severe for high order crossovers, we can as well just start with linear phase crossovers as they are much easier to work with. For example, I wanted to use an asymmetric shape in which the higher frequencies driver has more relaxed slope compared to the lower frequencies driver. This is beneficial for the LXmini design because the directional pattern of the full range driver yields more precise spatial cues than the omnidirectional woofer. This approach also helps for the pair of the woofer and the subwoofer because I only have one, so I would like to experience a stereo bass as much as possible. The asymmetric shape of crossover slopes at first yields a non-flat summed frequency response, however this is easy to compensate (again, with a linear phase filter), thanks to the fact that the phase shift, being always equal to zero, does not affect the summing of amplitudes of the crossover components.

Another interesting observation. The fact that I'm performing the tuning in a real room, not in an anechoic chamber, implies that I need to use windowing of the measured frequency response. As I have realized after brief experiments, the frequency dependent windowing (FDW) partially suppresses pre- and post-ringing of linear phase filters. However, as a result it also changes the shape of its frequency response by making it less steep. In my opinion, this is a good trade-off. In the next section I will show the shapes and IRs of the linear phase crossovers I have ended up with.

Crossover Preparation Details

The aforementioned Grimm Audio LS1 white paper has a suggestion on "ideal" crossover points. From the psychoacoustics data, the authors state that the directional pattern of the frequency response should be used down to 300 Hz. The original LXmini has its acoustic crossover point closer to 790 Hz, however it uses a 2nd order LR crossover thus the output from the full range driver actually goes quite low in frequency range. So the first thing I've done was to measure the raw response of the full range driver. Here it is together with an FDW processed version:

Looking at the natural roll-off of the driver I have chosen 366 Hz as the crossover point. At the high frequency end, the full range driver due to its relatively large size starts working in a breakup mode, thus losing efficiency. Plus, I'm not listening to it on-axis and that creates a natural roll-off at high frequencies. However, that's not a problem. Since the speakers are located quite close to my ears, there is no need to try to make the frequency response to be ruler flat at the high frequency end because that makes the sound too harsh. So I generated a LR2 linear phase crossover for 11 kHz and used its low frequency part to taper the response of the driver on the right side. This is how the final crossover component looks like, overlaid with the raw windowed response:

Similarly, for the woofer driver I have chosen 46 Hz as the crossover point. The slope on the left side is LR4, however on the right side I used Neville Thiele 1st order crossover as it has a sharp, "brick wall" slope. I passed it through the same frequency dependent window that I use for the in-room measurements, and this has made the shape of the slope more "relaxed". Below for comparison are the original NT1 slope overlaid with a windowed one:

There is not much difference in the time domain though:

And this is how the designed crossover component looks on top of the raw driver response:

The subwoofer was a bit interesting. Choosing the crossover did not require any thinking because the crossover point was already set from the woofer driver, and the type on the right side is also Neville Thiele 1st order. However, since it's an active subwoofer with servo (Rythmic F12G), it has some settings of its own. I experimented with different damping settings and low-end extension, and found that low damping and the extension down to 14 Hz creates a time domain response which looks close to the IR of the crossover if I invert its polarity. This is how these IRs look like overlapped (the polarity IR of the subwoofer is inverted):

And this is the final look on the crossover components that sum up into a flat frequency response (with the high frequency range trimmed down) and a zero phase response:

Visually this crossover reminds of the Bessel low-pass filter (used in the "RBessel" crossover type in Acourate) of a high order, however mine uses even steeper slopes on rights sides.

Driver Tuning Process

My tuning process has two major stages: the first to bring each driver as close as possible to the behavior of the corresponding band pass filter of the crossover (that also includes fixing the phase behavior), and the second stage is to combine these drivers into a proper acoustic crossover.

I was doing all the measurements from the single position—the listening position. Although it is possible to linearize drivers in the near field, I did not use this approach due to two reasons. First, the full range driver works as a dipole, and they must be measured from some distance. Second, since I was interested in the performance of the crossover at the listening position, this was the natural position to use for driver linearization as well.

For the driver linearization I used the "Room Macros" of Acourate, setting the "Target Curve" to be the desired crossover band pass behavior. Obviously, I used the same window for the FDW of the measured driver response as the one I used to process crossover parts during the preparation stage. I did not use "Psychoacoustic" smoothing at the driver linearization stage, instead I used more technical "1/12 Octave" smoothing. I was also limiting the amplitude correction to avoid creating a boost at the frequency bands where the response of the driver was naturally decaying below the intended crossover suppression level. As an example, below is the correction filter for the woofer driver, overlaid with the target:

After the correction filter has been generated by "Room Macro 4" and the result has been evaluated via a test convolution, I re-measure the driver with the filter applied. Then I check the phase behavior. Since the correction process of Acourate tries to bring the driver to the minimum phase behavior, it will leave out phase deviations that present in the minimum phase impulse response of the target curve. Note that when equalizing an entire full-range speaker to a mostly flat target curve, these phase deviations will end up outside the hearing range. However, for a driver, since it has a limited frequency range the phase deviations will typically end up near crossover frequencies, and this fact will make proper time alignment more problematic. For example, this is the phase response of the corrected woofer:

We can see that the phase gradually deviates from zero and "flips" over the 180° angle at 47 Hz. I treated these phase deviations using the same approach as in the Grimm Audio white paper, which are in essence the same approach as the one described by Dr. Brüggemann in his post "Time alignment of drivers in active multiway speaker systems" on the Acourate forum. That is, we need to "guess" an all-pass filter which has a similar shape as the form of the phase deviation of the speaker, and then put its reversal into the correction chain (that effectively means, we need to convolve the reverse all-pass filter with our existing filter). For example, for the woofer the corrected phase behavior looks like this:

Obviously, since it's an all-pass filter, the amplitude remains the same. There shouldn't be more than 1 or 2 all-pass corrections needed. Only the area within the driver working range must be corrected, and we must look at the windowed response to avoid correcting for the effects from reflections that very dependent on the mutual distances between the driver, the reflecting surface, and the measurement point.

Now with each driver being brought as close as possible to the desired crossover band-pass filter behavior, we need to "assemble" them into a speaker by aligning their levels and times of arrival. To do that, first I measured the speaker as is, and did a rough correction of driver levels. Then I used the the sine wave convolution approach first for aligning the full range driver with the woofer, and then the woofer with the subwoofer. At low frequencies, the convolved sines may initially be considerably shifted from each other. Also, the low frequency filter may be developing a bit slowly and have irregular sine amplitudes in the beginning. To ensure that the resulting time alignment of the drivers is proper, I had applied the same sine wave convolution step to the crossover components and used the produced overlapping picture as a reference. For example, this is how the sine waves of my crossover look like for the 46 Hz point:

And this is how the results of sine wave convolution was looking initially for the woofer and the subwoofer:

Compared to the image before, it becomes obvious that the subwoofer (the blue curve) needs to be shifted ahead in time of the woofer for a proper alignment.

After applying gains and delays to the driver filters, I have made another measurement and double-checked that the sine convolution on the measured IRs produces the expected result.

Target Curve Adjustments

Life would be too easy if we could just take the summed crossover response and use it as a target for the overall speaker tuning. I tried that first and was not impressed with how it sounded. The first problem was that the vertical positions of virtual sources were too high while I would prefer having them at the eye (or ear) level. The second problem was overall lack of "weight" in the sound. The target curve was definitely asking for some adjustments.

The first problem is a consequence of the fact that any virtual source, for example a rendering of the singer's voice, which is appearing to be in front of the listener, is created by a pair of stereo speakers that are physically located on the sides. In my case, the speakers are placed even wider than the conventional "stereo triangle." As S. Linkwitz explains in the paper "Hearing Spatial Detail in Stereo Recordings", if we consider the sound pressure on a very crude approximation of a human head—a sphere—we will find that physical sources located in front of the sphere and on the sides of it create very different sound pressure distributions across the frequency range. A more precise description of this distribution is of course the HRTF. Since the two audio streams that represent the virtual central source arrive from the sides, they do not have a proper frequency profile of a center source, and as a result, the hearing system places this virtual source higher. A simple solution used by Linkwitz is to apply a shelving filter which compensates for this effect.

And the second problem—overall lack of weight, or a bass-shy presentation from a flat target curve can be explained by the interaction with the room. Running a bit ahead, below are comparisons of the speaker quasi-anechoic response (FDW windowed) vs. the steady state room response, obtained from the same measurement position by taking an RTA measurement of pink noise playing continuously:

We can see that the room "eats" the bass but amplifies high frequencies. That's why adding more bass to direct sound as well as tapering the high end seem to make sense. So after some experiments with well recorded tracks, I have chosen the following target curve:

On this graph it is compared to the initial "tapered flat" crossover curve.

The Final Correction and Measurements

The final step in the tuning process is to apply "Room Macros" to the entire speaker using the target I have created. This time I used the "Psychoacoustic" smoothing. This step fixes any remaining discrepancies in the levels of the drivers. Below is the FDW response of the speakers after applying the correction, overlaid with the target:

And below is the phase response of the speaker—as we can see it is indeed close to the "zero phase" (this is also the windowed version which excludes phase deviations due to reflections):

I checked the group delays by using the "ICPA" function of Acourate ("Room Macro 6"), and found only one very high-Q group delay deviation, not worth correcting.

The step responses of the speakers look good:

Note that these are responses without any windowing, so they do not look fully identical due to reflections and asymmetry of the room. It can be clearly seen from the Energy Time Curve (ETC) graphs produced by RoomEQ Wizard (REW):

Since this is a small room, strong reflections start appearing quite early, but it's hard to do anything about that because there are windows behind my listening position—I can't put any acoustic treatment there.

Also using REW, I checked the distortion measurement and observed the known issue with Seas FU10RB drivers of the raised 2nd harmonic distortion level between 1 and 2 kHz, also noted in the "Erin's Audio Corner" review when he was measuring the LXminis:

Also there is a bit more distortion between 300–500 Hz probably because the full range driver is being pushed harder. The distortion in the right speaker around 100 Hz due to an interaction with a room mode—if I move it to a different position, this peak disappears. And I'm not sure why each harmonic trace ends up with a funny upwards curve—this must be a measurement artifact.

The resonances from room modes can be seen on the spectrogram:

I decided to order some more bass traps, will see if they actually help to reduce the effects from room modes.

Does Non-Ideal Summing Induce More Pre-ringing?

Now let's try to get back to one question from the beginning of the post. Recall the simulations of non-ideal summing of the acoustical linear phase crossover and the associated pre- and post-ringing. I decided to check what happens in reality. For that, I have moved the measurement mic by 17 cm to the right and re-did measurement. Below are the resulting step responses. This one is for the left speaker, overlaid with the original (where the crossover components are time aligned):

Note that since the causal part of the IR is dominated by the room reflections it is not possible to judge the effect on post-ringing. As for the pre-ringing, it seems that it is actually lower in the IR recorded from the microphone position shifted off the perfect alignment.

And this is the right speaker:

We can see that for this one there is indeed a bit more pre-ringing. Evidently, the real acoustic behavior of speakers is much more complicated than these ideal models. And for a proper evaluation of the crossover behavior off-axis an anechoic chamber should be used.

Does this all matter? Maybe not so much, after all. Anyway, there is no ideal solution when we are trying to combine a full-range speaker from several band-limited drivers. If we are striving to get a perfect solution, we actually need to avoid using crossovers at all, by using a single driver, for example, an electrostatic panel or the Manger Transducer. The Manger seems to me like a variation on a coaxial driver, however due to use of a single, specially engineered diaphragm it probably does not suffer from the Doppler effect. Anyway, that's a different price level.

To Be Continued

Of course, it's interesting to discuss how this setup sounds like, however this post has already ended up being quite long. I will write about listening impressions and other things separately.

Saturday, June 15, 2024

LXmini Desktop Version (LXdesktop)—Part I: Build

I would like to make a couple of posts regarding my project which was going on for a long time—a couple of years, actually—and finally got to a satisfying conclusion. As I wrote earlier (see this post and this post) I did build LXmini speakers in their original version, and then was experimenting a bit with different crossover settings and tunings. Then after a house move, I could not use them as floor speakers anymore, and decided to make a more radical experiment and turn them into a desktop near field version. In this post I will talk about the build differences versus the original design by S. Linkwitz, and in subsequent post (or posts?) I will explain the tuning process.

Major Differences

Let's start with the two most important changes.

Height Reduction

The first obvious difference is of course the reduced height of the speakers, since they have to stand on the desktop now. I always work standing, and this puts my head a bit higher above the surface of the desktop compared to a seated position. This is actually good because the vertical stand of LXmini also doubles as the enclosure for the woofer, and thus it shouldn't be too small in order to provide enough spring back action for the woofer driver. I think there is no need to remind everyone that LXmini are designed around simple plumbing pipes. When choosing the height for the vertical stand pipe, I was aiming to get the space between the bottom of the full-range driver and the woofer at my eye level:

The idea behind this requirement is that ears are about at the eye level, and in order to preserve the distances from both drivers to the ear as the head moves closer to or further from the speaker, the ear and both drivers must form an isosceles triangle:

Preserving the distances helps to maintain the time delay between the drivers and the ear, which is important for preserving acoustical summing of sound waves from the drivers.

The length of my vertical stand pipes ended up to be 43 cm (17 inches). Obviously, the change of the working volume for the woofer driver changes its frequency response, so I could not use the original equalizer settings. However, this wasn't a problem for me since I knew that I will end up using my own tuning anyway.

The Speaker Connector

Another major change was in the speaker connector. I'm a big fan of SpeakON connectors by Neutrik. Since LXmini uses line-level crossovers and an amplifier per driver, its speaker cable has 4 wires. SpeakONs exist in various versions including the 4-pole—the model NL4FC. SpeakONs are much more convenient and also safer than traditional speaker posts for "banana" plugs which the original LXmini design uses. The height of those posts and plugs allowed putting them at the bottom of the speaker, however SpeakON connectors are significantly longer and would not fit under the speaker. Thus, I had to move connector to the back side of the speaker. This also makes the connection process a lot easier.

The receptacle for the SpeakON at the rear side also serves as a lock holding the vertical pipe and its stand together. Because of that, only one complementing bolt is required on the front side.

Minor Tweaks

Besides these major changes, I also had a couple of ideas how to make the assembly process more convenient.

Threaded Bolt Holes

The first idea was to use 1/4-20 button cap Allen head bolts everywhere possible, and eliminate the need for nuts by threading the holes in pipes instead. This change makes the process of assembling the top part of the speaker much easier. For example, one step of the assembly procedure requires centering the full-range driver at the end of the horizontal pipe by means of 3 of 4 screws. The original design uses nuts screwed to the bottom of bolts—just enough to make a flat surface:

Aligning these bolts around the driver's magnet is a cumbersome procedure due to the strong interaction between them. In my design, the threaded boltholes prevent movement of the bolts, and screwing them in while preserving alignment becomes much easier:

Another fiddly step of the assembly procedure is alignment of the horizontal pipe to make the surface of the full-range driver to be at the right angle to the surface of the woofer. The horizontal pipe rests on 3 screws that have to be adjusted. The heads of two of them are easily accessible:

The third one in the original design must be a long one, and I did not have such a long bolt, so instead I used a short one and was adjusting it via a hole at the bottom of the woofer's plate:

Pipe Clamps

I have removed the top pipe clamp since it is purely decorative, and it can vibrate at high frequencies, creating undesired resonance. The bottom pipe clamp is functional since it holds the rubber collar on top of the vertical pipe. In order to prevent its possible vibration, first I have put a piece of rubber between the loose end of the clamp steel stripe and the rest of it. Then I used a wire in order to secure the loose end as much as possible:

Alignment Post (Not the best idea :)

Finally, this is the change which did not work as expected. I usually use laser guides while setting up speakers. Since surfaces that we are surrounded with in our homes are never ideal, the only true guides for setting up things straight are gravity and lasers. With regular rectangular boxes of convenient speakers, putting a laser level on top of them is a trivial task, but with the round pipes of the LXmini design it's not. So, my idea was to add a threaded post on top of the speaker in order to mount a laser level on it.

The idea did not work out because although the laser level holds well on the post, it's practically impossible to set it up aligned properly with the speaker itself. So, I had to abandon this idea and instead put marks at the baseplate of the speaker, and this is where I put the laser level during the set-up.

Overall Look

This is how the desktop setup looked like initially (yes, the monitor has exact square aspect ratio :):

Note that it is recommended to put dipoles at least 1 meter from the real wall, however in the desktop setup this is not possible. Instead, I have acoustic absorbers mounted around the desk.

After living with this classical 60 degrees stereo setup for a while, I decided to try to improve imaging by using the directional property of dipoles in order to suppress the sound going into the opposite (ipsilateral) ear. The speakers got moved closer to my side of the desk, and were rotated away from me:

Below is a schematic diagram, a view from above:

The process of configuring the speakers for this setup will be explained in the next post, stay tuned!

Sunday, March 10, 2024

Headphone Equalization For The Spatializer

As I have mostly finished my lengthy experiments with different approaches to stereo spatialization and finally settled up on the processing chain described in this post, I could pour my attention into researching the right tuning for the headphones. Initially I was relying on the headphone manufacturers to do the right thing. What I was hoping for, is that I could take headphones properly tuned to either "free field" or "diffuse field" target, and apply just some minor corrections in my chain. You can read about my attempts to implement "free to diffuse" and "diffuse to free field" transforms applied to the Mid/Side representation of a stereo track.

Why Factory Headphone Tuning Can't Be Used

However, since in the last version of my spatializer I actually came closer to simulating a dummy head recording, I realized that using any factory equalization of any headphones simply does not work—it always produces an incorrect timbre. The reason is that my spatializer simulates speakers and the room, as well as torso reflections, and even some aspects of the pinna filtering—by applying different EQ to "front" (correlated) and "back" (anti-correlated) components of a stereo recording. "3D audio" spatializers (they are also called "binaural synthesis systems") that use real HRTFs come even closer to simulating a dummy head recording. That means, if you play the spatialized audio over headphones with their original, say, "free field target" tuning, you will apply similar human body-related filtering twice: once by the headphones themselves, and then also by the spatializer.

At some point I used to think that using headphones calibrated to the diffuse field target is fine, because they are "neutral" (do not favor any particular direction), but this assumption is also false. That's because diffuse field equalization is an average over HRTFs for all directions—thus it's still based on the human body-related filter. This is why more advanced "3D sound" renderers have a setting where you specify the headphones being used—by knowing that, they can reverse the factory tuning. I also experimented with AirPods Pro and confirmed that they change their tuning when switching between "stereo" and "spatialized" modes.

By the way, in the reasoning above I use dummy head recordings and synthetic "3D audio" spatializatoin interchangeably, but is it conceptually the same thing if we view it from the listening aspect? I think so. Dummy head recordings are taken using a real person, or a head and torso simulator (HATS) with microphones mounted at the ear canal entrances. Sometimes it's not a complete HATS, but just a pair of ears placed at the distance similar to the diameter of a human head. In any case, the microphones placed into these ears block the ear canal (in the case of HATS the canal might not even exist). Now, if we turn our attention to HRTF-based 3D audio rendering, HRTFs are also typically measured with mics placed at the blocked ear canals of a person (see the book "Head-Related Transfer Function and Acoustic Virtual Reality" by K. Iida as an example). So basically, a dummy head recording captures some real venue (or maybe a recording playing over speakers in a room), while a "3D audio" system renders a similar experience from "dry" audio sources by applying HRTF filters and room reverb IRs to them.

It's actually a useful insight because it means that we can use the same equalization of headphones for all of these: dummy head recordings, "realistic" spatializers, and my stereo spatializer. Essentially, I can decouple the tuning of my spatializer from the tuning of the headphones, and validate the latter using other spatial audio sources. And I am also able to compare these different sources directly, switching between them instantly, which provides ideal conditions for doing A/B or blind comparisons.

Hear-Through Transfer Function

The problem is, finding correct equalization for playing any of the spatial audio sources over headphones is actually a bit tricky. As I've mentioned above, first we need to "erase" the factory tuning of the headphones, and then re-calibrate them. But to what target? If we think about it, it must be a transfer function from an entrance of a blocked ear canal to the eardrum. This actually sounds like a familiar problem. For example, if we consider AirPods Pro, by their design they isolate the user from outside noises, however one can enable the "transparency" mode which adds the outer world back in. In the audio engineering industry this is also known as the "hear-through" function.

There is a good paper "Insert earphone calibration for hear-through options" by P. Hoffmann, F. Christensen, and D. Hammershøi on the transfer function of a "hear-through" device. The TF they use schematically looks like this:

The curve starts flat, then it begins to elevate at 300 Hz, reaching a peak between 2–3 kHz, then a sharper peak between 8–9 kHz. It also has a downward slope after 4 kHz. Note that in the paper there is a discrepancy between the curve shown on the Fig. 2 and the target curve for headphone tunings which appears on the Fig. 4.

In practice, we should consider this curve as an approximation only, since for each particular earphone and ear canal the center frequencies of the peaks and their amplitudes will be different. The previously published paper "Sound transmission to and within the human ear canal" by D. Hammershøi, and H. Møller shows that rather high individual variation exists between subjects when the TF from the outside of the ear canal to the ear drum is measured.

Also note that this TF applies to in-ear monitors (IEMs) only. The paper describes re-calibration of different IEMs to this target. Some of them comply well with re-tuning, some are not. For my experiment, I decided to use EarPods because they are comfortable—I have no physical ear fatigue whatsoever, they use a low distortion dynamic driver (see measurements on RTINGS.com), and they are cheap! Plus, you can put them into and out of your ear very easily, unlike true IEM designs that require some fiddling in order to achieve good sealing which blocks outside noises. Obviously, the lack of this sealing with EarPods is one of their drawbacks. After putting them in, you still hear everything outside, and what's worse, the timbre of outside sounds is changed, although since we are not using them for augmented reality purposes, this is not an issue. Yet another annoyance is that the 3.5 mm connector is rather flimsy, however, I have worked around that by wrapping it into heat-shrink tubing:

These drawbacks are not really serious—since we are not implementing actual hear-through for augmented reality, there is no concern about hearing modified sounds from the outside, and we will partially override them with music. However, here is the main drawback of EarPods: due to their shape, it is not easy to make a DIY matching coupler for them. Basically, you need a coupler imitating a human ear. Luckily, I now have access to a B&K head 5128-B. Thanks to its human-like design, EarPods can be easily mounted on it and I could obtain reliable measurements.

Yet another challenge comes from the fact that I wanted to create an individual hear-through calibration. That means, I wanted to be able to adjust the peaks of the hear-through target on the fly. These peaks actually depend on the resonances of the ear canal and the middle ear, which are all individual. The B&K head has its own resonances, too. Thus, even if I knew my individual HT target, calibrating to it on a B&K head might not give me what I want due to interaction between an EarPod and the ear canal of the B&K head.

So I came up with the following approach:

  1. Tune EarPods to "flat", as measured by the B&K head.
  2. Figure out (at least approximately) the TF of the B&K ear canal and microphone.
  3. Add the TF from item 2 back to the "flat" calibration TF.
  4. On top of the TF from item 3, add the generic "hear-through" (HT) TF, and adjust its peaks by ear.

Why does this approach work? Let's represent it symbolically. These are the TFs that we have:

  • B&K's own TF, that is, how the head's mic sees a "flat" low-impedance sound source at the ear canal entrance: P_bk;
  • EarPods, as measured by B&K: P_epbk;
  • Hear-Through TF—according to the paper it transforms a source inside a blocked ear canal into a source at the non-blocked ear canal entrance: P_ht.

We can note that P_epbk can be approximated by P_eb * P_bk, where P_eb is the TF that we could measure by coupling an EarPod directly to a microphone calibrated to a flat response. What we intend to do is to remove P_eb, and replace it with a TF that yields P_ht when EarPods are coupled with my ear canal. That is, our target TF is:

P_tg = P_ht / P_eb

When I calibrate EarPods to yield a flat response on the B&K head, I get an inverse of P_epbk, that is, 1 / P_epbk, and we can substitute 1 / (P_eb * P_bk) instead. Thus, to get P_tg we can combine the TFs as follows:

P_tg = P_ht * 1 / (P_eb * P_bk) * P_bk = P_ht * P_bk * 1 / P_ebbk

And that's exactly what I have described in my approach above. I must note that all these representations are of course approximate. Exact positions of transfer function peaks shift as we move the sound source back and forth in the ear canal, and move outside of it—one can see real measurements in the "Sound transmission to and within the human ear canal" paper I've mentioned previously. That's why it would be naive to expect that we can obtain an exact target TF just from measurements, however they are still indispensable as the starting point.

Approximating B&K 5128 Transfer Function

So, the only thing which has to be figured out is the P_bk transfer function. As I understand it, without a purposefully built earphone which is combined with a microphone inside the ear, this is not an easy task. The problem is that the SPL from the driver greatly depends on the acoustic load. For example, if we use a simplest coupler—a silicon tube between the earphone and a measurement mic, the measured SPL will depend on the amount of air between them, and the dependency is not trivial. What I mean, it's not just changing resonances and power decay due to distance increase.

I checked B&K's own technical report "The Impedance of Real and Artificial Ears" with lots of diagrams for impedance and transfer functions. I also tried some measurements on my own using low distortion dynamic driver IEM (the Truthear Zero:Red I mentioned in my previous posts). The main understanding that I have achieved is that the ear canal and microphone of the B&K head exhibit a tilt towards high frequencies. I was suspecting that because the EarPods tuned to "flat" on the B&K (the 1 / (P_eb * P_bk) TF) did sound very "dark." So I ended up with creating a similar tilt, and the actual amount was initially tuned based on listening, and then corrected when measuring the entire rig of EarPods via P_tg on the same B&K head.

Individual Hear-Through Tuning

As I've already mentioned before, the hear-through TF from the paper could only be considered as a guideline because of individuality of ear canals and the actual depth of insertion of earphones into them.

In order to figure out the individual resonances of the left and right ears I used pink noise rendered into one stereo channel and spatialized via my processing chain. I had two goals: get rid of the "hollow" sound (pipe resonance), and make sure that the pink noise sounds like a coherent point source. When the tuning is not right, it can sound as if consists of two components, with the high frequency band placed separately in the auditory space.

For this tuning process, I have created a prototype HT TF in a linear phase equalizer (DDMF LP10), and then was tweaking center frequencies of the peaks and their amplitudes. My final settings for the left and right ears look like this:

As you can see, compared to the "generic" TF pictured before, some changes had to be made. However, they are only needed to compensate for the features that got "erased" by the "flat" tuning. I had to restore back some bass, and also diminish the region around 6 kHz to avoid hearing sharp sibilants. This is how the final tuning looks like when EarPods are measured by the B&K head:

It looks rather similar to the target described in the "Insert earphone calibration..." paper. Note that the difference in the level of bass between the left and the right ear are due to my inability to achieve the same level sealing on both artificial ears— I must say, coupling EarPods is really non-trivial. Yet another difference is addition of dips around 60 and 70 Hz, as otherwise I could feel the vibration of the drivers in my ears.

Listening Impressions

As I have mentioned above, since we can consider dummy head recordings, "3D audio" renderings, and my spatializer as equally footed sources which require the same "hear-through" earphone calibration, I started with the first two.

I have made some dummy head recording using MS-TFB-2-MKII binaural mics by Sound Professionals mounted on my own head. I tried both just recording sounds of my home, and also recording pink noise played over stereo speakers (left only, right only, correlated PN, anti-correlated PN, and correlated PN from the back). For the "3D audio" rendering I used Anaglyph plugin. This is a free and simple to use plugin which allows rendering a mono audio source via different HRTFs, and apply different kinds of room reverbs.

I must say, I was impressed by the realism achieved! When playing sources recorded or rendered in front of me, I was sure the sound comes from the laptop sitting on my knees. Note that in Anaglyph I have set the source height at -15° for achieving this effect.

Playing sources recorded from the back was a bit of a surprise—I have experienced a strong back-front reversal with my own dummy head recording. However, the rendering by Anaglyph and my own spatialization chain produced much more realistic positioning. Thus, either Sound Professional's mics have some flaw which makes behind sources less "real," or "reality" is not what we want to achieve with these renderings— perhaps the brain requires stronger "clues" to avoid FB reversals, and these clues for some reason are not so strong in a dummy head recording.

But other than that, the dummy head recordings done with these mics are very convincing. If you play a record back exactly at the same location where it was recorded, the mind automatically connects sounds with the objects around. Or you will experience a bit confusing feeling if the visual scene has changed, say, a person which was talking on the recording has left, and now the reproduced sound comes from an empty place.

I also noted how much the feeling of distance and location "collapses" when a dummy head recording is played in another setting—another room for example. Our visual channel indeed has a strong dominance over the audio channel, and seeing a wall at the location where the sound is "coming from" in the recording completely changes the perception of the source's distance to the listener.

Also, the feeling of the distance is greatly affected by the loudness of playback, and perceived dynamic range. Thus, don't expect a lot of realism from simple studio-produced stereo recordings, even when rendering them via a "spatialization" engine.

Fine-Tuning the Spatializer

Using these realisting reference recordings and renderings I was able to fine tune my spatializer. Basically, I used pink noise positioned at different locations, and was quickly switching between my spatializer and a dummy head recording, or a rendering by Anaglyph. This was much more precise and easier than trying to compare the sound in headphones with sound in speakers! Even if you train yourself to put the headphones very quickly on and off, and switching the speakers off and on at the same time, this still does not happen quick enough. I've heard from James "JJ" Johnston that switching time must be no longer than 200 ms. With instant switching, I could compare both positioning of sources and their tonality. Note that these attributes are not independent because the auditory system basically "derives" this information from incomplete data which it receives from ears.

I noticed that the perceived distance to the sources rendered by my spatializer is less than the distance of "real" or "3D audio" sources. When switching between my spatializer and a DH recording, or Anaglyph, the source was literally jumping between "far" and "near" while staying at the same lateral angle and height. Increasing the level of the reverb in my spatializer puts the source a bit further away, at the cost of cleanness. Similar thing with Anaglyph: with reverb turned off the sound is perceived as being very close to the head. Yet, reverb adds significant tonal artifacts. Since I'm not doing a VR, I decided that it's better to have a cleaner sound closer to the face. In fact, a well produced recording has a lot of reverb on its own, which helps to "feel" the sources as if they are far away.

As a result of these comparisons, I have adjusted the level of the cross-feed, and adjusted directional bands the Mid/Side equalization for better rendering of "frontal" and "behind the back" sources. As I have mentioned above, the sources from behind ended up sounding more "realistic" with my rendering, compared to my own dummy head recordings.

Another interesting observation was that my spatialization produces more "heavier" sounding pink noise than a dummy head recording. I actually noticed the same thing with Apple's spatializer. I think this heavier tonality is better for spatializing music, and maybe movies as well, so I decided not to attempt to "fix" that.

Bonus

While researching this topic, I found a paper by Bob Schulein and Dan Mapes-Riordan with a very long title "The Design, Calibration and Validation of a Binaural Recording and Playback System for Headphone and Two-Speaker 3D-Audio Reproduction" on their approach to doing dummy head recordings. The paper also suggests a headphone equalization approach which is similar to what I ended up using. In fact, the "Transfer function T2" from the paper has a lot in common with the "hear-through" target function I used. From the paper, I came to Bob's YouTube channel which features some useful dummy head calibration recordings, as well as recordings of music performances, both captured by the mannequin called Dexter which Bob built.

Some notes on these recordings. I could hear that the tuning applied to the recording is different on older vs. newer videos—it is indeed hard to come up with the right tuning! Then, on the "calibration" recordings where Bob walks around Dexter, I was experiencing front-back reversals as well—that's when listening to recordings with my eyes closed. Thus, one either needs to always look at the visual—this is what Bob proposes, or incorporate head tracking into the reproduction process. However, using head tracking requires replacing in the recording process the dummy head with ambisonics mic, which captures the sound field in a more "generic" form, and then rendering this ambisonics capture via a "3D audio" renderer.

Another weak point of "old style" dummy head recordings that I have noticed is that minor creaking and coughing sounds coming from the audience immediately distract attention from the performance. This happens unconsciously, probably because you don't see the sources of these creaks, and this forces you to look into the direction of the sound as it can be potentially dangerous. When sitting in a concert hall, this unconscious switching of the attention happens much less because you have a wider field of view, compared to the video recording. Whereas, in stereo recordings the performance is captured using mics that are set close to the performers, so all sounds from the auditory can be suppressed if needed during the mixing stage.

Conclusions

I'm really glad I have solved the "last mile" problem of the stereo spatialization chain. Also, the headphone tuning I came up with allows listening to dummy head and "3D audio" recordings delivered with a great degree of realism. It's also interesting that it turns out, for music enjoyment we don't always need the reproduction to be "real." Much like a painting or a photograph, not mentioning movies, actually represents some "enhanced" rendering of reality and reveal the author's attitude to it. Sound engineers also work hard on every recording to make it sound "real" but at the same time interesting. Thus, the reproduction chain must not ignore this aspect.