

In December 2025, Fortell went to market with its breakthrough hearing aids and a strong conviction, backed by clinical evidence, that our technology was the best in the world at solving the toughest challenge in hearing science: hearing a single voice clearly in a noisy environment.
Now, we're pushing our AI further because we're committed to delivering the best technology to our wearers as it becomes possible. That curve moves quickly, and we're staying on top of it.
The latest improvement, Fortell AI 2.0, marks the biggest advancement yet to our AI model since launch. But we didn't want to just tell you it's better. We wanted to prove it.
Traditional hearing aids process sound as an undifferentiated stream: every frequency, every voice, every clatter of a plate gets treated more or less the same way. Fortell takes a different approach.
Instead of amplifying all sounds indiscriminately, our AI uses a custom on-device AI processor to figure out two things at once: what a sound is, and where it's coming from. It draws on subtle differences in timing, spectral content, and amplitude across multiple microphones to localize each voice in a room, then selectively amplifies the ones that matter while suppressing the ones that don't in real time.
Fortell AI 2.0 doesn't change that underlying architecture, it improves it with an expanded set of real-world training data.
Our team went out and recorded training audio in the environments where our patients struggle the most: restaurants, cafes, construction sites, trains, buses, cars, reverberant offices, and living rooms. Fortell AI 2.0 is the result: a model retrained on that expanded, real-world corpus, alongside updates to several underlying signal-processing algorithms.
The question we set out to answer with our research: does training on an expanded set of real-world audio meaningfully improve speech clarity over our original AI model?
We compared Fortell’s original AI model and Fortell AI 2.0 head-to-head on identical acoustic scenes, using a KEMAR acoustic manikin (a standard tool in hearing research, essentially an ear-shaped microphone rig) wearing Fortell AI Hearing Aids.
Eighty challenging scenes were recorded: 40 with background noise from real environments like cafes and train stations, and 40 with three competing talkers alongside the target voice.
Both AI releases were tested on the same physical hearing aids, with identical settings, in both of Fortell's listening modes:
We measured four established metrics of speech clarity: signal-to-noise ratio (SNR), HASPI v2, PESQ, and STOI.
Fortell AI 2.0 improved every metric, in every mode, in every scene type we tested. We completed 16 comparisons in total and all 16 improved with 15 reaching statistical significance.
Decibels are hard to picture in the abstract, so a useful way to think about the 4.1 dB improvement is to compare it to distance. You can think of it as standing about 35% closer to the person you're trying to hear in a loud room.
The results above are a summary. For the complete methodology, statistical breakdown, and full findings, our research team published the underlying data. Read the full paper below.

David Pattinson, Nathan Agmon, Israel Malkin, Mark Berry, Sam Russell, Dimitiri Roumeliotis, Igor Lovchinsky
August 2026
Fortell AI Hearing Aids can receive updates to their AI models and other digital signal processing algorithms. This paper presents an evaluation of the AI 2.0 release against Fortell’s debut AI model, retroactively dubbed “AI 1.0”. We recorded 80 challenging speech-in-noise scenes (half with background noise, half with interfering talkers) on a KEMAR acoustic manikin wearing Fortell AI Hearing Aids and measured objective speech metrics. AI 2.0 improved every metric in both of Fortell’s AI listening modes (“All Voices” and “Front Voices”) and both scene types.
Takeaways
Fortell AI Hearing Aids run neural-network-based speech enhancement algorithms on a dedicated, in-house, machine-learning coprocessor. AI processing separates target speech from interfering sound and re-renders the scene with the noise attenuated. All-day AI is available through two listening programs: “All Voices” attenuates background noise while keeping voices from all directions audible, whereas “Front Voices” additionally suppresses talkers not in front of the user. In a pre-registered double-blind randomized controlled trial, Front Voices improved speech intelligibility (SNR-50) in challenging multi-talker noise by 9.2 dB over the next best hearing aid employing AI source separation (Fortell Research, 2026). In a blinded preference study, Front Voices was preferred to all top-spec hearing aids from all five major manufacturers in 99 of 100 comparisons (Morris et al. 2026).
How well the AI processing performs is bounded by the audio it is trained on. Everyday listening is acoustically messy in ways that simulated or narrowly sampled training data may not capture: overlapping talkers at conversational distances, sources and listeners that move, hard reflective surfaces, and background noise that is neither stationary nor spectrally simple. We have found significant improvements to our AI models when they are trained on diverse, high quality real-world data.
AI 2.0 was re-trained on real world audio our team captured in environments that are typically difficult for hearing aid users, e.g. restaurants, cafes, construction sites, trains, buses, cars, and reverberant offices and living spaces. In AI 2.0, the neural network was re-trained on this expanded corpus and several signal processing algorithms were also improved. This paper reports a systematic technical evaluation comparing AI 2.0 with the previous AI 1.0 release on identical acoustic scenes.
All recordings were made on a Knowles Electronics Manikin for Acoustic Research (KEMAR; GRAS Sound & Vibration, Denmark) placed at the center of a ring of eight Genelec 8010 loudspeakers (Genelec, Finland) of 1 m radius. Fortell AI Hearing Aids were physically worn on KEMAR and set to the program being tested. Signals were captured by KEMAR’s ear-canal microphones at 44.1 kHz.
Baseline recordings of every scene were also made with open (unaided) KEMAR ears. In addition, a recording of each scene’s target speech alone was made in conventional (non-AI) mode for the reference signal of the PESQ and STOI metrics. Unlike HASPI v2, which was designed to compare a processed signal against an unamplified reference, PESQ and STOI were not designed for situations where only the test signal has passed through WDRC (and other processing), and may score spectral and level differences as distortion. Recording target speech alone in the conventional program therefore provides a spectrally matched reference for PESQ and STOI.
The same pair of Fortell AI Hearing Aids was used for both conditions, running either AI 1.0 or AI 2.0. Both releases were tested with identical noise reduction aggressiveness. The feedback management system was disabled to conduct the Hagerman Olofsson phase inversion method (details below, Hagerman & Olofsson 2004). The hearing aids were fitted using NAL-NL2 targets for an N3 audiogram (Bisgaard et al. 2010).
Two sets of 41 scenes were recorded; the first scene of each set was a warm-up and was excluded, leaving 40 scenes per set. The target talker was always presented from the front (0°). In the first set, the target talker was mixed with background noise from the seven loudest environments in the Ambisonics Recordings of Typical Environments (ARTE) database: Cafe 1, Cafe 2, Food Court 1, Food Court 2, Train Station, Street Balcony and Dinner Party (Weisser et al. 2019). In the second set, the target talker was mixed with three simultaneous competing talkers presented from −135°, +135° and 180° relative to the front. Target and competing talker speech was taken from the VoxCeleb2 corpus (Chung et al. 2018). VoxCeleb2 is an attractive test set because it contains thousands of real world conversations and incorporates a wide variety of accents, emotional states, and conversational dynamics. 56 unique speakers (24 female and 32 male) were represented in the recordings. None of the test data was used to train the AI models. Scene mixtures were generated at SII-weighted signal-to-noise ratios (SNR) drawn uniformly at random from −7.5 to −2.5 dB. This SNR range was selected to be challenging but not impossible for someone wearing conventional hearing aids with beamforming enabled to understand.
SNR was the primary metric used to evaluate the two releases. The SNR at the ear was measured with the phase-inversion method of Hagerman & Olofsson (2004). In this method, summing and differencing the speech-plus-noise and speech-minus-noise blocks separates the processed speech from the processed noise in the output of the hearing aid, and their broadband energy ratio, computed per ear and averaged, gives the output SNR.
Summing speech-plus-noise and inverted speech-plus-noise cancels all linear processing; the residual (expressed relative to the recovered speech and noise) is the reconstruction error and measures the validity of the Hagerman decomposition. The average reconstruction error across all scenes was −18.7 dB, and the maximum for any scene was −9.0 dB.
Each scene lasted 36 s: the first 18 s of the mixture allowed processing to stabilize, followed by three 6 s phase-inversion blocks (speech-plus-noise, speech-minus-noise, and inverted speech-plus-noise). The SNR benefit reported here is the output SNR minus the open-ear SNR of the same scene.
Three additional metrics were computed per scene on the recorded speech-plus-noise block:
Like SNR benefit, each metric is reported as the difference between the aided condition and the open ear recording of the same scene, averaged over both ears, and the 40 scenes of each set.
Comparisons between the two releases use two-tailed paired t-tests across the 40 scenes of each set, pairing the AI 2.0 and AI 1.0 values of the same scene (each scene contributed its average value across left and right ears). p-values are reported uncorrected for multiple comparisons. Wilcoxon signed-rank tests were run as a distribution-free check and agreed with the t-tests on every comparison (the same 15 of the 16 reached significance at the 0.05 level, and the same one did not).

Figure 1. Background-noise (ARTE) scenes. Difference to the open (unaided) ear for AI 1.0 and AI 2.0 in All Voices and Front Voices modes. Panels show output SNR benefit, HASPI v2, PESQ and STOI; each bar is the mean over 40 scenes, averaged across ears.

Figure 2. Competing-talker scenes. Difference to the open (unaided) ear, as in Fig. 1.
In the background-noise scenes (Fig. 1), AI 2.0 increases the output SNR benefit from 8.9 to 10.6 dB in All Voices and from 9.1 to 10.8 dB in Front Voices. This corresponds to a paired, per-scene improvement of 1.6 dB (95% CI ±0.6, p < 0.001) and 1.7 dB (±0.7, p < 0.001) respectively. The other metrics also improved in both programs: PESQ by 0.11 and 0.13 (both p < 0.001), HASPI v2 by 0.03 (p = 0.009 and p = 0.015), and STOI by 0.015 and 0.023 (p = 0.014 and p < 0.001).
In competing-talker scenes (Fig. 2), AI 2.0 Front Voices reaches 11.0 dB output SNR benefit (1.0 dB over AI 1.0, p < 0.001) and a PESQ improvement of 0.87 over the open ear (0.23 over AI 1.0, p < 0.001). The largest change in SNR benefit with the update is in All Voices, from 1.9 dB to 5.9 dB, a paired improvement of 4.1 dB (±0.3, p < 0.001). HASPI v2 improved by 0.13, PESQ by 0.31 and STOI by 0.05 in this condition (all p < 0.001).
Across the 16 combinations of program, scene type and metric, AI 2.0 improved the mean in all 16. The only comparison that did not reach statistical significance was STOI for Front Voices in competing-talker scenes (0.004, p = 0.73).
On identical acoustic scenes, AI 2.0 outperformed AI 1.0 on SNR, quality, and intelligibility metrics, in both All Voices and Front Voices and in both background-noise and competing-talker scenes. All 16 comparisons improved; 15 reached statistical significance.
With AI 2.0, output SNR benefit over the open ear reaches 10.6 to 10.8 dB in background noise, and 11.0 dB for Front Voices with competing talkers, and the largest single change–All Voices with competing talkers– rises from 1.9 to 5.9 dB. These gains come from an expanded corpus of real-world training audio and updates to several signal-processing algorithms, not from new hardware. Both releases were measured on the same pair of Fortell AI Hearing Aids and AI 2.0 is delivered as a free over-the-air update, so every existing wearer receives the improvements reported here on the devices already in their ears.
Bisgaard N, Vlaming MSMG, Dahlquist M. (2010). Standard audiograms for the IEC 60118-15 measurement procedure. Trends Amplif. 2010;14(2):113–120. doi:10.1177/1084713810379609
Chung JS, Nagrani A, Zisserman A. (2018). VoxCeleb2: Deep speaker recognition. Proc. Interspeech 2018, 1086–1090. doi:10.21437/Interspeech.2018-1929
Fortell Research. (2026). Spatial AI improves speech intelligibility for hearing aid wearers in challenging multi-talker noise. Fortell Research white paper, VOL. 1. https://www.fortell.com/evidence
Hagerman B, Olofsson Å. (2004). A method to measure the effect of noise reduction algorithms using simultaneous speech and noise. Acta Acust. united Acust. 2004;90(2):356–361.
Kates JM, Arehart KH. (2021). The Hearing-Aid Speech Perception Index (HASPI) version 2. Speech Commun. 2021;131:35–46. doi:10.1016/j.specom.2020.05.001
Morris C, Lovchinsky I, Callahan C, Wallace K, Berry M, Lyngaas K, Agmon N, Malkin I, Nagler B, Montano J, Weinstein B. (2026). Spatial AI consistently preferred to state-of-the-art hearing aids in multitalker noise. Int. J. Audiol. 2026:1–13. doi:10.1080/14992027.2026.2663345
Rix AW, Beerends JG, Hollier MP, Hekstra AP. (2001). Perceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs. Proc. IEEE ICASSP 2001, 749–752. doi:10.1109/ICASSP.2001.941023
Taal CH, Hendriks RC, Heusdens R, Jensen J. (2011). An algorithm for intelligibility prediction of time-frequency weighted noisy speech. IEEE Trans. Audio Speech Lang. Process. 2011;19(7):2125–2136. doi:10.1109/TASL.2011.2114881
Weisser A, Buchholz JM, Oreinos C, Badajoz-Davila J, Galloway J, Beechey T, Keidser G. (2019). The Ambisonic Recordings of Typical Environments (ARTE) database. Acta Acust. united Acust. 2019;105(4):695–713. doi:10.3813/AAA.919349

