Abstract
BACKGROUND:
The performance of a cochlear implant (CI), especially in conveying pitch depends on its electrical stimulation strategy.
OBJECTIVE:
The present study proposes a variable-rate stimulation algorithm which improves speech emotion perception by using temporal fine-structure cues and electrophysiological parameters of the patient.
METHODS:
This method is based on the coding of the phase information at the peak time intervals of the band-passed signals. The stimulation pulse is generated at the time of peak occurrence, which is able to excite the number of fibers with a discharge probability above a threshold. Calculating the discharge probability is based on the excitable fiber model and taking into account the biological characteristics of the patient, such as the fiber threshold and the distribution of remaining intact fibers.
RESULTS:
The results of the emotion detection test on selective reconstructed sentences from the Persian emotional speech database (Persian ESD) indicated that the listeners have been able to detect the emotion by an average of 83.82% using the proposed stimulation algorithm while it was 75% and 48.03% for the zero-crossing and the continuous interleaved sampling (CIS), respectively. Furthermore, the number of pulses compared to the zero-crossing and the CIS has decreased by 76.3% and 75.4%, respectively.
CONCLUSIONS:
In this paper, a stimulation method was proposed for cochlear implants by considering the patient’s biological parameters. It has been successful in transmitting speech emotion despite the reduction of stimulating pulses. This has some advantages such as reducing the interaction of current fields between electrodes during stimulation and reducing battery usage.
Keywords
Introduction
The cochlear implant (CI) is known as the best treatment to restore hearing to people with severe to profound deafness [1, 2]. According to the statistical results released by the end of 2012, 324,000 people have already enjoyed the benefits of this technique worldwide [3]. However, they have difficulties in music perception, speaker recognition and prosody perception [4, 5].
A CI bypasses the hair cells and stimulates the intact auditory nerves directly by electrical impulses [6]. Cochlear auditory fibers stimulation conveys the messages to the brain. The pattern of stimulation determines the type and importance of these messages. Finally, the messages processing by the brain forms the sound perception in the patient. The performance of a CI is greatly dependent on the ability of its speech processor to extract and transmit the speech information. Generally, the speech processors use two categories of speech features in electrical stimulation: coarse features and fine features. Among these methods, those using the coarse features of spectral envelope such as
Investigations have shown that CI users need information beyond signal envelope cues for pitch perception, due to the low number of stimulus channels. This information is contained in the phase component of speech or on the other hand in the temporal fine structure [15]. Therefore, in some of the novel methods, usually the stimulation rate is several thousand pulses per second or it is variable and the time interval between pulses is not necessarily equal.
Nie et al. [16] proposed a frequency-amplitude-modulation encoding strategy that transformed the fast-varying fine structure into a slowly-varying FM signal to modulate the center frequency in each band. In another method, they used the interleaved pulses in the peaks of the time envelope after the frequency shift of the signal in each channel [17, 18]. Chen and Zhang [19] applied peak occurrence times in temporal fine structure extracted by a Hilbert transform. In another study [20, 21] they used zero-crossing times to represent the fine structure information. Liu et al. [22] used one-octave wavelet transform zero-crossing stimulation to encode more phase information. Sit et al. [23] proposed a bio-inspired asynchronous interleaved sampling by using a race-to-spike algorithm to encode the phase information. Zhang et al. [24] used a model of the peripheral auditory system to generate stochastic impulses more similar to the auditory neural firing.
The aim of this paper is to introduce a stimulation algorithm based on the phase coding at the time interval between the peaks of signal. The proposed stimulation algorithm tries to avoid the use of pulses with insignificant effect on auditory fibers by considering the patient’s biological conditions and the nonlinear behavior of auditory neurons. The performance of the proposed algorithm in Farsi emotional sentences identification is compared with the CIS and zero-crossing methods.
Method
The block diagram of the proposed strategy is shown in Fig. 1. After pre-processing, the audio signal is decomposed into N channels using a bank of band pass filters. The center frequencies of these filters have been distributed based on the Mel scale. Then, the envelope of each channel is calculated by rectifying and low-pass filtering. The obtained signals are compressed in amplitude to be within the narrow electric dynamic range.
Block diagram of the proposed stimulation algorithm.
The extracted amplitude is used in the amplitude modulation process. A peak detection modulus is involved in a parallel path for determining the peak points in each channel. The output of this modulus will enter the effective stimulation recognizing (ESR) block. The input of the ESR block is accepted as a stimulus pulse if it can produce a discharge probability above a threshold. Then, the generated pulses in each channel are modulated by the extracted envelope and the resulting pulse trains are sent to the stimulating electrodes.
In order to encode the temporal cues of the audio signal the peak times are used. The auditory nerve fibers (ANFs) tend to spike at or near the peak times of the stimulus which is known as the phase-locking phenomenon. Phase-locking is weak at higher frequencies than 5 kHz. The input of the ESR block is considered as a stimulation pulse if it could provide an ANF discharge probability above a threshold.
The discharge probability is calculated using the discharge estimator algorithm which was introduced previously [25] based on the excitable nerve fiber model of Bruce [26]. This algorithm is applicable even when the interval between pulses is not fixed. The model of Bruce was used since in this model, the neuron’s response to two consecutive pulses can be predicted by considering the refractory effect. Also, the membrane noise fluctuations are involved. The threshold determining approach is outlined below.
According to the literature [27], each inner hair cell (IHC) is connected with several nerve fibers. Based on microscopic studies, the number of hair cells and the ANFs in each section of the cochlea could be estimated. It was assumed that if the input of the ESR block could successfully trigger at least one fiber within one millimeter neighborhood, it is accepted as a stimulation pulse. In order to reduce the calculating time, the following criterion was applied. When an ANF discharge probability exceeds
If two impulses occur at the same time in two channels, the priority for pulse generation is applied to the lower one.
Emotion identification is chosen to evaluate the performance of the proposed algorithm in conveying pitch frequency since the pitch contours are different in various emotional sentences. Synthesized sentences which acoustically simulate the hearing process of implanted patients are presented to normal hearing listeners. The performance of the proposed algorithm is compared with the CIS and zero-crossing methods, which represent fixed rate and variable rate methods, respectively.
Signal processing
First, the signal is passed through a pre-emphasis filter. Then it is decomposed by 12 Butterworth band-pass filters with central frequencies of 100, 290, 524, 814, 1173, 1617, 2166, 2845, 3684, 4723, 6008 and 7596 Hz. The envelope extraction and compressing processes are the same for the three mentioned algorithms.
Speech synthesis
The reconstruction method is based on Vocoder synthesis. In this method, the same series of band-pass filters are used for synthesis which have been used in signal decomposition. The pulse train of each channel is convolved with the impulse response of the band-pass filter in its channel. The synthesized signal is formed by summing the reconstructed signal from all channels.
where
A validated database of Persian emotional speech (Persian ESD) was used in the experiment [28]. This database contains a set of 90 validated novel Persian sentences classified into five basic emotional categories (anger, disgust, fear, happiness, and sadness), as well as a neutral category. The sentences are created using the simple Persian grammatical structure and were recorded in a professional recording studio in Berlin, Germany.
The sentences are spoken by two native speakers (one male and one female) in three categories. (1) congruent: the lexical content of the sentences is in accordance with the expressed emotions, (2) incongruent: neutral sentences are expressed in one of the mentioned emotions and (3) baseline: all sentences with different lexical contents are expressed in neutral.
Seventeen different emotional sentences from three categories were selected randomly for each stimulation algorithm. In this case, it is also possible to examine the effect of the concept of sentences in the emotion recognition. In order to have no response influence, the sentences were not the same for three algorithms. Also, all of the synthesized sentences were randomly located in the electronic questionnaire. The answer options were anger, sadness, happiness, neutral, none of them and lack of recognition. The questionnaire included three control synthesized sentences in a randomized order for detecting inconsistent answers. Subjects who marked different answers to more than one pair of control sentences were excluded from the study. Before questions, subjects listened to some original sentences to be familiar with different emotions. Participants were asked to listen to the voices in a quiet place and with a headset preferably. 44 normal-hearing Farsi speakers from 20 to 38 years old participated in this test.
Results
The percentages of response distribution were obtained for each sentence and then averaged for sentences with similar conditions in three existing categories. The results are shown in the Table 1. The correct recognitions are indicated in bold.
Emotion identification for proposed stimulation strategy
Emotion identification for proposed stimulation strategy
The phrase “none of them” is an emotion other than neutral, happy, angry and sad. For this section of the test, fear and disgust sentences from the database were used. The results reveal that for the congruent condition all emotions were recognized very accurately (100%). The poor results obtained in “none of them” can be attributed to the participants’ unfamiliarity with such sentences (47.05%). As expected, the obtained results in baseline and incongruent categories are weaker than for the congruent category. The analysis of incorrect answers shows that this could be due to the attention to lexical content of sentences in the baseline category and the absence of lexical cues in the incongruent section. The average of correct recognition for baseline and incongruent were 82.35% and 77.94%, respectively.
The correct rates of emotion identification in three categories for three stimulation algorithms.
In Fig. 2, the correct rates of emotion identification of the proposed stimulation algorithm in congruent, incongruent and baseline categories are compared with two stimulation algorithms: Zero-crossing (stimulation with non-uniform time intervals) and CIS (stimulation at a constant rate of 5,000 pulses per second). Mean values and standard deviations show that, generally, the proposed algorithm and zero-crossing performed better than the CIS. The p-values obtained from t-tests with the Bonferroni correction showed that the differences between each pair of the three strategies in the congruent and baseline categories were not significant (all
The scores of three algorithms in emotion detection.
Figure 3 shows the results obtained from the emotional sentences classification test. Three stimulation methods have been compared with t-tests. The obtained p-values were 3.9
Another important point in comparing these algorithms is the number of stimulation pulses. To investigate this factor, 17 identical sentences composed of 8 to 10 words, were considered for all three algorithms. The average number of pulses used for stimulation of each channel is shown in Fig. 4. The ratio of the average number of pulses used for the zero-crossing and CIS methods compared to the proposed methods is 4.23 and 4.06, respectively.
The average number of pulses used for stimulation for each channel.
The electrical current injected by the stimulated electrode changes the electrical potential in the volumetric conductor. Based on studies [29], the electrical potential around the stimulated electrode is calculated based on the distance of the stimulated electrode as well as the conductivity of the volumetric conductor as follows:
Thus the electrical potential drops around the electrode in bell-shape form [30]. Therefore, the use of band- pass filters in the vocoder method is common in audio signal simulation in many studies.
In order to compare the performance of the proposed stimulation algorithm with two other stimulation algorithms, some observations can be made. As can be seen in Fig. 2, the average of the correct answers in the incongruent category for both CIS and zero-crossing algorithms has decreased compared to the other two categories. This could be due to the fact that the correctness of responses in the congruent category may be affected by the lexical content of the sentences. Thus better results have been achieved in that section. The reduction in the results for the incongruent compared to the congruent category for the CIS and zero-crossing algorithms was 71% and 28.49%, respectively while it was 11.6% for the proposed algorithm. The differences in accuracy of the results in the congruent and incongruent sections are noticeable, especially for the CIS algorithm. In addition, the superiority of the baseline to the other two categories in the CIS algorithm suggests that this algorithm does not have the capability to convey the pitch frequency that is necessary to understand the speech emotion. Also, the results of t-tests with the Bonferroni correction showed a significant difference in the performance of the CIS algorithm compared to the proposed algorithm (
Pitch trajectories of one frame of an original sentence (a), the signal processed by: the proposed stimulation strategy (b), the zero-crossing strategy (c) and CIS (d).
In general, the proposed stimulation algorithm and the zero-crossing were more successful in conveying the speech emotion to the listeners. The results of the emotion detection test on selective reconstructed sentences from the Persian ESD database indicated that the listeners have been able to detect the emotion by an average of 83.82% and 75% using the proposed and the zero-crossing algorithms, respectively. While it was 48.03% for the CIS algorithm. In general, the proposed stimulation algorithm has performed better in transmitting the speaker’s feelings, or in other words, in conveying the pitch frequency Fig. 5a shows a part of pitch trajectories of an original sentence. Sections (b), (c) and (d) of this figure show the pitch trajectories of the vocoder-synthesized signal after processing by the proposed algorithm, zero-crossing and CIS, respectively. The extracted pitch trajectory of the processed signal by the proposed algorithmis significantly more consistent with the extracted pitch trajectory of the original sentence than the pitch trajectories obtained by the two other algorithms.
Reducing the number of pulses sent to the electrodes as can be seen in Fig. 4, is another advantage of the proposed stimulation algorithm. The number of pulses compared to the zero-crossing and the CIS has decreased by 76.3% and 75.4%, respectively. Reducing the number of pulses and consequently reducing injection flow to electrodes will cause less damage to neural tissues such as spiral ganglions and inner hair cells. Furthermore, it causes less electrode corrosion [31]. Also, sending fewer pulses to the electrodes has the benefits such as extending the life of the battery. Reducing the number of pulses with preservation of excitation quality can also reduce the interaction of current fields between electrodes during stimulation, which is one of the important issues in the performance of the cochlear implants.
In this paper, an algorithm was proposed for electrical stimulating in cochlear implants. This algorithm works based on coding of the phase information of the signal at the time interval between the peaks of signal. Also, the patient’s biological conditions such as the fiber threshold and the distribution of remaining intact fibers are considered in stimulation since the same electrical stimulation in patients with different electrophysiological conditions leads to different firing patterns of neurons. By considering the patient’s condition, at the peak occurrence times, pulses are sent to the electrodes if they are able to excite the fibers with a discharge probability above a threshold. Current reduction will eventually decrease possible neural damage in the long term. The emotion detection test performed by normal hearing listeners on reconstructed sentences from the Persian ESD database demonstrates that the proposed stimulation algorithm obtained better results than the zero-crossing and the CIS. Moreover, the number of pulses and thus the total current flowing through the electrodes were significantly lower than with the other two algorithms.
Using the peak occurrence times improves the transmission of temporal cues and applying a threshold criterion to choose the effective stimulations reduces the total number of electrical pulses.
Footnotes
Conflict of interest
None to report.
