Abstract
Accurate pitch perception in cochlear implants (CIs) remains challenging, making psychophysical measures susceptible to procedural biases. In unilateral pitch ranking, subjects judge which of two successive stimuli is higher in pitch, typically yielding electrode order roughly consistent with tonotopy. In our study with 10 MED-EL CI users, each of the 12 electrodes was compared with its three next neighbors. Pitch ranking was influenced by the presentation order of the electrode pairs. To investigate this effect, we tested 9 normal-hearing (NH) listeners using bandpass-filtered noise (6 dB/oct slopes) to approximate current spread during CI stimulation. Comparable performances across CI and NH in all three spacing conditions indicated similar pitch perception difficulty. On average, both groups discriminated 8 of 12 channels based solely on pitch (discriminability via d’), suggesting that 6 dB/oct slopes approximate mean monopolar current spread at the group level. Pitch rankings of adjacent channels exhibited biases. A repeated-measures ANOVA revealed a significant interaction between presentation order and stimulation site in both groups (CI: p = .003; NH: p < .001): pairs below the stimulus range midpoint were more often reported as descending, whereas pairs above the middle were more often reported as ascending. This asymmetric bias may reflect a context-dependent decision process in which subjects compare the second stimulus to a weighted average of the first and previously presented stimuli. Presentation-order effects were mitigated by testing stimulus pairs in isolation, at the expense of reduced randomization. Our findings highlight the importance of controlling biases to ensure reliable psychophysical measures.
Keywords
Introduction
Cochlear implants (CIs) are the most successful neural prosthesis and the first capable of substituting a sensory organ. Hearing sensations are elicited via an electrode array implanted in the cochlea, which bypasses the damaged, non-functional hair cells in the inner ear and directly electrically stimulates the surviving auditory neurons. Currently, most CI users understand speech using the device alone and achieve high speech recognition scores in quiet environments. This high performance in speech perception leads to the desire for equally good levels of recognition in noisy environments and perception of non-speech sounds, especially music.
Improving pitch perception may potentially improve speech in noise and music perception (McDermott, 2004). Pitch, the perceptual dimension that allows a listener to rank a sound from “low” to “high” or from “dull” to “sharp” (McDermott & McKay, 1994; Schatzer et al., 2014; Zeng, 2002), is essential when analyzing auditory information. Changes in the fundamental frequency (F0) during an utterance, perceived as changes in pitch (McDermott, 2004; Pieper & Bahmer, 2019), constitute a prosodic cue that transmits emotion in speech and music (Gilbers et al., 2015; Jiam et al., 2017; Mitchell & Kingston, 2014; Moore, 2024; Picou et al., 2018) and convey whether a speaker is asking a question or making a statement (Jiam et al., 2017; Lieberman, 1966; Majewski & Blasdell, 1969). In tonal languages such as Mandarin and Vietnamese, subtle pitch changes convey important semantic structures and can change the meaning of words. In non-tonal languages, pitch can be used to indicate stress, which can also change lexical meaning (e.g., desert vs. dessert). Moreover, pitch cues help differentiate between sound streams, e.g., from different speakers, and (rate) pitch discrimination has been shown to correlate significantly with the ability to perceive speech in background noise (Pieper & Bahmer, 2019; Zhou et al., 2019). Beyond speech, pitch is a central component of music: melodies and harmonies consist of sequential pitch patterns. Due to CI users’ generally poor pitch perception, their performance in melody identification tasks is low, especially when verbal and rhythmic cues are unavailable (Gfeller et al., 2007; Kong et al., 2004; Looi et al., 2004; for a review, seeMcDermott, 2004).
In multi-channel CIs, pitch can be conveyed either by changing the stimulation place, stimulation rate, or a combination of both. When the stimulation place is fixed and the stimulation rate is varied, listeners perceive a change in pitch known as rate pitch (Landsberger et al., 2016; McDermott, 2004; Zeng, 2002). For most CI users, the perceived pitch increases with increasing stimulation rate up to around 300 pps (McKay et al., 2000; Shannon, 1983; Townshend et al., 1987), although some listeners can perceive pitch changes at higher rates (Kong & Carlyon, 2010). Interestingly, this limit is in the range of the saturated maximum firing rate of a single nerve fiber (Liberman, 1978; Taberner & Liberman, 2005); however, a direct relationship has not been established. Place pitch refers to the changes in pitch that occur with changes in the place of stimulation. Due to the tonotopic organization of the inner ear, stimulating different electrodes along the implanted array results in different pitch percepts (Donaldson et al., 2005; Townshend et al., 1987; Zeng, 2002). For conventional single-electrode stimulation, the number of discriminable place pitches is constrained by the number of available electrodes (12–22 in contemporary CIs), though not all electrodes necessarily provide distinct pitch percepts due to overlapping neural excitation patterns (Macherey & Carlyon, 2014).
Place pitch is commonly assessed using pitch ranking tasks, which employ a two-interval, two-alternative forced-choice (2I2AFC) procedure. Two electrodes are stimulated sequentially, and the subject indicates which electrode evokes the higher pitch percept. Unilateral pitch ranking of CI electrodes generally yields a tonotopic ordering, with more basal electrodes being judged as higher in pitch (Nelson et al., 1995; Townshend et al., 1987). Electrode pairs are compared multiple times, ideally in both apical-to-basal and basal-to-apical orders, and performance is often pooled across presentation orders for further analysis (Collins et al., 1997; Nelson et al., 1995; Reiss et al., 2011; Townshend et al., 1987). However, Huber et al. (2018) observed ambiguous responses depending on the presentation order of the electrode pair, a procedural bias known as “time-order error” first described by Fechner (1860) and extensively reviewed by Hellström (1985). Such time-order errors have also been reported in CI studies on frequency interval ranking thresholds (Luo et al., 2014) and melodic contour identification (Galvin et al., 2007). Alternative ranking procedures have also been proposed, such as the midpoint comparison method (Long et al., 2005), in which electrode comparisons are adaptively selected and typically measured only once. Nevertheless, repeated pairwise comparisons remain widely used in CI pitch-ranking studies.
Pitch ranking procedures can also be used to measure interaural pitch matches, either to a contralateral electrode or to an acoustic stimulus in the non-implanted ear. Goupell et al. (2019) and Jensen et al. (2021) reported procedural biases influencing such pitch ranking tasks in bilateral (BI) CI users: matches for E12Left shifted on average by three electrodes when the testing range spanned only the basal half (E2Right – E12Right) compared to the full array (E2Right – E22Right). This so-called “testing range bias” has also been observed in single-sided-deafness (SSD) CI users (Carlyon et al., 2010; Schatzer et al., 2014) and NH listeners (Carlyon et al., 2010). In addition to range biases, Reiss et al. (2011) and Reiss et al. (2012) reported that previous comparison pairs affected the judgment of subsequent stimulus pairs: in SSD CI users, pitch matches for a reference electrode were ambiguous depending on whether the acoustic comparison frequencies were presented in descending (2 kHz → 125 Hz) or ascending (125 Hz → 2 kHz) order (Reiss et al., 2011). Such a sequential effect was also observed in BI CI users, where pitch ranking responses differed depending on whether the contralateral electrodes were presented from base to apex or vice versa (Reiss et al., 2012). Among sequential effects, starting-point bias has also been observed: the first comparison stimulus in a test session can influence subsequent judgments, leading to different pitch matches depending on the initial stimulus (Adel et al., 2019; Carlyon et al., 2010; Goupell et al., 2019).
In this study, we address the bias in unilateral pitch ranking introduced by the presentation order of the stimulus pair, also known as time-order error. While time-order errors have been observed and occasionally noted in CI studies (Galvin et al., 2007; Huber et al., 2018; Luo et al., 2014), to our knowledge, they have not been systematically investigated, despite their potential to bias pitch ranking outcomes. Alongside other procedural biases (testing-range bias, sequential effects, starting-point bias), time-order errors may further distort pitch ranking data and compromise derived measures like the perceptual sensitivity (e.g., cumulative d’, see Collins et al., 1997; Nelson et al., 1995). Here, we quantify the magnitude of time-order errors, assess their implications, and explore strategies to mitigate them, with the goal of strengthening the methodological reliability of pitch ranking. To this end, we examined unilateral pitch ranking in CI users under three electrode-spacing conditions (adjacent, two-apart, and three-apart). In addition, we tested NH listeners with bandpass-filtered white Gaussian noise, which approximates the broad current spread of CI electrodes, to examine whether presentation order effects also occur in acoustic hearing. This approach enabled a dedicated assessment of time-order errors in pitch perception and their implications for both CI users and NH listeners.
Methods
Pitch ranking experiments were conducted on CI users and NH listeners. Pitch, as well as timbre, are likely to contribute to the pitch-like perception (Schatzer et al., 2014). The term “pitch” in this study was defined in the broad sense, describing the perceptual dimension used by the subjects to rank a stimulus from “low” to “high” and may encompass cues relates to pitch in the musical sense as well as timbre. Part of the CI dataset has been reported previously in conference proceedings (Huber et al., 2018).
Pitch Ranking in CI Users
Subjects
Ten CI users (CI1–CI10; age range: 22–55 years, mean age: 44 years) participated in the study. All users were implanted with 12-channel electrode arrays from MED-EL (Pulsar, Sonata, Concerto) and had at least two years of CI experience. A total of 13 ears were measured. Demographic information is listed in Table 1. All CI users provided written informed consent and received monetary compensation for their participation. Measurements for CI users and NH listeners were conducted in accordance with the Declaration of Helsinki and approved by the medical ethics committee of the Klinikum rechts der Isar (Munich, 2126/08).
Cochlear Implant (CI) Users’ Demographic Information at Time of Testing. M: Male, F: Female, L: Left, R: Right, SNHL: Sensorineural Hearing Loss.
Stimuli
The speech processor of the CI was bypassed by directly stimulating the electrode array with the Research-Interface-Box RIB2 (Institute of Ion Physics and Applied Physics, University of Innsbruck). The individual experimental designs were programmed using MATLAB (Mathworks). Stimuli were 300-ms pulse trains consisting of biphasic monopolar pulses (40-μs phase duration), with no interphase gap, and a leading cathodic phase. All stimuli were presented at a fixed pulse rate of 1250 pps at 70% dynamic range (DR). For the pitch ranking, subjects compared the pitch of two successive stimuli, each at a different electrode location. The interstimulus interval between the two stimuli was 500 ms.
Experimental Procedure
Thresholds and Maximum Acceptable Levels
First, thresholds (THR) and maximum acceptable levels (MAL) were determined via the method of adjustment for all electrodes. CI users could change the loudness of the stimuli by adjusting the stimulus amplitude in coarse and fine steps using a keyboard marked with thick and thin up and down arrows. CI users adjusted the THR such that the stimulus was as soft as possible but still reliably audible. For the MAL, subjects were asked to set the stimulus as loud as possible without it being uncomfortable or painful. THR and MAL were adjusted 3 times each. In later measurements, the adjustable range for the stimulus amplitude always remained between THR and MAL.
Loudness Balancing
To minimize the effects of level on pitch, the loudness of all electrodes was equalized using the adjustment method. CI users adjusted the loudness of one stimulus (the target stimulus) to match that of the other (the reference stimulus). The reference stimulus was presented once, followed by the target stimulus presented twice. This sequence was repeated until the equal loudness criterion was satisfied. The adjustment of the target stimulus amplitude was possible in coarse and fine steps. Each target stimulus was equalized 4 times to the reference stimulus at E6 with a level of 70% DR.
Pitch Ranking
The main experiment was a unilateral place pitch ranking using a two-interval, two-alternative forced-choice (2I2AFC) procedure. Subjects were presented with two sequential stimuli and indicated on a keyboard which one, the first or the second, was higher in pitch. No feedback was given. All 12 electrodes (E1–E12, numbered from apex to base) were used to form comparison pairs. For each electrode Ei (i = 1–11), pairs were formed with electrodes Ei + 1, Ei + 2, and Ei + 3 when available, yielding three distance conditions: 1-electrode distance (11 adjacent pairs), 2-electrode distance (10 pairs), and 3-electrode distance (9 pairs). Each electrode pair was presented 5 times in an apical-to-basal order (ascending pitch pattern according to the cochlea's tonotopy) and 5 times in a basal-to-apical order (descending pitch pattern). For simplicity purposes, the two presentation orders will be referred to as ascending and descending conditions, respectively. Stimulus pairs were randomized over all trials and combinations. Subjects could repeat (replay) the stimulus pair as many times as needed before making a decision. The number of additional repetitions beyond the initial stimulus presentation will be referred to as “replays”. At the beginning, the subjects had one trial run with 9 randomized 3-electrode distances to get used to the procedure; no feedback was provided.
Pitch Ranking in NH Listeners
Subjects
Nine NH listeners (NH1–NH9; age range: 23–57 years, mean age: 29 years) participated in this study. Absolute hearing thresholds in quiet were measured in a soundproof cabin (IAC 350; IAC Acoustics, Winchester, UK) at the standard audiometric frequencies between 125 Hz and 8 kHz with the SENTI DESKTOP FLEX (PATH MEDICAL) device. The better of the two ears was chosen for subsequent measurements. If both ears had equally good hearing thresholds, the subject could pick the side for later stimulation. All but one participant showed thresholds in quiet of ≤ 20 dB HL in the selected ear. NH4 presented a mild hearing loss at 8 kHz (<30 dB HL). All NH participants provided written informed consent for their participation.
Stimuli
We used bandpass-filtered white Gaussian noise to simulate the monopolar-intracochlear stimulation of CI electrodes (Adel et al., 2019; Bingabr et al., 2008; Oxenham & Kreft, 2014). Bandpass noise was generated by applying first- and tenth-order Butterworth bandpass filters to white Gaussian noise (Horbach et al., 2018; Zwicker & Fastl, 2013). A filter bank comprising 12 first-order filters (Figure 1) was designed to simulate the 12-electrode array implanted in all CI participants. Per participant, white Gaussian noise was generated once prior to filtering. Stimuli were 300 ms long, including 30-ms raised-cosine ramps at signal onset and offset. During pitch ranking, an interstimulus interval of 500 ms separated the stimuli to be compared, analogously to the CI study. All stimuli were generated digitally at a sampling rate of 192 kHz using MATLAB and were presented monoaurally through an external audio interface (Fireface UC RME, Germany) and headphones (HD 650, Sennheiser electronic GmbH, Wedemark). The headphones were calibrated to 60 dB SPL at 1 kHz. NH subjects conducted all experiments in the soundproof cabin to attenuate ambient noise and reduce visual influences.

Filterbank for the shallow condition to approximate the current spread in the cochlea produced by monopolar stimulation. Center frequencies specified in Table 2, Q10dB = 0.17, and 6 dB/oct slopes on both sides of the spectrum.
Center Frequencies (in Hz) for Each of the 12 Filters for the Approximation of the 12-Electrode Array.
Center Frequencies
Based on acoustic pitch matches from SSD CI users (Baumann & Nobbe, 2006; Boëx et al., 2006; Dorman et al., 2007; Peters et al., 2016; Schatzer et al., 2014; Vermeire et al., 2008; Vermeire et al., 2015), 12 center frequencies (fc) were spaced logarithmically between 241 and 5334 Hz (Table 2).
Bandwidth
The filters used to simulate the CI array had a fixed relative bandwidth (Q10dB) of 0.17, which lies within the range of relative bandwidths found with forward-masked psychophysical spatial tuning curves (fmSTCs) in CI users (Nelson et al., 2008; Nelson et al., 2011). We verified the selected bandwidth to exceed the critical bandwidth of each of the corresponding center frequencies (Table 2), resulting in stimuli with a consistent stimulus pitch strength of (less than) 20% (Horbach et al., 2018; Zwicker & Fastl, 2013).
Slopes
The pitch strength of bandpass-filtered noise also depends on the steepness of the filter slopes, decreasing as the slopes become shallower (Zwicker & Fastl, 2013). Previous work has shown that steeper spectral slopes are required to elicit simple and complex pitch percepts in vocoder simulations (Mehta & Oxenham, 2017). To examine the role of pitch salience, we implemented two filterbanks comprising first- and tenth-order Butterworth bandpass filters. The first-order Butterworth bandpass filters (Figure 1), with a slope of 6 dB/oct on both sides of the spectrum, were chosen to approximate the current spread in the cochlea produced by monopolar stimulation. Considering a mean current decay rate for monopolar stimulation of αc = 1.27 dB/mm (Nelson et al., 2008; Nelson et al., 2011) and using Greenwood's equation (Greenwood, 1961, 1990) to relate the distance along the basilar membrane to characteristic frequencies in an adult human cochlea, the current decay rate αc can be approximated by a 5.95 dB/oct slope. It should be noted that for the conversion from current decay rate to filter slope we did not adjust for the different dynamic ranges between acoustic and electric stimulation. Adjusting for the larger acoustic dynamic range would result in steeper filter slopes; however, the 6 dB/oct was purposely chosen to reflect the broad current spread at the level of the auditory nerve fibers and the resulting spectral overlap associated with monopolar stimulation (Nelson et al., 2008). The stimuli generated with first-order Butterworth bandpass filters (6 dB/oct slope, Figure 1) will henceforth be addressed as the shallow condition (F1shallow–F12shallow).
Filters with steeper slopes and a Q10dB of 1.34 were selected for a second filterbank with otherwise identical parameters as the shallow filterbank. Stimuli generated with tenth-order filters (60 dB/oct slope), from now on referred to as the steep condition (F1steep–F12steep), served as a control measurement to assess the reliability of the NH subjects’ pitch judgments. The aggregate pitch ranking performance (grand average, nNH = 9) for the steep condition is shown in Figure 2. The results suggest that ranking F1steep–F12steep was an easy task, reflected in the more than 90% correct responses at all three measured distance conditions (1 filter: 91.41 ± 9.02%; 2 filters: 92.19 ± 8.43%; 3 filters: 92.84 ± 9.25%).

Aggregate pitch ranking performance for the steep filter condition in normal-hearing (NH) listeners. The performance for each distance condition (1-, 2-, and 3-filter distance) was computed by averaging over both presentation orders (ascending and descending) and across all filter pairs. Error bars represent one standard deviation from the mean. nNH = 9.
Experimental Procedure
Before proceeding with the experiment, subjects received written instructions for the corresponding task. After reading the instructions, subjects could ask questions that could be answered strictly with “yes” or “no” to ensure standardized information across participants. No further instructions or advice were given.
Loudness Balancing
All stimuli were scaled to have the same root mean square (RMS) value as a 1 kHz sine tone at 60 dB SPL. To further minimize the effect of loudness on pitch, the loudness of the lowest frequency stimulus F1 and the highest frequency stimulus F12 were equalized to the mid-frequency stimulus F6 for the steep and shallow conditions, respectively. The procedure was analogous to the loudness balancing in the CI study. Each target stimulus (F1 and F12, respectively) was equalized 4 times to the reference stimulus (F6). Linear interpolation in the log-domain (dB SPL) was used to balance the rest of the stimuli in between.
Pitch Ranking
Analogous to the CI study, NH subjects were presented with a stimulus pair and indicated on a wireless numeric keypad whether the first or the second stimulus was higher in pitch. All 12 stimuli (F1–F12, numbered from apex to base) were used to form comparison pairs within each filter condition (shallow and steep). For each stimulus Fi (i = 1–11), pairs were formed with Fi + 1, Fi + 2, and Fi + 3 when available, yielding three distance conditions: 1-filter distance (11 adjacent pairs), 2-filter distance (10 pairs), and 3-filter distance (9 pairs). Each stimulus pair was compared 4 times in ascending (stimulus with the lowest center frequency first) and 4 times in descending (stimulus with the highest center frequency first) order. Stimulus pairs were presented in randomized order. Randomization did not occur between shallow and steep or between different distance conditions. Subjects could replay the stimulus pair as many times as needed before reaching a decision. As in the CI group, the number of additional presentations beyond the initial playback will be referred to as “replays”. Subjects familiarized themselves with the procedure in a single training session (steep, 3-filter distance). No feedback was given during training or main sessions.
Statistical Analysis
From the 13 CI ears, 12 were used for statistical analysis due to missing data from E10, E11, and E12 from subject CI5. All 9 NH subjects were included in the statistical analyses.
Aggregate Pitch Ranking Performance
A binomial generalized linear mixed-effects model (GLMM) with a logit link was fitted using maximum likelihood in R (lme4 package) to assess overall performance differences between CI and NH groups pooled across presentation order and location. The dependent variable was the number of “correct” responses out of the total number of trials. The model had fixed factors of group (CI vs. NH), distance condition (1-, 2-, 3-electrode/filter), and their interaction, with a random intercept for subject to account for inter-subjective variability.
Pitch Ranking of Adjacent and Non-Adjacent Electrodes and Filters
To examine presentation order effects, repeated-measures analyses of variance (ANOVA) were conducted in IBM SPSS Statistics separately for each distance condition (1-, 2-, and 3-electrode/filter distance) and for each group (CI and NH). For the 1-electrode/filter distance, a two-factor repeated-measures ANOVA was performed to test the main effects of presentation order (ascending, descending), and electrode/filter location (E1–E12/F1–F12), as well as their interaction at a 95% confidence interval. Presentation order, electrode/filter location, and their interaction were within-subject variables tested with the Greenhouse-Geisser correction. Post-hoc Bonferroni pairwise comparisons were conducted when significant effects were observed. For the 2- and 3-electrode/filter distance conditions, analogous two-factor repeated-measures ANOVAs were computed.
Data Visualization
Unless stated otherwise, the plots represent the mean over 12 CI ears and 9 NH subjects. Reported error bars show one standard deviation from the mean. The data was acquired, analyzed, and plotted in MATLAB.
Reporting Guidelines
We used the STROBE reporting checklist when editing (Von Elm et al., 2007).
Results
Unless explicitly specified otherwise, the NH results in this section refer to the shallow (F1shallow–F12shallow, 6 dB/oct slopes) condition. For simplicity purposes, when describing the results of the CI group, “performance” refers to the percentage of responses that correlated with the cochlea's tonotopic organization, i.e., the basal electrode being judged higher in pitch would be “correct”. The performance reflects the number of “correct” responses out of 5 measurements (4 in the NH case) for a particular electrode pair and presentation order, i.e., ascending and descending.
Aggregate Pitch Ranking Performance
Figure 3 shows the aggregate pitch ranking performance for the three distance conditions (1, 2, and 3 electrodes/filters) for CI users (black) and NH listeners (gray). The aggregate performance for each distance condition was calculated by pooling over both presentation orders (ascending and descending) and across all electrode/filter pairs. CI users and NH listeners exhibited highly similar performances at all three distances, improving as the distance between the stimuli increased: 73.69 ± 10.15% (CI) vs. 74.07 ± 13.50% (NH) at 1 electrode/filter distance, 85.77 ± 8.75% (CI) vs. 87.58 ± 13.48% (NH) at 2 electrodes/filters distance, and 90.73 ± 8.57% (CI) vs. 93.45 ± 7.91% (NH) at 3 electrodes/filters distance.

Aggregate pitch ranking performance for cochlear implant (CI) users and normal-hearing (NH) listeners for the shallow filter condition. The performance for each distance condition (1-, 2-, and 3-electrode/filter distance) was computed by averaging over both presentation orders (ascending and descending) and across all electrode/filter pairs. Error bars represent one standard deviation from the mean. nCI = 12. nNH = 9.
The binomial GLMM revealed a significant main effect of distance, with a strong positive linear trend (β = 0.94, SE = 0.09, z = 11.04, p < .001), corresponding to an odds ratio of 2.57. The quadratic component was not significant (β = −0.13, SE = 0.08, z = −1.52, p = .13), suggesting a linear increase in performance with greater electrode/filter distance. There was no significant effect of group (β = 0.15, SE = 0.09, z = 1.61, p = .11; odds ratio = 1.16) and no group × distance interaction (linear component: β = 0.26, SE = 0.15, z = 1.70, p = .09; quadratic component: β = 0.04, SE = 0.15, z = 0.30, p = .77), suggesting comparable (distance-dependent) performances for CI users and NH listeners.
The random intercept variance for subject was 0.41 (SD = 0.64), reflecting substantial inter-subject variability. Both groups exhibited large variability across individuals. For example, within the CI group, CI5 scored over 85% in all three distance conditions and even performed perfectly at 3 electrodes distance, while CI7 only reached 80% performance at the widest (3-electrode) distance condition. Nevertheless, all CI users performed better as the distance between the compared electrodes increased. No systematic trends in performance were observed related to electrode spacing or array length. Notably, some of the highest and lowest scores were found among users with the most closely spaced electrodes (FLEX24).
Pitch Ranking of Adjacent Electrodes and Filters
The pitch ranking results in Figure 4 show the performance of individual electrode/filter pairs for CI users and NH listeners. The compared electrode (filter) pairs are indicated on the abscissa, while the percentage of responses that correlate with the cochlea's tonotopy (correct responses) is on the ordinate. Electrode/filter pair labels in the figure correspond to the ascending order. The light blue/green bars represent the descending condition (more basally located electrode/filter stimulated first), and the dark blue/green bars the ascending condition (more apically located electrode/filter stimulated first). A two-factor repeated-measures ANOVA (presentation order × electrode/filter location) was conducted separately for CI users and NH listeners at the 1-electrode/filter distance. Neither the presentation order (CI: F(1, 11) = 0.67, p = .43, η2 = .06; NH: F(1, 8) = 0.84, p = .39, η2 = .10) nor the electrode/filter location (CI: F(4.53, 49.78) = 2.21, p = .07, η2 = .17; NH: F(3.68, 29.42) = 1.53, p = .22, η2 = .16) significantly affected the pitch ranking measurements of CI users or NH listeners. However, there was a significant interaction between presentation order and stimulation place along the cochlea in both groups (CI: F(5.09, 55.95) = 4.16, p = .003, η2 = .27; NH: F(4.07, 32.59) = 8.04, p < .001, η2 = .50). While at apical locations, the descending pitch pattern yielded higher number of correct responses, more basally located electrodes/filters performed better with the ascending presentation order.

Pitch ranking performance by electrode/filter pair for adjacent electrodes/filters for
Post-hoc tests revealed the electrode pair E1/E2 at the apex performed significantly better when presented in descending order than when judged in ascending order (E2/E1: 73.85 ± 26.31% vs. E1/E2: 50.77 ± 34.27%, p = .03). On the contrary, the ascending presentation of E7/E8 performed significantly better than the descending order (E7/E8: 89.23 ± 13.20% vs. E8/E7: 67.69 ± 30.04%, p = .02). This interaction was even stronger in NH listeners. In the apical region, the descending order achieved significantly higher performances: F1/F2 (F2/F1: 94.44 ± 7.77% vs. F1/F2: 41.67 ± 37.5%, p = .002), as well as F2/F3, F3/F4, and F5/F6, which dropped by 41.67% (p = .04), 38.89% (p = .01), and 22.22% (p = .02) when presented in ascending order. Conversely, in the basal region, ascending pitch patterns performed significantly better: F9/F10 (F9/F10: 91.67 ± 12.50% vs. F10/F9: 61.11 ± 35.60%, p = .04) and F11/F12 (F11/F12: 94.44 ± 16.67% vs. F12/F11: 58.33 ± 30.62%, p = .03).
Individual pitch ranking results are shown in Figure 5. CI8 (Figure 5

Pitch ranking performance by electrode pair for adjacent electrodes for
Pitch Ranking of Non-Adjacent Electrodes and Filters
For CI and NH pitch ranking at the 2- and 3-electrode/filter distances, no significant effects of presentation order, electrode/filter location, or the interaction between both factors were observed (see Table 3). The pitch ranking performances at 2- and 3-electrode/filter distance showed high performance scores and good agreement between the descending and ascending presentation orders. At the 2-electrode distance, the mean maximal deviation between the two presentation orders was around 15% (E7/E9, E9/E11). Increasing the distance to 3 electrodes reduced the mean maximal difference between the descending and ascending conditions to around 9% (E2/E5, E4/E7). In NH, pitch judgments at 2- and 3-filter distances did not exhibit significant differences between ascending and descending presentation orders. The largest deviation between presentation orders was 19%, observed for F1/F3 and F7/F10 at 2- and 3-filter distances, respectively.
Two-Factor Repeated-Measures ANOVA Results for the 2- and 3-Electrode/Filter Distance Conditions for the Cochlear Implant (CI) Users and Normal-Hearing (NH) Listeners.
Channel Discriminability
To estimate the total number of discriminable steps across the stimulus range, we converted the percent correct scores into a sensitivity index
The number of discriminable channels for the descending and the ascending conditions for each of the 12 CI users and 9 NH listeners is shown in Figure 6. Mean values are shown in filled red symbols. Two open symbols connected by a line represent a subject, showing that the number of discriminable channels differed depending on the presentation order, sometimes by as much as 5 electrodes (nCI = 2) and up to 6 filters (nNH = 1). Concentric circles (CI) and squares (NH) show that more than one subject was able to discriminate, e.g., 9 out of 12 electrodes in the descending condition (nCI = 5). The average difference was 1.83 ± 1.47 electrodes and 2.40 ± 2.08 filters, with slightly more channels being discriminable in the ascending condition. However, the effect of presentation order on the number of discriminable channels was not significant for CI users or NH listeners. On average, CI users could discriminate 8.17 ± 1.99 out of 12 electrodes. The lowest number of discriminable electrodes was 5 (descending, nCI = 2); the best performers could discriminate up to 11 electrodes (nCI = 4). NH listeners showed higher intersubjective variability, being able to discriminate between 3 (descending, nNH = 1) and all filters (nNH = 4). On average, NH listeners could discriminate 8.55 ± 2.75 out of 12 filters.

Number of discriminable channels (out of 12) for the ascending and the descending presentation order. Each pair of open symbols connected by a line represents one subject. Concentric circles/squares denote that more than one subject was able to discriminate a particular number of channels (out of 12) when presenting stimulus pairs in the corresponding presentation order. Means for each presentation order shown with filled symbols. Left: cochlear implant (CI) users. Right: normal-hearing (NH) listeners. nCI = 12. nNH = 9.
Further Pitch Ranking Experiments in NH Listeners
To detect and eliminate possible confounding effects, three additional experiments were conducted on three additional groups of NH subjects. The first two experiments investigated whether presentation order effects were driven by loudness cues, specifically addressing two potential limitations: first, that white Gaussian noise increases in perceived loudness at higher frequencies due to the widening of critical bandwidths (Zwicker & Fastl, 2013). Second, that the initial loudness balancing procedure, when performed by inexperienced subjects, might inadvertently introduce loudness cues. The third experiment aimed to mitigate the presentation order effects by modifying the experimental procedure. The general experimental design remained identical to the main experiment, except for specific modifications described below.
In the first follow-up experiment, white Gaussian noise was replaced with uniform exciting noise (UEN) to compensate for the increased perceived loudness of white noise at higher frequencies due to the increased critical bandwidths in the inner ear (Zwicker & Fastl, 2013). 32 NH listeners (≤ 20 dB HL, mean age: 25 years) participated in the experiment. Despite using UEN, results shown in Figure 7

Pitch ranking performance by filter pair for adjacent filters in normal-hearing (NH) listeners for
In the second experiment (Figure 7
Finally, after suspecting a “regression to the mean” effect as the underlying cause for the interaction between presentation order and filter location (see Discussion), we conducted a third follow-up experiment consisting of three consecutive testing sequences on a new group of 8 NH subjects (≤ 20 dB HL, mean age: 28 years). Sequence 1 (immediately after training) included only the outermost filter pairs (F1/F2 and F11/F12). In this configuration, no significant difference was found between ascending and descending orders (Figure 8

Pitch ranking performance by filter pair for adjacent filters and uniform exciting noise in normal-hearing (NH) listeners when
Taken together, the follow-up experiments confirm that the presentation order effects in NH listeners were not due to loudness cues from the noise type or loudness balancing procedure, but rather to exposure to the full stimulus range, consistent with the regression to the mean effect.
Discussion
CI Pitch Ranking Difficulty Matched
CI and NH aggregate pitch ranking results (Figure 3) yielded highly comparable performances between both groups for all three tested spacing conditions (1-, 2-, and 3-electrode/filter distance). In addition, discriminability measures based on signal detection theory showed that, on average, CI users and NH listeners could discriminate 7–9 out of 12 channels, based solely on place pitch (Figure 6). Taken together, these results indicate that monopolar CI-electrode pitch ranking difficulty was effectively approximated at the group level for NH listeners by using shallow (6 dB/oct) bandpass-filtered white Gaussian noise. Our slopes are shallower than those proposed by Bingabr et al. (2008), who suggested values between 14 and 110 dB/oct to simulate monopolar stimulation. They are also shallower than the slopes used in other CI studies, such as 36 dB/oct in Adel et al. (2019) for electric-to-acoustic pitch matches in SSD CI and 24 dB/oct in Wang et al. (2015) to study loudness context effects. In a study similar to ours, where CI-electrode pitch ranking was compared to NH pitch ranking of noise bands, relatively steep slopes of 40 dB/oct were used (Laneau et al., 2006). Our results align more closely with studies showing that much shallower slopes (6–12 dB/oct) can approximate CI performance in speech recognition and speech-in-noise tasks (Fu & Nogaki, 2005; Oxenham & Kreft, 2014; O’Neill et al., 2019) We conclude that shallow 6 dB/oct slopes are suitable to approximate the mean pitch-ranking difficulty CI users experience in monopolar place pitch discrimination (Figure 3) and speech recognition (Fu & Nogaki, 2005; Oxenham & Kreft, 2014; O’Neill et al., 2019).
The aggregate pitch ranking performance (Figure 3) also shows that increasing the distance between electrodes (shallow filters) led to improved performance in CI users and NH listeners. These results align with research using acoustic pure tones, where performance increases with semitone interval between stimuli (Gfeller et al., 2007). Our findings are also consistent with previous CI studies showing that discrimination tends to improve with greater spatial separation between electrodes (Baumann & Nobbe, 2004; Collins et al., 1997; Eddington et al., 1978; Kwon et al., 2011; Nelson et al., 1995), suggesting reduced task difficulty as the excitation regions become more distinct. In NH listeners, no such distance-related improvement was observed with the steep condition (Figure 2), where performance for adjacent filters was already very high. This is likely because the pitch differences were sufficiently salient due to greater pitch strength (Zwicker & Fastl, 2013) and reduced stimulus overlap resulting from the 60 dB/oct filter slope (Mehta & Oxenham, 2017).
Presentation Order Effect in Pitch Ranking
Neither the presentation order nor the stimulus location alone significantly influenced pitch ranking results in CI users or NH listeners. Pitch ranking being independent of cochlear location has also been reported previously (Baumann & Nobbe, 2006; Nelson et al., 1995). However, in our study, we observed a significant interaction between presentation order and stimulus location for adjacent electrodes/filters (Figure 4). Specifically, descending pitch patterns performed better in apical (low-frequency) regions, whereas ascending patterns yielded higher performance at basal (high-frequency) locations. Notably, the significant order effect observed at E7/E8 (Figure 4
Frequency-dependent asymmetries between ascending and descending pitch patterns have also been reported in NH listeners: Raviv et al. (2012; 2014) reported a “contraction bias” (Poulton, 1979) in two-tone discrimination tasks, whereby large stimuli are underestimated and small stimuli are overestimated, Russo and Thompson (2005) found pitch register-dependent differences in interval-size judgments, and Greenwood (1997) described hysteresis effects when listeners equisected pitch intervals across a wide frequency range. Presentation order effects (time-order errors) have also been observed in CI tasks such as interval ranking (Luo et al., 2014) and melodic contour identification (Galvin et al., 2007), with a reportedly general higher difficulty when judging descending pitch contours.
The observed interaction between presentation order and stimulus location (Figure 4) can be interpreted in terms of a context-dependent decision process in which the second stimulus is compared not only to the first stimulus but to a weighted average of the first and the previously presented stimuli, with more recent inputs contributing more strongly (Raviv et al., 2012). This weighted integration effectively shifts the first tone in a pair toward the mean of the experienced stimulus range (central tendency), reflecting listeners’ expectations based on (recent) stimulus history. Consequently, descending pairs in the apical region (e.g., F2/F1) and ascending pairs in the basal region (e.g., F11/F12) resulted in higher performance. In the case of F2/F1, combining the first stimulus (F2) with previous predominantly higher-frequency stimuli shifts F2 towards higher frequencies, enhancing the perceived difference when F1 follows. Conversely, in the basal region, where prior stimuli are predominantly lower in frequency, the first tone in the pair is shifted towards lower frequencies, increasing the effective contrast for ascending sequences (e.g., F11/F12), resulting in higher performances. This history-dependent model based on weighted integration of recent stimulus history provides an explanation of what has previously been described as “regression to the mean” (“central bias” in Poulton, 1979; “regression effect” in Stevens & Greenbaum, 1966).
Similar biases have been demonstrated in interaural pitch matching, where the testing range influenced results. Restricting the available electrode or frequency range shifts pitch matches toward the center of the testing range (Carlyon et al., 2010; Goupell et al., 2019; Jensen et al., 2021; Schatzer et al., 2014), showing that listeners’ judgments are influenced by the stimulus set the subjects learn and adapt to throughout the experiment. In addition to the stimulus set, the listeners’ general auditory experience might also influence responses (Hellström, 1985; “adaptation level theory” in Helson, 1947; Russo & Thompson, 2005). Our findings are consistent with the testing-range bias: the pitch estimate of the first interval of each trial is biased towards the center of the stimulus range, thereby influencing pitch matches.
Sequential effects reported by Reiss et al. (2011; 2012) showed that previous pitch comparisons bias subsequent ones, leading to ambiguous matches if stimuli are not randomized (Table 1, Run 1 vs. Run 4 in Reiss et al., 2011; Table 1, Trials 3–10 and 90–98 in Reiss et al., 2012). In line with Raviv et al. (2012; 2014), these findings support that recent stimulus history contributes to the bias. A preliminary analysis of our data following Raviv et al. (2012) revealed trial-by-trial recency effects were present but comparatively small and showed substantial inter-subjective variability. Although these effects were modest, they indicate that recent stimulus history may have influenced responses. However, our study design did not explicitly control for sequential or starting-point biases (Adel et al., 2019; Carlyon et al., 2010; Goupell et al., 2019).
Finally, our results confirm that procedural biases are not specific to electrical stimulation or CI users (Carlyon et al., 2010). In our study, we replicated the interaction between presentation order and stimulus location in NH listeners using noise-band stimuli, where the effect was even more pronounced. The greater ambiguity in NH listeners may arise from the noise-band stimuli, which might have provided less salient pitch cues than electrical stimulation and represented an unfamiliar stimulus, making NH listeners more susceptible to bias. Moreover, the turning point where the higher-performing presentation order changed from descending to ascending (from apex to base) occurred at around F7/F8 in NH (Figure 4). Interestingly, the turning point in NH listeners lies approximately in the middle of the Bark scale (fc,F7 = 10.34 Bark, fc,F8 = 12.08 Bark), a psychoacoustical scale that reflects the way our auditory system processes loudness in critical bands (Zwicker & Fastl, 2013; Zwicker & Terhardt, 1980). The scale grows linearly from 0 to 24 Bark, dividing the audible 16 kHz frequency range into 24 bands, each of width one Bark. The NH turning point is also located in the middle of the mel scale (fc,F7 = 1186 mel, fc,F8 = 1402 mel), which relates to the human perception of changes in pitch (pitch ratios) (Beranek, 1949; Stevens & Volkmann, 1940). It goes from 0 to 2400 mel and is a power function of Greenwood's frequency-to-place map (Greenwood, 1990). In CI users, the turning point occurred earlier (E3/E4, E4/E5) than in NH. This discrepancy in turning points reflects that the acoustic frequency range used to approximate the CI electrode array may have been shifted toward lower frequencies, consistent with the electrode array's limited access to low-frequency nerve fibers.
Mitigation of Procedural Biases
Fechner (1860) already noted that experimental results can be affected by time-order errors and recommended testing both presentation orders so that the effects cancel out when combined. However, pooling across presentation orders does not guarantee compensation of time-order effects since order-dependent biases may interact with stimulus location in a non-symmetric manner. At a minimum, the experimental design should explicitly balance the presentation order of the stimulus pair, and gathered data should be checked for time-order effects before blindly pooling across presentation orders.
Beyond balancing the presentation order, mitigation strategies should also consider biases associated with the testing range. Our results suggest that stimulus pairs at the edges of the range were particularly affected. One approach is to expand the testing range, thereby redefining the “center” and discarding contaminated, irrelevant edge data (Poulton, 1979). In NH listeners, this can be easily done by including stimuli below F1 and above F12. In CIs, however, the entire electrode array was already used. To extend the pitch testing range further, lower pitches could be elicited by using phantom electrode techniques available in Advanced Bionic devices (de Jong et al., 2020; Macherey & Carlyon, 2012; Saoji & Litvak, 2010). Whether this approach leads to a reduction in biases in apical locations remains to be investigated. Another strategy to mitigate presentation order effects is to test stimulus pairs in isolation. Even testing only a limited number of pairs (e.g., F1/F2 and F11/F12) may already reduce ambiguities (Figure 8
Interestingly, several participants reported perceiving each stimulus pair as a pitch pattern (“up” or “down”) rather than two discrete pitches. They then translated this perception into terms of whether the “first” or the “second” stimulus was higher in pitch. Future studies might investigate whether directly asking for the direction of change instead of identifying the higher stimulus reduces the presentation order biases.
Sequential effects can also influence paired comparison outcomes. While we randomized our stimuli, pseudo-randomized designs may provide better control over sequential biases. Reiss et al. (2011, 2012) used a Latin-square sequence where the second half mirrored the first, reducing sequential effects. In Jensen et al. (2021), multiple starting points were measured to check for potential starting-point biases (Carlyon et al., 2010).
Pitch Ranking at Larger Frequency Intervals
The pitch ranking of non-adjacent electrodes and filters led to a tonotopic ordering of the stimuli, which was mostly independent of the presentation order. The significant interaction between stimulus presentation order and stimulation place found in the pitch ranking of adjacent stimuli (Figure 4) disappeared when ranking non-adjacent electrodes and filters (2- and 3-electrode/filter distance). This suggests that the procedural biases only influenced the outcome of difficult tasks, in which the perceptual difference between the compared stimuli was small. Consequently, the pitch rankings at the 2- and 3-electrode/filter distance conditions were more robust against presentation order effects (as was also the steep condition in the NH case). Nevertheless, individual subjects in both groups still exhibited order effects at the 2- and 3-electrode/filter distance, suggesting that the level of difficulty varied among individuals and might be affected, for example, by their level of musical training (Jiam et al., 2019; Leal et al., 2003; Sucher & McDermott, 2007) and their degree of neural survival.
Our findings are consistent with the observation that difficult or somewhat ambiguous tasks are more likely to be affected by non-sensory biases (Carlyon et al., 2010; Goupell et al., 2019; Reiss et al., 2012), as are judgments with small stimulus differences (Müller, 1903; Raviv et al., 2012). Carlyon et al. (2010) and Reiss et al. (2012) attributed strong, non-sensory biases in SSD CI to the perceptual mismatch between electric and acoustic stimulation. However, our data and recent results in BI CI users (Jensen et al., 2021) show that procedural biases can also affect can also affect stimuli which are presumably similar. These findings suggest that presentation order-related biases in CI pitch ranking likely reflect the inherently difficult discriminability of place pitch in CIs, which becomes most problematic for small electrode spacings.
Conclusion
We evaluated presentation order effects in unilateral CI-electrode pitch ranking and reproduced them in NH listeners using 6 dB/oct bandpass filtered noise as stimuli. Pitch-ranking difficulty was comparable between groups, with both discriminating approximately 8 of 12 channels. Presentation order effects (time-order errors) in pitch ranking were most pronounced for adjacent electrodes or filters, where perceptual differences were smallest. This procedural bias was explained by a context-dependent decision process in which the second stimulus is compared with a weighted sum of the first stimulus and previous stimuli. The bias was mitigated by testing on a reduced stimulus set; however, at the expense of giving up random stimulus pair presentation. These findings highlight the need to explicitly control and check for biases when designing and interpreting psychophysical measures.
Footnotes
Acknowledgments
We are particularly grateful to the study participants for their important contribution. We thank the Institute of Ion Physics and Applied Physics, University of Innsbruck for providing the RIB2. We also thank Andrew Oxenham, who pointed out that a bias to the center might be the cause for our observed results, and Lina Reiss and Matthew Goupell for their valuable insights and support. We are grateful to Bernhard Gleich and Barbara Gleich from the Munich Institute of Biomedical Engineering, Technical University of Munich, for their assistance with the generalized linear mixed-effects analysis. Finally, we thank Andrew Oxenham and the two anonymous reviewers for their comments, which helped improve the clarity and interpretation of this manuscript. This work was partially presented at the 20th International Symposium on Hearing (ISH 2025) in Vienna, Austria, June 2025.
Ethical Considerations
Measurements were conducted in accordance with the Declaration of Helsinki and approved by the medical ethics committee of the Klinikum rechts der Isar (Munich, 2126/08).
Consent to Participate
All subjects provided their written informed consent. CI study participants received monetary compensation.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This project was funded by a grant from MED-EL (Innsbruck) and the German Research Foundation (DFG, Deutsche Forschungsgemeinschaft) [project number 415658392].
Declaration of Conflicting Interests
The authors declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: W. H.'s institution has received research grants from MED-EL, a leading cochlear implant manufacturer.
Data Availability
The datasets generated during and/or analyzed during the current study are available from the corresponding author on request.
