Abstract
The article explores the transnational circulation of sound-based expertise that came together in the establishment of audio forensics in Cold War Poland and Czechoslovakia. It is the first to study the Polish Phonoscopy Lab, the first ever audio forensic police department, founded in 1963 to carry out original research on speaker identification and distorted speech. It shows the complexity of sound dissection science behind the Iron Curtain, where detailed methodology, the use of precision technological equipment, and the application of up-to-date phonetic, acoustic, and linguistic knowledge, as well as practical guidelines for trained group listening, were developed. It argues that by developing such methodology, the Polish and Czechoslovak phonoscopy labs were leading the way in reformulating the notion of sound-based objectivity in forensics. By examining the transfer of methods, theories, technological devices, and visions between the Cold War East and West, the article shows that from the point of view of the development of expertise, there were specific advantages of the embeddedness of forensics in the structures of the totalitarian state, which granted access to technologies, ensured a steady supply of cases, and simplified investigation procedures.
Introduction
In the Forensic Technology section of the Czech Police Museum in Prague, a peculiar apparatus has been on exhibit since the 1990s, informing visitors of otherwise unknown investigations into sound conducted by Czechoslovak forensic experts in the mid-1970s. The so-called ‘Voiceprint’ machine, which bears the trademark ‘Voiceprint Laboratories, Inc.’, was made in Somerville, New Jersey in 1964 by a private company founded by Lawrence Kersta, a U.S. engineer who commercialized previously classified research on speaker identification done at the Bell Telephone Laboratories during World War II (WWII). The device promised to produce spectrographic visualizations of the human voice with a biometric value equivalent to that of matching fingerprints, and it found its way to the Prague Institute of Criminalistics in 1975, when a specialized department devoted to audio forensics was established there by the Communist state police. But how did the device travel so far across Cold War ideological borders to Czechoslovakia? And what kind of story about forensic voice identification behind the Iron Curtain can it tell?
The article explores the transnational circulation of sound-based expertise that came together in the establishment of audio forensic police departments in Cold War East-Central Europe, in particular, in Poland and Czechoslovakia. 1 It argues that whereas in the US forensic voice analysis was performed primarily by engineers and audiologists and was strongly influenced by the voiceprint methodology, the European research on the individuality of voice was arguably more sophisticated. The article shows that criminalistic sound dissection in Communist East-Central Europe was not defined by Soviet research on speaker identification but was shaped primarily by pioneering audio forensic work conducted in Warsaw.
This study is the first to pay systematic attention to criminalistic investigations of sound in Communist Poland, where original research on speaker identification and distorted speech was carried out starting already in the early 1960s, making the Polish lab one of the first audio forensic institutes on the continent. The Polish laboratory of forensic phonoscopy was established to identify unknown voices in wiretapped conversations and threatening calls, and to develop workable means of dissecting sound into component parts that could be analyzed objectively. In doing so, it mobilized different kinds of knowledge along with specialized audio equipment, some of which originated in Western Europe and the US, while others were developed and tested behind the Iron Curtain.
While the US voiceprinting methodology and the development of forensic phonetics and acoustics in Western Europe have been explored in a number of studies (Foulkes and French, 2012; French, 2017; Hollien, 1990; Li and Mills, 2019; Mopas, 2023; Nolan, 1997), until recently, historians had paid no attention to voice identification programs in the Soviet Union and countries of the former Eastern Bloc. This article draws on recent pioneering studies in the history of science and science and technology studies (STS) of the German Democratic Republic (GDR) and Czechoslovak audio forensic programs (Bijsterveld, 2021; Kvicalova, 2023a, 2023b), and considers them in light of the development of sound-based criminalistics in Poland – which was instrumental for the establishment of the audio forensic lab in Prague in the mid-1970s – but also in connection to what can be reconstructed about Soviet speaker identification research. The primary sources analyzed in this article include forensic handbooks and journals, police information bulletins, conference reports and proceedings, as well as unpublished archival documents, most importantly, those from the Institute of Criminalistics in Prague and the Security Services Archive. The study stems from the author’s previous investigation of the criminalistic Department of Phonoscopy in Cold War Prague, in which a wide range of source materials, including interviews and films, were used to shed light on the nature and status of acoustic expertise in the legal and security system of Communist Czechoslovakia, thereby contributing primarily to sound studies and STS literature concerned with the use of ‘sonic skills’ in the natural and social sciences (Kvicalova, 2023a, 2023b). This article adopts a different perspective and focuses on the nature of the circulation of audio forensic knowledge in the Cold War era more broadly. By combining Czechoslovak sources with those from Communist Poland against the backdrop of the evolution of forensic phonetics and acoustics on both sides of the Cold War divide, the study is the first to provide a more systematic insight into the development of sound-based forensics behind the Iron Curtain.
In what follows, I will examine the transfer of methods, theories, technological devices, as well as plans and visions between the Cold War East and West, and among countries of the former Eastern Bloc, to shed light on the nature of knowledge transfer in the period. By approaching the history of audio forensics from the vantage point of the transnational circulation of expertise, the article builds on recent scholarship that has shown the multidirectional movement of expert knowledge in the Cold War human sciences, which shaped scholarly, social, and political practices on both sides of the East–West divide (Christian, Kott, and Matejka, 2018; Iacob et al., 2018; Lišková and Fisher, 2024; Solovey and Dayé, 2021).
In the first circle: Voiceprint and visible speech in the USSR
In the novel by Russian Nobel Prize winner Aleksander Solzhenitsyn, In the First Circle (V kruge pervom), a group of scholars and technicians arrested during Stalin’s purges after WWII is imprisoned in a specialized research bureau (sharashka) to work on various technical and scientific tasks. The novel revolves around a speaker identification case, which is requested from the Acoustic Lab by the secret police. Depending on which edition of the book one reads, the secret police wiretap a call made to either a family doctor (in the censored version) or the US embassy by an unknown person, whose identity needs to be disclosed. 2
When the character Lev Rubin, a philologist, is asked to identify the voice, we learn several intriguing details about the nature of speech analysis he was performing: he categorized voices under different types, focused primarily on voices heard via telephone, and made use of sound spectrograms, which are also referred to in the text as visible speech (видимая речь – vidimaja rech’). The Acoustic Lab is also said to be equipped with ‘a visible speech device for recording speech visually’, which prints off ‘voiceprints’. Rubin sets out to identify the speaker by ‘voiceprinting the telephone message and comparing the results with the voiceprints of the suspects’ (Solzhenitsyn, 2008: 21, 235).
The events in the book are set in December 1949, while Solzhenitsyn himself wrote the text between 1955 and 1958. As such, the autobiographical novel provides a unique historical glimpse into the Soviet knowledge of forensic sound analysis in the post-war period, which is otherwise not available for study. Notwithstanding that it is a work of fiction, the novel testifies to the fact that the Soviets were familiar with the spectrograph, an apparatus for graphically representing time, frequency, and loudness of voice, originally designed at the US Bell Telephone Laboratories in the 1940s. 3 The description of sound dissection performed in the Acoustic Lab shows striking methodological similarities to Russian speaker identification research and that conducted in the US during WWII and after. While this paper is not primarily concerned with the study of audio forensics in the Soviet Union, a close examination of Solzhenitsyn’s novel clearly exhibits the relationship between Russian and American investigations into the individuality of voice, and thus sets the stage for a better understanding of the originality of sound-based criminalistics as it was developed in East-Central Europe amid the Cold War bipolarity.
The term visible speech has its origin with Alexander Maleville Bell and refers to historical attempts to teach hearing-impaired people to articulate, first by reading phonetic symbols, and later – in a system devised by Bell’s son, Alexander Graham – by means of visual translation of speech made by the spectrograph. The project of ‘visual hearing’ aimed to enable the deaf to read real-time telephone conversations, and focused on the transmission of speech content as opposed to individual characteristics of speech (Li and Mills, 2019; Mills, 2010; Potter, Kopp, and Green, 1947). The latter became the primary focus during WWII, when the US military initiated research on speaker identification at the Bell labs. The project, which sought to identify individual voices with the spectrograph, was classified during the war, and although it remains unclear to what extent it actually enabled the Americans to identify the voices of German radio operators (Tosi, 1979: 67), interest in ‘voiceprint’ technology prevailed in the security context of the second half of the 20th century. The ‘voiceprint technique’ was used by the FBI for ‘investigative support’ from the 1950s (Koenig, 1986: 2089), although its forensic application was seriously considered only after Lawrence Kersta, a former engineer at the Bell labs, popularized the research on speaker identification in the early 1960s (Kersta, 1962).
In Solzhenitsyn’s novel, the similarity of the Soviet efforts at sound dissection to the American project of visible speech is conspicuous in regard to the terminology, the methods of voice analysis, and the technological equipment used. Although one might speculate as to whether the Soviets developed original research on speaker identification during WWII, a close reading of Solzhenitsyn’s uncensored version of the novel leaves no doubt about the origin of the voiceprint apparatus used in the Russian lab. While the Deputy Minister is assured by the head of the Acoustic Lab that the visible speech apparatus was ‘designed in our laboratory’, the reader is assured that he ‘had already forgotten the extent to which the design had been borrowed’ (Solzhenitsyn, 1968: 454). The uncensored version further specifies that ‘they had lifted it from an American journal’ (Solzhenitsyn, 2008: 236; emphasis added). The motive of the presence of American, British, and German equipment along with ‘twenty thousand new acquisitions of … technical literature’ (Solzhenitsyn, 2008: 52) appears in the novel more than once and reflects an actual practice of industrial, technical, and intellectual property espionage during the Cold War.
If we take Solzhenitsyn at his word, the novel suggests that the functioning of the apparatus built on the American model did not quite measure up to expectations. The voice spectrograms it produced were hard to interpret and could only be used in conjunction with other information about the suspects.
4
To please the Deputy Minister and to get political support for further developing voice identification techniques, Lev Rubin pretends to be able to read detailed information, including the exact content of the utterances, from the sonograms, while he privately confesses to his colleague that the job is going to be very tough as the research on voice has only just started: This was, in effect, a new science: identifying a criminal by a voiceprint. Till then, criminals had been identified by their fingerprints … and [the technique] had evolved over the centuries. The new science could be called … ‘phonoscopy’. And it had to be created in a matter of days. (Solzhenitsyn, 2008: 247)
The novel is surprisingly accurate in its description of the nascent field of audio forensics in the second half of the 20th century: it relied heavily on sound spectrograms, although they were hard to interpret; it made frequent use of the analogy to fingerprints and tended to overstate the informative value of ‘voiceprints’; and it had to deal with the complexity of the human voice, which could not be fully scrutinized using electroacoustic methods only.
By the late 1960s, the initial enthusiasm regarding the military and, by extension, forensic use of voice spectrograms for identification purposes was replaced by moderate disillusionment. Kersta’s Voiceprint Laboratories marketed ‘voiceprinting’ based on a crude side-by-side pattern matching of voice spectrograms as a quick and reliable identification method. 5 A committee of the Acoustic Society of America, however, concluded that the analogy between voice spectrograms and fingerprints was inaccurate and that their use in legal disputes was unreliable and potentially misleading (Bolt et al., 1969, 1970; Committee on Evaluation of Sound Spectrograms, 1979). Despite the reduced status of the visual voiceprint method, the notion of vocal fingerprints proved enduring and fueled the development of forensic speaker identification behind the Iron Curtain well into the 1980s. 6 Notwithstanding the lingering influence of the voiceprint metaphor in the Eastern Bloc, this article shows a distinct East-Central European tradition of forensic sound analysis that differed considerably from the criminalistic sound dissection as it developed in the US. One of the first audio forensics labs on the continent was established in Communist Poland, which, as I argue in what follows, developed methods of sound dissection that not only borrowed – both technologically and methodologically – from the West but formulated its own original procedures and conclusions.
The ‘new science of phonoscopy’
In the beginning of the 1960s, the Polish Supreme Court first recognized a tape recording as valid evidence in a criminal trial, and further specified that if a sound recording was admitted as material evidence, it ‘in turn requires a proof of the identity of both the recorded voices and the tape itself as well as a proof that no alterations have been made in it’. 7 The rulings gave impetus to the establishment of audio forensics as a specific field of criminalistics which was to deal not only with voice identification, but should also be able to perform a comprehensive analysis of the recorded material, including the authenticity of the tape.
Forensic fonoskopia, as the branch of expertise was called in Poland, was systematically developed at the Department of Forensic Science at the Citizens’ Militia Headquarters (Zakład Kryminalistyki Komendy Głównej MO) in Warsaw, where the first phonoscopic lab was established in early 1963, headed by Stanisław Błasikiewicz (Błasikiewicz, Miściuk, and Wójcik, 1967). The term ‘phonoscopy’, which curiously appeared already in Solzhenitsyn’s novel from the late 1950s, was widely adopted in Poland and later on in Czechoslovakia. 8 Unlike the English terms ‘forensic acoustics/phonetics’, or their German equivalents ‘Forensische Akustik/Phonetik’, phonoscopy does not specify the method used, only the object of study, that is, sound. The Warsaw Phonoscopy Lab worked on different kinds of audio materials which were increasingly submitted for analysis by criminal law enforcement agencies and the judiciary in the 1960s and 1970s. This stemmed not only from the new legal status of the ‘audio-document’ but also from a dramatic increase in general availability of tape recorders among the public, which meant that evidence recorded on magnetic tapes became more and more common in civil disputes. 9
When voiceprinting gained momentum in the US after the publication of Kersta’s article in Nature in 1962, Polish forensic phonoscopy started developing complex voice dissection methodology, which combined expertise from acoustics, phonetics, linguistics, phoniatry, and dialectology, along with the use of precision laboratory instruments which were both purchased abroad and constructed specifically for audio forensic purposes. The nature of phonoscopic sound dissection, which was being developed by the newly established lab, required interdisciplinary cooperation, and from the very beginning it was performed by expert teams as opposed to individual specialists (Błasikiewicz, Miściuk, and Wójcik, 1967).
To compare, in the UK, the first expert evidence regarding someone’s identity based on the sound of his voice recorded in a phone call was presented by a linguist named Stanley Ellis in 1967 (Foulkes and French, 2012: 560). While in the UK early audio forensics relied exclusively on auditory-phonetic analysis, and the American discussions regarding the admissibility of sound-based evidence in court revolved around Kersta’s voiceprint technique well into the 1970s (French, 2017), Polish experts found a way to integrate phonetic, linguistic, and acoustic approaches into one analytical framework in the 1960s.
When looking back at the first decade of Polish phonoscopy, Stanisław Błasikiewicz noted that when establishing the lab in Warsaw in 1963, they first studied ‘the scarce domestic and foreign literature on the subject’ (Błasikiewicz, 1976: 498). Most of the available literature was, however, considered insufficient as far as practical guidance for forensic work was concerned, as it formulated purely theoretical insights based on studies carried out in experimental settings. Forensic phonoscopy, in contrast, had to deal with complicated real-life situations when analyzing material provided by investigative and judicial institutions. Most of the sound recordings submitted for examination were of poor quality, either on account of bad recording equipment or its mishandling, or due to bad acoustic conditions of the original event. The speech was often unintelligible, and the recorded voices could be affected by emotional distress or deliberately altered. To find a workable way of analyzing the submitted audio material, the lab had to find its own means of sound dissection that would satisfy the demands of police and legal authorities and guarantee the ‘objectivity of the research’ (ibid.: 500). This involved not only the development of theoretical–methodological approaches but also guidelines for selection and training of personnel, including the identification of best listening practices, a detailed technical and acoustic set-up of the phonoscopic studio, and the specification of equipment for recording, measuring, playback, correction, and visualization of sound.
The documents pertaining to the development of the Polish Lab of forensic phonoscopy, including the papers, reports, and guidelines produced by Stanisław Błasikiewicz and his team, thus represent a key source of information for the study of the development of audio forensics behind the Iron Curtain, whose image has long been shaped solely by the reading of the censored version of Solzhenitsyn’s In the First Circle. When interpreted in conjunction with other East-Central European sources, the Polish audio forensic practice emerges as formative for the development of the field in the region.
Cooperation and secrecy behind the Iron Curtain
Starting in the early 1960s, forensic practitioners in the Eastern Bloc, including from Poland, Czechoslovakia, the GDR, Bulgaria, Hungary, and the USSR, would meet for annual forensic symposia to share expertise in various fields of criminalistic practice. The topic of speaker identification occasionally appeared on the program, especially when Polish forensic experts presented their papers on the subject. 10 Actual institutional collaboration, in which forensic practitioners would share their know-how with colleagues from friendly socialist states, was a different matter, though. In this respect, the foundation of the Czechoslovak Phonoscopy Lab (oddělení fonoskopie) in the mid-1970s provides us with a rare glimpse into the nature of international expert exchange in audio forensics.
Before establishing the Phonoscopy Lab as part of the Prague Institute of Criminalistics, the Czechoslovak delegation visited Warsaw in 1975 to learn about the new field. 11 The Polish experts were willing to share rather detailed knowledge about how they performed sound analysis at the Main Headquarters of the Citizens' Militia (KGMO): Jan Málek, the head of the Prague Lab, remarked that they were welcomed by Stanisław Błasikiewicz, who was very ‘friendly’ and, together with his colleagues, provided them with much needed advice regarding both methodology and equipment. 12 The Prague Lab focused on the same areas of investigation, including the (a) authenticity of sound recording and identification of the recording device, (b) identification of speakers, (c) analysis of background noises, and, although less profoundly than Polish phonoscopy, (d) the reconstruction of distorted speech; and also drew on the same methodology, which combined spectral (acoustic), phonetic, and linguistic analysis. The delegation left Warsaw with detailed information on methodology as well as with instructions regarding the practical arrangement of the phonoscopy studio, including a comprehensive list of equipment along with a clear idea of the personnel it wanted to hire. 13 Málek spent two intense weeks at the Polish lab in 1975, but the cooperation continued throughout the 1970s and 1980s.
Although the professional contacts with Poland were crucial for the development of Czechoslovak phonoscopy, Jan Málek also drew inspiration from a visit to East Berlin in 1981. As recently shown by historian of science Karin Bijsterveld, the East German secret police, the Stasi, had its own voice analysis program in which it closely cooperated with the Kriminalistische Akustik department at Humboldt University in Berlin. Málek visited the leader of the university research effort, Christian Koristka, with whom he discussed the German method of standardized aural analysis, the issue of background noises, and, above all, automatic speaker recognition. 14 Despite their differences in size and institutional set-up, East German, Polish, and Czechoslovak audio forensics shared basic methods and premises of sound analysis. The documents from the Czechoslovak visits to Berlin and Warsaw represent a unique source of information on the nature of the exchange of sound-based expertise behind the Iron Curtain, which has not yet been studied. However, whereas there is direct evidence of actual cooperation especially between the Poles and Czechoslovaks, collaboration with the Soviets was an entirely different matter.
Despite a host of indirect evidence about the existence of Russian speaker identification research, there are no publicly available archival sources that inform us about the exact institutional and methodological background of audio forensics in the USSR. The existence of a speaker identification program in the Soviet Union is confirmed by Czechoslovak, Polish, and GDR sources, all of which, however, express general disappointment about the extent of actual cooperation between Moscow and the criminological institutes in Prague, Berlin, and Warsaw. The KGB was apparently not willing to share information about its voice identification research, and although there is evidence of technological exchange between the USSR and Eastern Bloc countries, especially in regard to sound recording and eavesdropping devices, a more intense scientific and institutional collaboration was repeatedly called for. The East German Stasi set out to exchange knowledge with the KGB already in the 1960s, but no real cooperation took place until the end of the Cold War. 15 The limited knowledge of the actual state of research in the Soviet speaker identification program is exemplified in a reference from the Czechoslovak phonoscopy handbook written in the late 1980s, mentioning quite vaguely that it was ‘going in the right direction’ and that it was based on spectral dissection as opposed to auditory methods of analysis (Málek and Musilová, 1989: 7). 16
In the archives of the Prague Institute of Criminalistics, there is a note from August 1980, mentioning a conference on voice identification to be held in Kiev in October of the same year, with a request that the head of the Phonoscopy Lab in Prague, Jan Málek, should participate in the event. It further specifies that, unlike already existing cooperation with Poland and the GDR, ‘it has not yet been possible to organize a research internship’ for a Czechoslovak phonoscopy expert in the USSR. 17 The note is quite telling as it not only confirms the lack of actual cooperation between the two Communist countries, but it also shows that international declarations of expert exchange often met with practical difficulties. Málek never went to Kiev and no further information about the conference is to be found in the materials of the Prague phonoscopy lab, including no further reference to Professor Saltěvecký, a Soviet voice identification expert whose name is mentioned in the letter.
Expertise travels: Technology, literature, and dreams of an automated future
The study of both capitalist and socialist literature, along with participation in international science exhibitions and fairs, was cited as the official strategy for the development of intelligence and forensic technology behind the Iron Curtain. Despite continuous efforts of Eastern Bloc countries to develop original eavesdropping devices that would work on novel technological principles, including acoustic pressure or laser, the most up-to-date technology was often purchased abroad – typically in cash through an intermediary in neighboring countries such as West Germany or Denmark. 18 The Czechoslovak ‘Voiceprint’ made by the Voiceprint Laboratories (based in New Jersey) that is now housed by the Czech Police Museum was probably acquired in a similar way. The practice of copying designs from Western literature, mentioned by Solzhenitsyn, was also part and parcel of the development of surveillance technologies, which relied either on publicly available sources or on technological espionage. 19
Technology clearly did not respect ideological borders and routinely travelled to the East. The list of essential acoustic equipment which the Czechoslovaks brought from Warsaw before establishing an audio forensic department in Prague included both domestic technology, such as the Tonette tape recorder and frequency correctors made in Poland, and apparatuses of Western origin, including those constructed by Danish company Bruel and Kjaer, the West German portable UHER reel-to-reel tape recorder, and US-made Missilyzer spectrographs. 20
Both Czechoslovakia and Poland also saw the development of original technological devices, some of which found their way to the West. Most importantly, this concerned research done in phonetics, which informed both Polish and Czechoslovak approaches to the study of voice from the very outset. At Charles University in Prague, Přemysl Janota performed many original experiments in the 1950s and 1960s, in which he analyzed the speech spectrum of individual speakers and investigated synthetic speech. He designed and constructed the ‘segmentator’, an apparatus that could slow down a speech sample without distorting it, and a ‘vowel synthesizer’, which was used for experimental purposes at the Institute of Phonetics from the early 1960s (Janota, 1967). Forensic phonoscopy in Prague drew on this strand of university phonetic research, which, however, was never done in direct institutional conjunction with criminal investigations of sound as was the case in the GDR.
The Phonoscopy Lab in Warsaw could, in turn, draw on the work of renowned Polish phonetician Wiktor Jassem. Jassem was a pioneer in the research of speech signal analysis, which he based on discrimination of the speaker’s individual formant frequencies. He was a world-famous phonetician active in the International Phonetics Association from the early 1950s, whose research in phonology, experimental phonetics, and speech technology influenced specialists from all over the world, including Harry Hollien from the University of Florida (Hollien, 1990; Jassem, 1968, 1984; Windsor Lewis, 2003). Jassem was based at Adam Mickiewicz University in Poznan, where he was able to acquire up-to-date technologies at that time, including spectral analysis filters which allowed him and his team to develop automatic speech segmentation methods for which he was awarded the international Kay Elemetrics Prize in 1983. Although based in Poland, he received international visits in Poznan and spent eight months at the University of Calgary, Canada in 1977–1978 (Gibbon, 2020; Windsor Lewis, 2003). When Harry Hollien, a phonetician by training, took interest in developing semiautomatic speaker identification systems in the 1970s, he also called on the expertise of Polish engineer Wojciech Majewski from the Wrocław University of Technology, who worked on the analysis of the speech spectrum (Hollien and Majewski, 1977).
Throughout the 1960s and 1970s, one of the top priorities of the Polish Phonoscopy Lab was to reduce the labor intensity and excessive time consumption of the analytical process (Błasikiewicz and Bednarczyk, 1979: 716; Błasikiewicz, Miściuk, and Wójcik, 1967: 326). By the end of the 1970s, Błasikiewicz still estimated that a full analysis of one speech parameter, such as a Fo (a fundamental frequency of the respiratory tone), for two speakers, including measuring, calculating, and interpreting the results, took about 12 hours (Błasikiewicz and Bednarczyk, 1979: 716). To compare, the model of Kersta’s voiceprint spectrograph deposited at the Police Museum in Prague came with a leaflet advertising a ‘Voiceprint identification service’ which promised to perform voice identification on a ‘24-hour schedule’ after receiving a tape recording of the unknown voice and tapes of voices of the suspects. 21
Audio forensic experts in Western Europe and the US, as well as their Eastern Bloc counterparts, all invested in the development of automated speech dissection, which promised to radically reduce the time required for the analysis. Although an automated system for finding matching voiceprints looked appealing, both Polish and Czechoslovak experts were well aware that voiceprint identification as described by Kersta could not be performed when analyzing distorted recordings or when comparative material did not exactly match the original speech situation, which was always the case in forensic practice (Błasikiewicz and Bednarczyk, 1979: 714). Despite considerable progress in the research of automatization programs of speech recognition and verification of speakers in the 1970s and 1980s, identification in real-life conditions was much more complicated. In examining the subject, Polish phonoscopy experts reviewed a host of available foreign and domestic literatures. Apart from the aforementioned research in phonetics (Hollien and Majewski, 1977; see also Jassem, 1966), the Polish investigation into automated speech recognition was best represented by the work of acousticians Janusz Kacprowski, Czesław Bazstura, and Wojciech Majewski (Basztura and Majewski, 1978), while the (mostly) English-language literature on the subject was quite varied (Rosenberg and Sambur, 1975; Sambur, 1972). 22 At the International Criminalistic Conference in Oxford in 1970, for example, the AUROS (automatic recognition of speakers by computer) system was introduced, which was developed at the Philips Research Institute in Hamburg together with the Heinrich Hertz Institute in Berlin (Bunge, 1977). Polish phonoscopy experts, who could not take part in the event, thoroughly studied the descriptions of the technology in the conference proceedings, but concluded that its use for forensic purposes was still limited. As for the transnational exchange of acoustic knowledge more generally, it is noteworthy that both Polish and Czechoslovak experts frequently referred to American acoustic research, most notably, the Journal of the Acoustical Society of America.
Whereas in the US and the UK audio forensics was performed by independent experts – engineers, acousticians, and, in the case of the UK, university academics working in phonetics, dialectology, or sociolinguistics (French, 2017) – in Communist Poland and Czechoslovakia, forensic sound analysis was developed and tested primarily in police criminalistic departments. From the point of view of the development of audio forensics, the institutional backing by the totalitarian state had its advantages: whereas experts in the private sector often found themselves in a difficult position when they wanted to acquire technological equipment under governmental control, such as cepstral analyzers or certain types of sound (Kvicalova and Bijsterveld, 2023: 58), their Communist counterparts could count on the support of the Ministry of the Interior (Málek and Musilová, 1989: 78). 23
Another advantage from the point of view of the development of sound-based forensic expertise was the steady supply of cases the totalitarian state needed to analyze. While in the UK the number of cases using forensic speaker comparison in the early 1980s was insignificant (Foulkes and French, 2012: 560), the police phonoscopy experts in Poland and Czechoslovakia routinely analyzed voices of dissidents, foreign diplomats, corrupt state officials, hoaxers, and frauds for different police departments. The Phonoscopy Lab in Warsaw had already analyzed hundreds of cases by the mid-1970s (Błasikiewicz, 1976: 501), while in the UK the number of criminalistic voice analyses increased only after the introduction of the Police and Criminal Evidence Act (PACE) in 1984. The PACE legislation ruled that all interviews with suspects at police stations should be tape recorded, which meant that, from that point, there was always a sample of speech that could be readily compared with the original voice recording (Kvicalova and Bijsterveld, 2023: 53). Prior to that, a person could only be asked to give a voluntary voice sample, which, in contrast, was not an issue for the Communist police, who could summon suspects for speech tests without their explicit consent.
Although some of the essential technological equipment in the phonoscopy labs in both Warsaw and Prague came from the West, the methodological expertise was not simply copied from American, British, or French journals, but was developed in its own right, and the analysis performed by Polish and Czechoslovak experts rivaled that done in American or British laboratories. One such example is the Polish research on whispering and severely distorted speech, which shall be discussed in the following section.
Distorted sounds
Stanisław Błasikiewicz’s experimental research on whispering and distorted speech was an original attempt to formulate new theoretical insights as well as practical recommendations for analyzing poorly intelligible recordings. The initial research, which was done under experimental conditions in the Warsaw Lab, was based on the examination of 152 recordings that were submitted for phonoscopy analysis by law enforcement and criminal justice agencies between 1963 and 1970 (Błasikiewicz, 1971). This concerned voices recorded in the cabins of airplanes and in cars parked in noisy streets, conversations secretly recorded in busy cafes or over the phone, or conversations captured by microphones placed at great distance from the speakers.
In recovering the content of disturbed recordings, exact knowledge of the nature and type of interference and the proper application of corrections using frequency response equalizers was essential (Błasikiewicz, 1971: 173–4, 177–8). The documents of Polish phonoscopy are, however, quite unique in the amount of detail they give not only on the best equipment for recording, measuring, and filtering sound, but also in regard to specific directions on how it should be arranged in the room. The ‘listening studio’ (studio odsluchowe) and the main room equipped with recording apparatus were to be connected visually and via intercom, and specific recommendations were also given on lighting and furniture (ibid.: 169–70).
Even when making the best use of available technology, the analysis would not have been possible without skilled listening. Despite being an indispensable component of any auditory-phonetic analysis, listening techniques as such are rarely discussed in primary sources. Polish phonoscopy, in contrast, drafted several rules concerning techniques for listening (technika odsluchu) to sound recordings for its staff to follow, which give us precious insights about forensic bodily practices and ‘sonic skills’. 24 The most important direction specified that listening was to be done several times and, ideally, by a group of experts. According to Błasikiewicz, ‘the best results in reconstructing the content of whispering and intensely distorted speech are obtained by team listening’ (odsłuch zespołowy) (Błasikiewicz, 1971: 177). Before the listening task, phonoscopy experts should be isolated from their surroundings and external noises for some time, which should enhance their concentration. Depending on the degree of clarity and intelligibility of the utterance and the nature of the interference, listening should not exceed five to seven hours a day, with sufficiently long breaks between sessions to allow the perceptual organ to rest. To prevent mishearing and making false associations, individual listeners should write down the content they listened to independently of each other and then compare results. If different listening results were obtained, the analysis of the disputed passages was repeated. Among the many requirements of a phonoscopy expert, very good hearing and auditory memory along with the ability to focus and distinguish high tones were listed – a requirement that also appears in the documents of Czechoslovak phonoscopy, which further specified that the experts should be ‘musically gifted’. 25
The report on the analysis of whispering and distorted speech from 1971 specified that more that 90% of speech content in the analysis of the 152 cases was recovered by phonoscopy experts. By 1976, they claimed to have reconstructed 94% of more than 500 audio documents submitted for analysis. This was possible thanks to a ‘skillful use of the special electroacoustic equipment’; a correct application of the rules of phonetic, linguistic, and acoustic analysis; a thorough knowledge of ‘all psychophysiological phenomena occurring in the speech process’; and the creation of appropriate listening conditions. If all the conditions were met, it was possible to reproduce the content ‘quickly and flawlessly’ (Błasikiewicz, 1971: 169–70; 1976: 506–7). The tone of confidence, along with the reported accuracy of the method, was part of the process of establishing the new branch of forensic expertise in Communist Poland. It is important to emphasize, however, that the high success rate claimed by Błasikiewicz’s team only concerned the recovery of speech content; the identification of unknow speakers based on their recorded voices was an entirely different matter.
Sound at a trial
Although sound recordings were admitted in courtrooms in both Poland and Czechoslovakia, their legal status remained complicated, and they had to go through a complex ‘translation’ process to be accepted as valid legal evidence. 26 This had to do with the status of hearing and sound not only in criminalistics but in the sciences in general. While objective knowledge in the forensic sciences was historically linked to sight and visual methods of recording, such as photography or fingerprints, knowledge gained through listening was associated with subjectivity (Adam, 2020; Kvicalova, 2023b; Neale, 2020). The vision of a fully automated dissection done by computers, which promised objective analysis free from subjective human interference, was shared across forensics since the 1960s (Adam, 2020; Bijsterveld, 2021). 27 However, computer technology that would allow for such analysis was not available at the time, and different national traditions disagreed on how to achieve valid acoustic knowledge using methods other than purely mechanical ones. Historians of science Lorraine Daston and Peter Gallison describe a discourse of ‘trained judgement’ which supplemented the ideal of ‘mechanical objectivity’ in the 20th century (Daston and Galison, 2007). As far as Polish and Czechoslovak phonoscopy was concerned, the judgment of a trained expert, who could not only operate the acoustic equipment but also interpret speech sonograms and listen analytically to the recorded voices, was indispensable. The phonoscopy authorities tried to get the subjective element under control by drafting specific guidelines for specialized team listening, which were to be used in conjunction with other analytical methods.
The validity of phonoscopic analysis and its use in legal proceedings was openly debated in Polish legal discourse. The criminal law journal Problemy prawa karnego published an overview paper on phonoscopy in 1980, quoting a host of English but also French, German, and Russian publications on the subject, connecting the emergence of the field to the developments in acoustics in the 1950s and 1960s, and to Kersta’s voiceprint methodology (Legień, 1980). The paper takes a rather critical stance towards audio forensics. Despite its ‘undeniable achievements’, including the original research of Stanisław Błasikiewicz regarding the improvement of audibility of severely disturbed speech and whispering using frequency filters, the article concludes that the ‘recognition of it [phonoscopy] as a fully useful technique in forensic science, and in particular in evidentiary proceedings, is still a difficult matter’(ibid.: 71). The reason for this incomplete confidence in phonoscopy was the ‘ever-emerging identification difficulties and interpretative doubts, arising as experimental research continues’. The author of the article advises ‘great caution and restraint in the widespread introduction of this technique into evidentiary proceedings’ (ibid.: 71–2).
The doubts regarding the validity of forensic voice identification did not only concern the methodology of the field as such, but also more specifically the work of the Phonoscopy Unit at the Citizens’ Militia in Warsaw. The author points out that the absence of other phonoscopy labs in Poland made proper peer review impossible and raised questions regarding ‘the evidentiary value of the expertise performed by the existing center' (Legień, 1980: 71). He warns against the ‘exceeding enthusiasm’ of some domestic researchers – quoting Błasikiewicz (1976), and Błasikiewicz, Miściuk, and Wójcik (1967) – who, in contrast to ‘foreign literature’, present phonoscopy as a fully reliable technique. The legal article not only displays a surprisingly high level of engagement with the available literature, but is also unusually critical of the Phonoscopy Lab. The criticism was partly misplaced as the article insufficiently recognizes the distinction between different national audio forensic traditions and did not fully appreciate the complex methodology developed at the Warsaw lab; and in particular, it fails to discriminate between Błasikiewicz’s conclusions regarding the high success rate in the reconstruction of distorted speech as opposed to speaker identification in general, for which the effectiveness was not numerically quantified. At the same time, it faithfully presents the reservations about the clarity of the method and the validity of its use in the courtroom. Similar to Poland, there was just one department of criminalistic phonoscopy in the whole of Czechoslovakia and, although I found no evidence of open criticism of the method similar to that formulated by Legień, the initial plan to establish other regional labs was never realized, owing likely to a moderate disappointment regarding the ability of the method to clearly identify unknown voices (Kvicalova, 2023b: 391). Unlike the initial enthusiasm regarding the biometric value of the – soon to be discredited – voiceprint methodology in the 1960s, the strengths of conclusions to be expected from phonoscopic voice dissection was much weaker.
Although conclusions concerning speaker identity were initially expressed in binary categorical terms, deciding whether the voice on the recording belonged or did not belong to the suspect, this practice was gradually abandoned in favor of nuanced and more realistic probability scales. The identity of a speaker was thus no longer expressed in a binary ‘yes/no’ framework but by verbal expression of likelihood. Whereas in the UK the binary conclusion framework was used until the end of the 1980s (French, 2017), in Czechoslovakia the probability scale was adopted earlier in that decade. In 1979, Błasikiewicz and Bednarczyk claimed – too enthusiastically, as their critics would put it – that if a comparative sound recording was available, it was possible to make a categorical identification (1979: 714). In the course of the 1980s, however, the notion that the speaker’s identity was either confirmed or ruled out by the analysis evolved into a scale ranging from ‘no’ to ‘it cannot be ruled out’, ‘it might be possible’, ‘probably’, ‘with high probability’, and ‘with the highest probability’ (Málek and Musilová, 1989: 76). The probability formulations found in the Czechoslovak expert reviews thus closely resemble those used in the UK or Germany in the early 1990s. 28 Although it has been assumed that the move from categorical conclusions to probabilistic statements was prompted by experience with DNA methods (Cole, 2002: 290), sound-based forensics was clearly at the forefront of this more general shift in the expression of expert opinion before court. In the scales used by audio forensics, likelihood was expressed verbally instead of numerically because – unlike in DNA analysis – no population statistics were available for individual components of speech. The adoption of probability conclusions together with the development of detailed methodology for analytical listening was an original attempt by East-Central European forensic phonoscopy to reformulate objectivity standards in the courtroom to account for elements of manageable uncertainty. Similar concerns regarding the admissibility and presentation of expertise in legal trials were shared by forensic experts on the western side of the Iron Curtain, whose efforts to incorporate probability statements into their expert reviews were also prompted by investigations of sound. As shown in this article, however, both Polish and Czechoslovak attempts at drawing up standards of objectivity for sound-based evidence in forensics were not modeled on foreign examples or simply derived from Western or Soviet sources, but they were rooted in a relatively independent and original forensic tradition.
Conclusion
Forensic science in East-Central Europe was firmly planted in international criminalistic networks since the 19th century. Although institutional and personal ties were suspended in the Cold War era, the transfer of expertise was never completely interrupted: technologies, methods, and theories continued to travel across ideological borders, although their movement was now restricted. The new branch of sound-based forensics, which was called phonoscopy behind the Iron Curtain, followed from earlier investigations of the human voice using the spectrograph as they were developed for military purposes during WWII. A detailed investigation of the practices of the Polish phonoscopy lab – the first ever police audio forensic department, founded in 1963 to identify unknown voices in sound recordings and to develop practical means of sound dissection – reveals an original approach to sound analysis that went beyond the research of ‘voiceprint’ as it was conducted in the US.
In establishing phonoscopic expertise in Poland and Czechoslovakia, technological equipment from the West was instrumental, although in both countries original devices for speech analysis were developed by phoneticians and specialists in electroacoustics. Most frequently, however, sound-based expertise travelled in the form specialized literature – American acoustic research was well studied in the East, and British, French, and German literature was also cited. The reception of Western theoretical research on voice identification and automated speech recognition behind the Iron Curtain was rather nuanced: phonoscopy experts did not simply accept or use its results but subjected it to well-founded criticism in the light of the practical demands of forensic work. Although some of the apparatuses required for phonoscopic dissection, such as cepstral analyzers, were designed and constructed in the West, experts from Polish and Czechoslovak forensic police departments – somewhat paradoxically – sometimes had easier access to semi-classified technology than their Western counterparts, who worked as independent experts with limited access to government-controlled technologies.
Polish and Czechoslovak sound-based expertise rivaled and even surpassed that in the West in formulating complex methodological approaches to forensic voice analysis, which was based on the dissection of hundreds of cases submitted to the labs by the Communist police and other state institutions. The phonoscopy lab in Warsaw developed detailed instructions regarding the proper use of precision technological equipment and application of up-to-date phonetic, acoustic, and linguistic knowledge, along with practical guidelines for trained group listening, all of which together was meant to guarantee the objectivity of research.
The question of the objectivity and accuracy of phonoscopic analysis was most pressing in the judicial context, where tape recordings had to be translated into the category of legal proof. In this respect, both Polish and Czechoslovak audio forensics led the way in reformulating the notion of objectivity in the courtroom: although relying on subjective hearing in the analysis, practitioners found a way of incorporating the uncertainty element in the expression of their expert reviews. The verbal probability scales introduced in Czechoslovakia in the second half of the 1980s resembled similar efforts by Western forensic experts, who had recourse to probability claims at the turn of the decade, yet they stemmed from a relatively autonomous forensic tradition whose considerations of sound-based objectivity went back to 1960s Polish forensic practice.
As far as knowledge transfer across the Cold War divide is concerned, the study shows that – contrary to mainstream opinion – Soviet influence on the development of forensic expertise in the Eastern Bloc countries was rather limited. The international criminalistic symposia held behind the Iron Curtain did not generate any intense cooperation between the USSR and East-Central European countries, which would be based on actual knowledge exchange between the Moscow lab and audio forensic police departments in Warsaw and Prague. This complements recent findings about the scope of cooperation between the East German Stasi and the KGB, which was marked by long-term reluctance of the Russians to share detailed information about their research on criminalistic voice identification (Bijsterveld, 2021: 220). Such evidence suggests a more nuanced notion of the Cold War bipolar division than traditionally assumed: even in the areas of expertise pertaining to state security and surveillance matters, the Eastern Block was not simply defined by Soviet influence. As far as audio forensics is concerned, both Czechoslovak and Polish sources analyzed in this article testify to the existence of a largely independent East-Central European strand of applied research that both refined existing acoustic knowledge as formulated in foreign literature and brought about original theoretical insights and novel technological solutions.
This article not only shows a greater – albeit mostly indirect – influence of Western over Soviet expertise in the development of sound-based criminalistics behind the Iron Curtain, but, more importantly, it challenges the monolithic notion of the Eastern Bloc as a sphere of Soviet influence. Although the Soviet Union exercised significant political and ideological control over the region, as for the history of expert knowledge, the reality was far more complex depending on the type of expertise concerned. The evolution of audio forensics in East-Central Europe shows that the region’s history of applied research – and sciences and humanities more broadly – is not best understood simply by locating it within the Cold War polarity: the development of criminalistics was neither preconditioned by Soviet science and state security nor simply derived from the reading of Western literature; thus, it needs to be interpreted in its own terms.
Footnotes
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The writing of this article was funded by the Czech Science Foundation under the research project ‘Sound as Evidence: Listening in the Field and the Laboratory in 20th-Century Czech Science’ (25-18562S).
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
