Interpreting Chaos: Observations on STAARs Performance in Complex Speech Scenarios
Ana Luísa Reis – Portuguese
Parliament STAAR is the Whisper-based AI supported speech-to-text solution used in the Portuguese Parliament. Read more about STAAR in issue 2/2025 of Tiro. STAAR has repeatedly proved its high performance when a single speaker contributes clear and structured speech, relieving parliamentary reporters from painstakingly transcribing speech to text and enabling them to put more effort in delivering a perfect edited text. How does STAAR behave in complex speech scenarios, especially in the case of overlapping speakers? And what does it take for the parliamentary reporters to work with this output? STAARs performance declines in quality when multiple speakers intervene, in the case of interruptions, heated discussions and overlapping remarks, which is common in parliamentary debates. Factors such as spontaneous reactions, side comments, multiple speakers talking simultaneously or background noise can significantly reduce the quality of transcription. Error patterns are consistent but unpredictable, with missing or inaccurate words or phrases, merged interventions, or incorrect automated attribution of speakers. STAARs behavior under nonstandard and non ideal conditions demands a particular skill of the human in the loop, the parliamentary reporter. A new source of error and misinterpretation is introduced when the automatically generated transcript is approached without critical thinking and professionalism. Reporters have to train new listening and interpretation skills to address this bias, with more emphasis on interpretation of parliamentary situations, editing, and proofreading. With its sometimes hilarious mistakes, STAAR has even proven to be a valuable tool for de-stressing through comic relief.
The Effect of Computer-Aided Stenography in Korea
Yeonhwan Choi – Korea Steno
Despite the brief history of Korea since 1948, stenography has offered various textualization services in Korea these days. In 1994, the first steno-machine using computer programming released for Korean by Korea Steno. From minutes in parliament to subtitles in video, computer aided stenography is widely used as valid technology and about 20,000 stenographers work for several services. The normal way of working involves a stenographer attending a meeting and taking notes and comments and working out minutes later, in the office, with the help of the audio files. This procedure has two major drawbacks. First, a Seoul-based stenographer must be physically present at the meeting and the time to correct requires a lot of time. Korea Steno has therefor introduced the solution of Smart-I and Steno Editor which enable working with ASR without having the need of a business trip and furthermore play audio files simultaneously, which speeds up the entire process. One of the biggest plans Korea Steno has for the future is to close caption live TV. Developments are aiming for more widely using ASR technology there and using the stenographers more efficiently: We can see the ambition of going from 4 to 2 to 1 staff member working at the same time while keeping up a high accuracy rate of 98%. This reduction is still very much a work in progress which requires further discussions on a social level as the impact on the working process is significant: Stenographers wouldn’t be required to be at the working place anymore and also the tolerance range for mistakes in TV live caption should be included in that discussion.
How Do We Transcribe Dialect Speech in Parliamentary Recordings?
Tatsuya Kawahara – Kyoto University
Whereas “standard” speech is usually expected in parliamentary meetings, MPs occasionally utter dialect speech. This is often intentional, to show regional identity and emphasis, or for rhetorical effect. Professor Kawahara investigates how parliamentary reporters in the National Parliament (Diet) and local assemblies of Japan transcribe dialect speech in their official records. To this end, he conducted interviews with stenographers of the National Parliament and reporters who work on the records of meetings of local assemblies. Professor Kawahara identifies three aspects of dialect: variation in pronunciation (e.g., /k/ turns into /g/), differences in vocabulary (e.g., “subway”, “underground” or “tube”), and morphological variation (e.g., “I ain’t”). Transcribing this dialect speech always requires balancing accuracy and readability. In the Japanese National Parliament, pronunciation variations are normalized in the written record. Morphological variations are also changed into formal styles. Dialect vocabulary, however, is transcribed as it is, except for minor corrections. The main guideline dictates that dialect words are never replaced by other words, even if many people will not understand them. Local assemblies have more permissive or even actively inclusive policies towards dialect. Some regional guidelines are more detailed than others, but the way in which this dialect speech is transcribed mostly depends on the person in charge of the records. In general, pronunciation variations are normalized and the vocabulary is kept as it is. The most significant difference with reporting at the National Parlement is that morphological variations are not corrected in local assemblies. Ten years ago, representatives in local assemblies often used dialect speech to demonstrate locality and express emotions. However, the use of dialects has declined over the past decade; it is hardly observed anymore. Professor Kawahara suggests that this might be a side effect of video streaming or recording.
Improving Dutch ASR Transcripts through Rule-Based Post-Processing
Max van Winden – House of Representatives of the Netherlands
Max van Winden, a parliamentary reporter, presented an experiment using rule-based post-processing to improve automatic speech recognition transcripts. The aim was not to generate a finished parliamentary report automatically, but to remove predictable friction before a reporter begins editing. This includes automatically correcting recognition errors and features of spontaneous speech and implementing editorial conventions of the reporting office. The experiment distinguished between corrections that can be applied safely and changes that require human interpretation. Fixed spellings, compound words, capitalization, and number notation usually have one preferred form and can therefore be automated. Repetitions, sentences structures and other changes that may affect emphasis or political meaning require more attention. The guiding principle was to automate stable forms, flag uncertain patterns and leave interpretation to the reporter. Using Python, the system combined Word autocorrect entries, a glossary, parliamentary style guide rules, and a reference dictionary. All changes were recorded in a log, making the process transparent and testable. This also revealed the risks of over-broad rules. Fuzzy matching, for example, changed the Dutch word for “confidence figures” into “confidence crises”. The words were similar in spelling, but entirely different in meaning. To evaluate their system, 47 minutes of a parliamentary debate was transcribed, recording 281 automatic changes. Compared to the official edited report, the word error rate decreased from 28.95% to 26.76%. Relative to the original error rate, which is a reduction of about 7.6%. Although numerical improvement was modest, the corrected transcript was easier to follow. Van Winden therefore proposes a layered approach in which safe corrections are automated, uncertain cases are presented as suggestions and the reporter remains responsible for correctly reflecting meaning, attribution, emphasis, and political nuance.
Training and Work of Judicial Stenographers – Practice in China
Chen Yang – Beijing College of Politics and Law
Professor Yang presented the education and professional practice of judicial stenographers in China and explained how the profession is adapting to the increasing use of artificial intelligence. The presentation highlighted the essential role of judicial clerks, the educational model used to prepare them, and the future collaboration between human professionals and AI. In China, judicial stenographers are officially known as judicial clerks and are indispensable members of the court system. Besides recording court proceedings, they manage case files, prepare legal documents, organize hearings, and assist judges throughout the judicial process. Due to the heavy workload of Chinese courts, where some district courts handle more than 100,000 cases each year, judicial clerks have become key contributors to the efficient functioning of the legal system. To prepare students for these responsibilities, Beijing College of Politics and Law applies a “Three-in-One” educational model. This approach combines targeted recruitment, comprehensive skills training, and extensive internships. Students receive instruction in legal procedure, judicial secretarial work, courtroom shorthand, and modern stenography software. Equal attention is given to professionalism, confidentiality, integrity, and ethical responsibility, ensuring that graduates possess both technical competence and strong professional values. The college also makes extensive use of technology during training. A virtual court simulation system allows students to practice realistic courtroom situations before entering the workplace. Long internships, beginning during the second year of study, help students gain practical experience and make the transition to full-time employment much smoother. Finally, Professor Chen discussed the growing impact of artificial intelligence. Rather than replacing judicial stenographers, AI-supported speech recognition and transcription systems are viewed as tools that improve efficiency and accuracy. According to the speaker, future professionals will need not only strong shorthand and legal knowledge but also the ability to work alongside AI by monitoring, correcting, and validating automated transcripts. Human judgment, responsibility, and ethical decision-making will therefore remain essential to ensuring fair and reliable judicial proceedings.
What Is the Role of Humans when Using AI in Professional Reporting?
Eero Voutilainen – Reporting Office of the Parliament of Finland
Eero Voutilainen, chief senior specialist, examined how automatic speech recognition has changed professional reporting at the Finnish Parliament. ASR has been used to produce draft reports since 2019. A new large language model introduced in 2025 also performs limited pre-editing based on earlier reports. Drawing on observations and discussions with colleagues through a semi-structured interview, Voutilainen considered how these developments have affected reporters’ work, roles and agency. The introduction of ASR has changed the workflow from a two-stage process to a single-stage process. Previously, document secretaries produced the initial draft, which was then edited by senior specialists. Both groups now edit ASR-generated text. This reduces the time required to produce a report from around 90 minutes to 30 minutes. Document secretaries have moved from typing to editing, while senior specialists still perform essentially the same editorial task. ASR nevertheless requires different forms of attention. It produces unfamiliar and sometimes deceptive changes, while editorial suggestions may over-edit, omit information or leave irregular structures unchanged. This has introduced some elements of post editing in the process, meaning that editors must supervise and review some pre-editorial suggestions by AI. In the editing stage, more technical and procedural tasks have also been concentrated. Reporters must manage different ways of incorporating listening, editing and proofreading with regards to working with draft transcripts. Voutilainen argues that ASR-based reporting should still be regarded as a computer-aided rather than human-aided process. The system provides preliminary textual material, but trained professionals interpret, contextualize, and edit it according to office guidelines. They control the workflow and remain responsible for the final report. The human professional should therefore not merely be “in the loop” but “on top of the loop”.
AI-Supported Parliamentary Reporting: There’s More to AI than Meets the Eye
Katrien Van Mulders – Flemish Parliament
The Flemish Parliament has recently integrated AI into their working processes in a pilot form, but not yet in their redacting tool. It combines speech to text technology with generative AI. In technical terms it consists of three parts: 1. Pyannote for speaker recognition, 2. Whisper for transcription and 3. generative AI, the LLM called ChatGPT for postprocessing. The customizable prompt used for ChatGPT includes general instructions and specific editorial rules. In a pilot conducted in 2026, 85% of the reporters used the tool and 60% of the 5-minute takes were completed with the tools. That number is significantly lower because the AI tool is useless when, for instance, the voting in the meeting starts. The results show that about 20% of the text generated by ChatGPT still must be changed, but overall, the result with is better than without. Working with AI provides a number of challenges. For instance, Pyannote is not yet integrated in the editing tool. Also, misrecognitions, misinterpretations or deletions of AI which are not recognized or corrected by reporters can endure in the process because of the AI trust paradox. AI is getting good at imitating human language, producing fluent text. It can give the impression of quality and make errors harder to find. However, the choices AI makes are sometimes puzzling. Why would it change “shitshow” in “chaos” or “resting bitch face” in “neutral facial expression”? ChatGPT also sometimes can come up blank or simply refuse to produce a new text. What are the lessons here about making a parliamentary report? The audio comes first, the context comes second, the draft third, always followed by revision. AI can draft the report, but humans must guarantee the record. The promise of AI is real, but human judgement is indispensable.
Hybrid Transcription Systems and Parliamentary Reporting: An Evaluation Study
Fabio Angeloni and Paolo Antonio Michela Zucco – Senate of the Republic of Italy
Fabio Angeloni and Paolo Antonio Michela Zucco discussed so called hybrid forms of speech transcription in the field of parliamentary reporting. The term “hybrid” is used for using a combination of different technologies working together to achieve better results than any single technology can achieve on its own. Goals are higher accuracy, better quality, and greater efficiency. Over time, the reporting sector has seen the use of various tools and technologies. Common ways of reporting went from manual shorthand to mechanical stenography, analog and digital recording, speaker-dependent speech recognition system to, digital stenography, and widely used automatic speech recognition (ASR). In 2020 stenographic vendors in the U.S. started developing systems in which steno and voice no longer compete with each other but operate in an integrated manner. The first integration attempt involved using ASR in the editing phase. The software compared the stenographic transcription with that produced by the speech engine, highlighting any inconsistencies, missing words, or translation conflicts. In 2022, in what the presenters call the “second phase”, the next development ensured that ASR was no longer only used for reviewing stenotype, but also did the real-time comparison. The stenotype output was instantly compared with the ASR text and AI-routines suggested corrections in real-time, based on so-called confidence scores for words or phrases. The third phase, starting in 2023, meant that the speech recordings were automatically reported through an AI/ASR-combination under human supervision. That may be called steno-assisted ASR-transcription. This phase can be defined as “human-on-the-loop”, as the operator supervises and edits the ASR/AI-generated text in real time. The editor can confirm, modify, validate, or add any missing words with the stenotype keyboard. In their experimental study, Angeloni and Zucco set ASR on “always-on”. They used a light AI-mode, so automatic correction is limited to missing words, untranslated terms, names, and conflicts. AI-intervention was only activated on request of the operator. AI suggestions were set off for the absence of untranslated segments or conflicts. Their goal was to provide AI-assistance while preserving full human control over the transcription process. Operators can now choose between two different working modes: ASR-assisted stenography and the steno-assisted AI/ASR-combination. The key challenge is that successful hybrid transcription depends not only on AI, but also on operator expertise and system quality. This involves non-negligible costs, but delivers more reliability, efficiency, and confidence. In the future, further applications could be subtitling, accessibility or the use of this system in other transcription areas.