SenateOfCanada
KoreaSteno
VivaVoceReporting
ad

IPRS Meetings & Diamesic 9 Conference

55th Intersteno Congress, July 25th – 30th, 2026

IPRS Meetings & Diamesic 9 Conference

Diamesic 9 Conference – Monday, July 27th, 2026

The Case for Shorthand in the Age of AI

Andrew Hill – Financial Times

Generative AI can record, transcribe, and interpret speech with sufficient speed and accuracy to deter new learners of pen-shorthand and render the “winged art” redundant. What will we lose if we allow shorthand to become a dead language, consigned to unread archives and practiced only by enthusiasts and eccentrics? Andrew Hill sees three important reasons for nurturing and developing the skill before it is too late. Shorthand is a key to our past, it is fuel for our brains, and it is a means of securing a future beyond the screen-bound digital gloom of our current age.

Shorthand is deeply embedded in our past, for instance as a language to record parliamentary debates. Understanding snippets of shorthand could prove to be a key to unlocking the words around moments that shaped history through digitization and transcription of these archival records. Shorthand also might find its place as an analogue writing tool in an increasingly digital world, for the same reason people who do not want to get lost in and do not want to rely on the cloud switch back to listening vinyl records and rediscover analogue cameras.

Academic research suggests learning and mastering shorthand has several cognitive benefits, from strengthening short term memory to even slowing down brain aging. Another beacon of hope, familiar to people mastering the skill, is the practice of shorthand for the sheer fun of it. The lively shorthand community on Reddit proves there is a future for shorthand as a hobby.

Andrew Hill will further investigate the past, present and future of shorthand in his upcoming book, Take Note, to be published in November 2026.

The Significance of Shorthand in the Early Finnish Society

Jari Niittuinperä – Finnish Shorthand Association

Jari Niittuinperä presented the historical importance of shorthand in Finland and its contribution to the development of Finnish society, public administration, and parliamentary democracy. His presentation focused on the introduction of shorthand in the nineteenth century, the organization of stenographic offices, the role of women, and the lasting value of verbatim parliamentary records.

Shorthand was introduced in Finland in the 1850s, when government officials recognized the need for accurate parliamentary reporting. Influential figures such as August Schauman promoted the idea after studying shorthand systems abroad, while several young Finns were sent to Leipzig to receive professional training. Among the first trained stenographers were Johan Edward Swan, Wasilii Margunoff, and Svante Dalström, who later played important roles in the Finnish Parliament and public administration. German shorthand teacher Karl Albrecht also contributed by educating the first Finnish stenographers.

The presentation explained how shorthand became connected with the growing status of the Finnish language. During the nineteenth century, legislation increasingly required official documents and parliamentary minutes to be available in Finnish as well as Swedish. Publishing verbatim parliamentary records in both languages supported the development of Finnish as an official language and improved public access to political debates.

Niittuinperä also discussed the position of women in the profession. Hilma Aminoff became the first female parliamentary stenographer during the 1877–1878 Diet, followed by several other women in later years. Although women gradually entered the profession, leadership positions within the stenographic offices and the Finnish Shorthand Association remained almost exclusively occupied by men.

In conclusion, the speaker emphasized that verbatim parliamentary records remain an invaluable historical source. While modern technology has changed the way speeches are recorded, written minutes continue to provide researchers with one of the fastest and most reliable ways to understand how political discussions and decisions develop over time.

The Note-Taking System of Consecutive Interpreters – a Multilingual Shorthand System?
A matter of multilingual speech capturing

Boris Neubauer – Aachen University of Applied Science

Speech capturing is not only used and developed in contexts where Intersteno is active, but also in a sector such as that of consecutive interpreters: translators who translate what has been spoken into another language not simultaneously, but retrospectively. Unlike simultaneous translators, who rely primarily on their memory when translating spoken text into another language, consecutive translators sometimes use a notation system to later correctly convert spoken texts into another language.

Consequently, examples can be found of consecutive translators who have developed their own notation systems. This development appears to have originated in the period after the Second World War, when texts were translated from French to English and other languages. The method of listening to sound recordings and translating based on them, shifted towards a method of figurative notation.

These notation systems are combinations of symbols, which have the meaning of words, and figurative signs (arrows, lines) intended to indicate the connections between words. An advantage of this type of system is that it makes it possible to reproduce extensive texts using simple notation. Due to the graphical representation, this notation is also not dependent on language. A major disadvantage is that the number of symbols the reporter must memorize can be large.

How does this graphical system compare to existing shorthand systems? A comparison reveals striking differences. Unlike existing shorthand systems, the graphical systems are not intended for verbatim reproduction of the spoken word. Furthermore, the abbreviations in the graphical system are based on words in longhand, not on stenographic characters. If no symbol is known for a spoken word, it is written in longhand. The systems are very inefficient because they are not based on an alphabet. And finally: they are individual systems, which limits their use.

Reporting without Barriers: Philippine Court Reporting in a Global Context

Christian Dabeh T. Clerigo – PhilSteno

In his presentation, Christian Dabeh Clerigo examines how the Philippines had a low barrier to assimilate writing systems but high barriers of challenges in converting the spoken word to transcript efficiently.

The presentation was started by showing a timeline of Philippine writing systems. Before English and Filipino became the two dominant languages in the 20th century, Kawi (9th–10th century), Baybayin (14th–16th century) and Latin/Castillan (16th–19th century) were the common written languages. Nowadays, more than 175 dialects are spoken across the archipelago. Filipino and English are the official languages of the government, but English is the accepted language of the official court record, which is used in nearly all proceedings. This means that witnesses often testify in a regional Filipino dialect. They are interpreted into English by a court interpreter and transcribed into English by the court stenographer. This means that words pass through two different people’s translations before they become the official written record. As a result, the meaning of what someone actually said sometimes changes.

Philippine stenography is used as a profession by court stenographers, legislative officers, and non-court/executive stenographers. In the courtroom, different technologies and innovations lead from Gregg Shorthand and Takigrapiya in the past to contemporary early trials with AI-assisted transcription. This is however still in the pilot stage.

7.000+ stenographers are currently serving the Philippine judiciary nationwide. In 2019 the A-to-Z Program was introduced to start re-training machine stenotyping with computer-aided transcription software. In this program, in partnership with the U.S. National Court Reporters Association, an introduction to a career as stenographer is given in six sessions. The program promotes awareness, develops future talent, and supports educational institutions to adopt modern stenography.

Dabeh Clerigo ended his presentation by naming the current barriers to a reliable record. The earlier mentioned language diversity may lead to losing accuracy, while the caseload leaves stenographers under time pressure and a nation of islands like the Philippines creates challenges in delivering training, technology, and opportunities.

Crafting Craftsmanship: Selecting and Training Parliamentary Reporters in the Netherlands

Laura van den Boogaard and Erik PJ Vrinds – House of Representatives of the Netherlands

Producing parliamentary proceedings is a highly specialized job that requires excellent language and editing skills as well as political sensitivity and knowledge of legislation. In the absence of a specific degree in parliamentary reporting in the Netherlands, the selection process for parliamentary reporters at the House of Representatives is carefully designed to identify candidates with the potential to develop the level of craftsmanship necessary to produce high-quality reports.

The selection procedure at the Parliamentary Reporting Department involves screening application letters, an assessment, and a job interview. The assessment tests candidates’ general and political knowledge, language proficiency and editing skills, meanwhile giving applicants a clearer picture of the demands of the job.

New parliamentary reporters start with a year of intensive in-house training, during which they already contribute to the official edited verbatim reports of plenary sittings and committee meetings. They work under the supervision of experienced co-workers, who give them feedback on every text they produce. As their accuracy and speed improve and their confidence grows, the length of the audio fragments to be processed in a set time is increased.

In-house sessions to further their understanding of the roles and procedures of the parliament complement this one-on-one training and help them to make well-informed editing choices. One’s progress is monitored through tutor reports, self-evaluation and progress meetings between the reporter in training, the tutors, and the line manager.

At the end of the presentation topics for discussion are proposed, including challenges such as labor shortages, the impact of artificial intelligence and technological innovation, and changing expectations regarding long careers and lengthy training programs.

Exploration and Practice of Stenography Education and Talent Cultivation in China

Biao Xu – Chinese Information Processing Society of China

Over years of exploration and practice, the Chinese judiciary has gradually established a distinctive system for education, integrating stenography skills with the training of court clerk talents. The aim is to cultivate high-quality, application-oriented technical and skilled legal talents who meet the needs of judicial system reform. Even in the era of rapidly advancing AI, stenography continues to play an irreplaceable role in court proceedings, government meetings, media communications, and business negotiations.

The growing demand for professionals continues to drive the improvement and upgrading of China’s stenography educational system. Through a school-court cooperation scheme, an integrated teaching model combining post requirements, course competitions, and certification, the curriculum aligns with legal work processes. With a multi-party integrated teaching evaluation mechanism, and the integration of legal education and moral cultivation, current issues such as the disconnect between talent cultivation objectives and market demands, as well as mismatched curriculum standards, can be addressed.

At the same time, new models of human and AI collaboration should be explored. According to Biao Xu, this is the defining mission of stenography education in China. Stenography competitions help shape an innovative mechanism for the cultivation of highly qualified professionals, integrating job requirements with the curriculum. The curriculum ensures alignment between education and professional requirements.

Bridging Continents: Bringing Stenographic Standards to Emerging Markets

Rachel Harris – Sopherim & Associates LLC

Rachel Harris, a stenographic educator and legal documentation advocate, explained what is needed to achieve stenographic excellence in regions where formal systems are lacking and stenographic infrastructure is limited or non-existent. Harris discussed this topic using a practical example: the case of Nigeria.

Nigeria’s courts and legal institutions needed reliable, real-time records, but the certification, curriculum, and mentorship infrastructure to train stenographers to a professional standard did not exist locally. The approach was to use a proven American stenography training standard and adapt it to Nigeria’s legal context, language landscape, and educational infrastructure. The aim was to get a standard imported: the American stenographic curriculum and accuracy benchmarks.

Harris explained to the audience what the four key challenges are to adapt to a rigorous standard without compromising it. First, she mentioned cross-cultural implementation. This involves introducing an unfamiliar profession within existing legal, linguistic, and educational norms. The second challenge is educator development by building a local core of trainers who can sustain the standard after the founders step back. The third challenge is to maintain accuracy standards, holding the line on speed and precision benchmarks as the program scales. The last key challenge is the local context adaptation. In the case of Nigeria that meant fitting delivery to its realities such as access, language, and infrastructure without compromising quality.

Harris explained that an educational program only outlives its founders if local trainers can carry it forward. First, they get trained to have the same accuracy benchmarks as the future students. They should be trained to be capable, certified trainers. Then they practice teaching by giving supervised instruction and getting feedback from senior mentors. Before leading a classroom independently, they get a formal sign-off and are certified to teach. Their students could subsequently form a new generation of legal stenographers and enter the field.

The program creates a repeatable pipeline for identifying, training and certifying new stenographers. It provides access to accurate legal records. Courts gain reliable, real-time transcription, strengthening due process and the rule of law. This system can also create international partnerships and provide cross-border collaboration that advances the profession well beyond one country. Harris ended her presentation by mentioning that stenography grows strongest when it crosses borders – carried by educators, protected by standards, and shared through partnership.

Analysis and Comparison between Automatic Subtitles and the Ones Made by Subtitlers in Italian Language on Mobile Phones and Computers: Updating with AI Technologies, Quality and Contest

Giacomo Pirelli – onA.I.R. Intersteno Italia

Giacomo Pirelli works at the University of Turin and is the deputy president of onA.I.R. Intersteno Italia, an Italian association and working group for respeaking and subtitling, both real-time and pre-recorded. The organization focuses on accessibility for deaf and hard-of-hearing people. The topic of his presentation is automatic live transcripts/automatic speech recognition (ASR) and AI based solutions in everyday life.

He is prelingual deaf and uses a cochlear implant in daily life and is aware of the difficulties deaf people face every day, in school or at work. In school for instance, he had to rely on other people making notes of what was said. Later, during the Covid 19 pandemic, because of people wearing safety masks, he could not follow any livestreams.

When using his smartphone for a call or making a videocall with his laptop, Giacomo Pirelli must depend on apps for automatic live transcripts, which are not easy to operate. These apps provide live subtitles or live captions to enable interaction via phone or laptop. The transcript in general is fast and accurate, with few errors, but there is still room for improvement. There is the issue of delayed punctuation. Also, sometimes the software does not recognize the voice of Giacomo Pirelli, because his spoken Italian differs from that of a hearing person. When this is the case, bad translation can lead to misunderstanding.

It is Giacomo Pirelli’s firm belief that in last 5 years ASR has been making huge progress in accuracy. This is crucial in making informal contexts accessible for the deaf and hard of hearing. However, on a more practical note, he points out that poor pronunciation, crosstalk, and an unreliable internet connection can ruin your day.

Designing Accessible Speech-to-Text Content: Insights from User-Centered Research with Older Adults with Mild Cognitive Impairment

Paula Hernandez Burguete – Cardenal Herrera University

Older adults with mild cognitive impairments struggle to process information. The question arises as to whether this also applies to information presented via media channels. Digitalized content relies heavily on sensory input-specifically sound and visuals. Little was previously known about the cognitive accessibility of media-delivered information. This prompted Ms. Paula Burguete to conduct research to gain greater insight into the matter.

The study’s overarching goal was to analyze the accessibility of media information for older adults with mild cognitive impairments. Specific research questions addressed the barriers to understanding information, the strategies employed to comprehend it, and the perspectives of care professionals, the older adults and their families, and subject-matter experts. The hypotheses were that the main barriers lie not in the availability of information but in its comprehensibility, that simple language and a structured offering improve the accessibility of information, and that the three categories of participants would largely agree on barriers and solutions.

The findings indicate that the delivery of digital information is too rapid and unstructured for older adults with cognitive impairments. Furthermore, the language used is overly complex. The study participants require assistance from others – such as family members and care professionals – to process the information provided. Failure to understand the information leads to psychosocial consequences, including frustration and withdrawal from the information source.

This is particularly unfortunate, given that mental stimulation slows the decline of cognitive processes in older adults. Maintaining such stimulation is crucial. Consequently, the study’s recommendations are highly relevant: information should be presented with the user at the center, using simple language and straightforward audiovisual templates.

Diamesic 9 Conference – Tuesday, July 28th, 2026

Interpreting Chaos: Observations on STAARs Performance in Complex Speech Scenarios

Ana Luísa Reis – Portuguese Parliament

STAAR is the Whisper-based AI supported speech-to-text solution used in the Portuguese Parliament. Read more about STAAR in issue 2/2025 of Tiro. STAAR has repeatedly proved its high performance when a single speaker contributes clear and structured speech, relieving parliamentary reporters from painstakingly transcribing speech to text and enabling them to put more effort in delivering a perfect edited text. How does STAAR behave in complex speech scenarios, especially in the case of overlapping speakers? And what does it take for the parliamentary reporters to work with this output?

STAARs performance declines in quality when multiple speakers intervene, in the case of interruptions, heated discussions and overlapping remarks, which is common in parliamentary debates. Factors such as spontaneous reactions, side comments, multiple speakers talking simultaneously or background noise can significantly reduce the quality of transcription. Error patterns are consistent but unpredictable, with missing or inaccurate words or phrases, merged interventions, or incorrect automated attribution of speakers.

STAARs behavior under nonstandard and non ideal conditions demands a particular skill of the human in the loop, the parliamentary reporter. A new source of error and misinterpretation is introduced when the automatically generated transcript is approached without critical thinking and professionalism. Reporters have to train new listening and interpretation skills to address this bias, with more emphasis on interpretation of parliamentary situations, editing, and proofreading. With its sometimes hilarious mistakes, STAAR has even proven to be a valuable tool for de-stressing through comic relief.

The Effect of Computer-Aided Stenography in Korea

Yeonhwan Choi – Korea Steno

Despite the brief history of Korea since 1948, stenography has offered various textualization services in Korea these days. In 1994, the first steno-machine using computer programming released for Korean by Korea Steno. From minutes in parliament to subtitles in video, computer aided stenography is widely used as valid technology and about 20,000 stenographers work for several services.

The normal way of working involves a stenographer attending a meeting and taking notes and comments and working out minutes later, in the office, with the help of the audio files. This procedure has two major drawbacks. First, a Seoul-based stenographer must be physically present at the meeting and the time to correct requires a lot of time. Korea Steno has therefor introduced the solution of Smart-I and Steno Editor which enable working with ASR without having the need of a business trip and furthermore play audio files simultaneously, which speeds up the entire process.

One of the biggest plans Korea Steno has for the future is to close caption live TV. Developments are aiming for more widely using ASR technology there and using the stenographers more efficiently: We can see the ambition of going from 4 to 2 to 1 staff member working at the same time while keeping up a high accuracy rate of 98%. This reduction is still very much a work in progress which requires further discussions on a social level as the impact on the working process is significant: Stenographers wouldn’t be required to be at the working place anymore and also the tolerance range for mistakes in TV live caption should be included in that discussion.

How Do We Transcribe Dialect Speech in Parliamentary Recordings?

Tatsuya Kawahara – Kyoto University

Whereas “standard” speech is usually expected in parliamentary meetings, MPs occasionally utter dialect speech. This is often intentional, to show regional identity and emphasis, or for rhetorical effect. Professor Kawahara investigates how parliamentary reporters in the National Parliament (Diet) and local assemblies of Japan transcribe dialect speech in their official records. To this end, he conducted interviews with stenographers of the National Parliament and reporters who work on the records of meetings of local assemblies.

Professor Kawahara identifies three aspects of dialect: variation in pronunciation (e.g., /k/ turns into /g/), differences in vocabulary (e.g., “subway”, “underground” or “tube”), and morphological variation (e.g., “I ain’t”). Transcribing this dialect speech always requires balancing accuracy and readability.

In the Japanese National Parliament, pronunciation variations are normalized in the written record. Morphological variations are also changed into formal styles. Dialect vocabulary, however, is transcribed as it is, except for minor corrections. The main guideline dictates that dialect words are never replaced by other words, even if many people will not understand them.

Local assemblies have more permissive or even actively inclusive policies towards dialect. Some regional guidelines are more detailed than others, but the way in which this dialect speech is transcribed mostly depends on the person in charge of the records. In general, pronunciation variations are normalized and the vocabulary is kept as it is. The most significant difference with reporting at the National Parlement is that morphological variations are not corrected in local assemblies.

Ten years ago, representatives in local assemblies often used dialect speech to demonstrate locality and express emotions. However, the use of dialects has declined over the past decade; it is hardly observed anymore. Professor Kawahara suggests that this might be a side effect of video streaming or recording.

Improving Dutch ASR Transcripts through Rule-Based Post-Processing

Max van Winden – House of Representatives of the Netherlands

Max van Winden, a parliamentary reporter, presented an experiment using rule-based post-processing to improve automatic speech recognition transcripts. The aim was not to generate a finished parliamentary report automatically, but to remove predictable friction before a reporter begins editing. This includes automatically correcting recognition errors and features of spontaneous speech and implementing editorial conventions of the reporting office.

The experiment distinguished between corrections that can be applied safely and changes that require human interpretation. Fixed spellings, compound words, capitalization, and number notation usually have one preferred form and can therefore be automated. Repetitions, sentences structures and other changes that may affect emphasis or political meaning require more attention. The guiding principle was to automate stable forms, flag uncertain patterns and leave interpretation to the reporter.

Using Python, the system combined Word autocorrect entries, a glossary, parliamentary style guide rules, and a reference dictionary. All changes were recorded in a log, making the process transparent and testable. This also revealed the risks of over-broad rules. Fuzzy matching, for example, changed the Dutch word for “confidence figures” into “confidence crises”. The words were similar in spelling, but entirely different in meaning.

To evaluate their system, 47 minutes of a parliamentary debate was transcribed, recording 281 automatic changes. Compared to the official edited report, the word error rate decreased from 28.95% to 26.76%. Relative to the original error rate, which is a reduction of about 7.6%. Although numerical improvement was modest, the corrected transcript was easier to follow. Van Winden therefore proposes a layered approach in which safe corrections are automated, uncertain cases are presented as suggestions and the reporter remains responsible for correctly reflecting meaning, attribution, emphasis, and political nuance.

Training and Work of Judicial Stenographers – Practice in China

Chen Yang – Beijing College of Politics and Law

Professor Yang presented the education and professional practice of judicial stenographers in China and explained how the profession is adapting to the increasing use of artificial intelligence. The presentation highlighted the essential role of judicial clerks, the educational model used to prepare them, and the future collaboration between human professionals and AI.

In China, judicial stenographers are officially known as judicial clerks and are indispensable members of the court system. Besides recording court proceedings, they manage case files, prepare legal documents, organize hearings, and assist judges throughout the judicial process. Due to the heavy workload of Chinese courts, where some district courts handle more than 100,000 cases each year, judicial clerks have become key contributors to the efficient functioning of the legal system.

To prepare students for these responsibilities, Beijing College of Politics and Law applies a “Three-in-One” educational model. This approach combines targeted recruitment, comprehensive skills training, and extensive internships. Students receive instruction in legal procedure, judicial secretarial work, courtroom shorthand, and modern stenography software. Equal attention is given to professionalism, confidentiality, integrity, and ethical responsibility, ensuring that graduates possess both technical competence and strong professional values.

The college also makes extensive use of technology during training. A virtual court simulation system allows students to practice realistic courtroom situations before entering the workplace. Long internships, beginning during the second year of study, help students gain practical experience and make the transition to full-time employment much smoother.

Finally, Professor Chen discussed the growing impact of artificial intelligence. Rather than replacing judicial stenographers, AI-supported speech recognition and transcription systems are viewed as tools that improve efficiency and accuracy. According to the speaker, future professionals will need not only strong shorthand and legal knowledge but also the ability to work alongside AI by monitoring, correcting, and validating automated transcripts. Human judgment, responsibility, and ethical decision-making will therefore remain essential to ensuring fair and reliable judicial proceedings.

What Is the Role of Humans when Using AI in Professional Reporting?

Eero Voutilainen – Reporting Office of the Parliament of Finland

Eero Voutilainen, chief senior specialist, examined how automatic speech recognition has changed professional reporting at the Finnish Parliament. ASR has been used to produce draft reports since 2019. A new large language model introduced in 2025 also performs limited pre-editing based on earlier reports. Drawing on observations and discussions with colleagues through a semi-structured interview, Voutilainen considered how these developments have affected reporters’ work, roles and agency.

The introduction of ASR has changed the workflow from a two-stage process to a single-stage process. Previously, document secretaries produced the initial draft, which was then edited by senior specialists. Both groups now edit ASR-generated text. This reduces the time required to produce a report from around 90 minutes to 30 minutes. Document secretaries have moved from typing to editing, while senior specialists still perform essentially the same editorial task.

ASR nevertheless requires different forms of attention. It produces unfamiliar and sometimes deceptive changes, while editorial suggestions may over-edit, omit information or leave irregular structures unchanged. This has introduced some elements of post editing in the process, meaning that editors must supervise and review some pre-editorial suggestions by AI. In the editing stage, more technical and procedural tasks have also been concentrated. Reporters must manage different ways of incorporating listening, editing and proofreading with regards to working with draft transcripts.

Voutilainen argues that ASR-based reporting should still be regarded as a computer-aided rather than human-aided process. The system provides preliminary textual material, but trained professionals interpret, contextualize, and edit it according to office guidelines. They control the workflow and remain responsible for the final report. The human professional should therefore not merely be “in the loop” but “on top of the loop”.

AI-Supported Parliamentary Reporting: There’s More to AI than Meets the Eye

Katrien Van Mulders – Flemish Parliament

The Flemish Parliament has recently integrated AI into their working processes in a pilot form, but not yet in their redacting tool. It combines speech to text technology with generative AI. In technical terms it consists of three parts: 1. Pyannote for speaker recognition, 2. Whisper for transcription and 3. generative AI, the LLM called ChatGPT for postprocessing. The customizable prompt used for ChatGPT includes general instructions and specific editorial rules. In a pilot conducted in 2026, 85% of the reporters used the tool and 60% of the 5-minute takes were completed with the tools. That number is significantly lower because the AI tool is useless when, for instance, the voting in the meeting starts. The results show that about 20% of the text generated by ChatGPT still must be changed, but overall, the result with is better than without.

Working with AI provides a number of challenges. For instance, Pyannote is not yet integrated in the editing tool. Also, misrecognitions, misinterpretations or deletions of AI which are not recognized or corrected by reporters can endure in the process because of the AI trust paradox. AI is getting good at imitating human language, producing fluent text. It can give the impression of quality and make errors harder to find. However, the choices AI makes are sometimes puzzling. Why would it change “shitshow” in “chaos” or “resting bitch face” in “neutral facial expression”? ChatGPT also sometimes can come up blank or simply refuse to produce a new text.

What are the lessons here about making a parliamentary report? The audio comes first, the context comes second, the draft third, always followed by revision. AI can draft the report, but humans must guarantee the record. The promise of AI is real, but human judgement is indispensable.

Hybrid Transcription Systems and Parliamentary Reporting: An Evaluation Study

Fabio Angeloni and Paolo Antonio Michela Zucco – Senate of the Republic of Italy

Fabio Angeloni and Paolo Antonio Michela Zucco discussed so called hybrid forms of speech transcription in the field of parliamentary reporting. The term “hybrid” is used for using a combination of different technologies working together to achieve better results than any single technology can achieve on its own. Goals are higher accuracy, better quality, and greater efficiency. Over time, the reporting sector has seen the use of various tools and technologies. Common ways of reporting went from manual shorthand to mechanical stenography, analog and digital recording, speaker-dependent speech recognition system to digital stenography, and widely used automatic speech recognition (ASR).

In 2020 stenographic vendors in the U.S. started developing systems in which steno and voice no longer compete with each other but operate in an integrated manner. The first integration attempt involved using ASR in the editing phase. The software compared the stenographic transcription with that produced by the speech engine, highlighting any inconsistencies, missing words, or translation conflicts.

In 2022, in what the presenters call the “second phase”, the next development ensured that ASR was no longer only used for reviewing stenotype, but also did the real-time comparison. The stenotype output was instantly compared with the ASR text and AI-routines suggested corrections in real-time, based on so-called confidence scores for words or phrases.

The third phase, starting in 2023, meant that the speech recordings were automatically reported through an AI/ASR-combination under human supervision. That may be called steno-assisted ASR-transcription. This phase can be defined as “human-on-the-loop”, as the operator supervises and edits the ASR/AI-generated text in real time. The editor can confirm, modify, validate, or add any missing words with the stenotype keyboard.

In their experimental study, Angeloni and Zucco set ASR on “always-on”. They used a light AI-mode, so automatic correction is limited to missing words, untranslated terms, names, and conflicts. AI-intervention was only activated on request of the operator. AI suggestions were set off for the absence of untranslated segments or conflicts. Their goal was to provide AI-assistance while preserving full human control over the transcription process. Operators can now choose between two different working modes: ASR-assisted stenography and the steno-assisted AI/ASR-combination. The key challenge is that successful hybrid transcription depends not only on AI, but also on operator expertise and system quality. This involves non-negligible costs, but delivers more reliability, efficiency, and confidence. In the future, further applications could be subtitling, accessibility or the use of this system in other transcription areas.

1887 - 2025 All Rights Reserved. Intersteno - International Federation for Information and Communication Processing