Fast-Tracking Speech Recognition

The Open Speech Initiative Begins

Bruce Byfield

OSI seeks to bring advanced speech processing to free software.

Over the years, free software has seen at least a dozen projects to implement speech recognition. However, even the most advanced of these projects, such as CMU Sphinx and Festival, often lag behind commercial equivalents, largely because of a lack of resources. To close this gap, Peter Grasch, a KDE developer from Austria, launched the Open Speech Initiative (OSI) in October 2013 with the goal of assembling “a team of developers looking to bring first class speech processing to the world of free software.”

This is an ambitious project, but Grasch argues strongly for its importance. “Speech recognition,” he says, “is a greatly underutilized input method. Over the recent years we can at last see it slowly being adopted in mobile applications, and I think this trend will surely continue. I am not saying it can or will replace the other input methods we have now, but there are many use cases where speech input can simplify things. At least, after the first Iron Man movie came out, I think many would agree,” he adds, referring to the extensive speech processing capabilities that the movie’s protagonist has embedded throughout his home and office as well as in his combat suit.

Given the growing importance of speech recognition, Grasch describes as “troubling” the fact that “speech technology was and to a large extent still is entirely in the domain of big business.” Not only are the three main commercial developers – Nuance, Microsoft, and Google – proprietary, but so are the implementations for Android and Tizen.

According to Grasch, the reason for this situation is that “speech recognition requires significant upfront investment to acquire the necessary data, and tedious speech modeling that does not generalize well across languages – requiring countless more hours if you intend to support multiple languages.” The result is a combination of resources, expertise, and effort that is extremely difficult to organize and sustain in a volunteer project.

However, Grasch suggests that such obstacles are beginning to become less important because of crowd-sourcing and the gradual accumulation of existing data. Now, he suggests, “with community engagement, we can build more accurate speech recognition systems and create better integrated solutions for more devices – and the use cases are truly endless.”

Since 2010, Grasch has been explaining this rationale at conferences, including KDE’s annual Akademy and the Desktop Linux Summit, as well as writing about it in academic papers. Now, with the OSI, he plans to put the rationale into practice.

The Path to OSI

Grasch became interested in free software while still in high school. “I was always a tech enthusiast and had been following the Linux movement with a bit of interest for some time when, in the tenth grade, I managed to throughly wipe my Windows installation from my home computer,” he recalls. “At that point, I had never even tried a Live CD, but, for some reason, I decided to just install SUSE 9.3 instead of going back to XP. Since then, I have not owned a single Windows system.”

The transition was not always smooth, but in getting his system up and running, Grasch discovered that he enjoyed both hacking and the community he discovered in the Linux forums. “While I had already dabbled in writing small programs back on Windows, I never realized that merely writing code is just the start of it; I wanted to become part of the free software community and give back,” he says.

His opportunity to get involved came about a year later, in a class in which students worked on projects suggested by professionals. When Franz Stieger, a special needs teacher, wanted to study the best speech recognition software for children with speech impediments, Grasch and three other students volunteered.

Grasch remembers, “We quickly realized that there was no commercial offering that was flexible enough to cope with non-standard speech patterns. So we drafted the concept of Simon, an extremely flexible open speech recognition system and set to work. To us, it was always clear that the result would be free software.”

As Grasch went on to university, he continued the development of Simon as a KDE project. Building on CMU Sphinx, Julius and HTK, Simon is designed for both Linux and Windows.

At the 2013 Akademy, the annual meeting of KDE developers, Grasch delivered a talk in which he described the progress he was able to make on one challenge in speech processing in one week, using only free software data and technologies.

“The point,” Grasch said, “was to show off how close – or, indeed, far away – we were from being able to implement current and next gen speech recognition using free software. The experiment was a big success and proved that by even just investing a handful of days, it was possible to further the state of the art in open source speech recognition.”

Grasch continues, “A large part of the reason for conducting such an experiment was to show interested third parties that investing time in free speech technology is viable. Quite a few enthusiasts and even companies responded and showed interest. As a response, we set up the Open Speech Initiative both to give our newly formed team a common label and to formalize the cooperation.”

Setting Priorities

The OSI is still in its earliest stages of development – so early that the website is still under construction, and the project lists only half a dozen members.

“From a consumer perspective, there is little to see at this point,” Grasch admits. “But behind the largely desolate end user software landscape for anything but command and control applications, there are some promising efforts. Grasch singles out CMU Sphinx for “mature speech recognition engines for a variety of use cases” and the KALDI toolkit, which he describes as “working on state of the art neural-network-based decoding.

However, the main challenge is to develop an efficient speech model, which is necessary to teach the recognition software what each language sounds like. “Creating such speech models requires careful and painfully time-consuming planning and tons and tons of data,” Grasch explains. “Because of this, even projects using open source speech recognition engines mostly rely on proprietary speech models. This is why creating high quality open source speech models is on the forefront of the Open Speech Initiative’s agenda.”

Given this background, the immediate goals of the OSI are well-defined, particularly for Grasch’s continued work on Simon. As described on Grasch’s blog, much of this work is a necessary prelude to providing speech recognition – such as reviewing existing resources and working to improve acoustic and language models, as well as dictation capabilities. The ideal is “to create systems that aid with this and automate the process as much as possible,” according to Grasch.

Eventually, the OSI intends to create the infrastructure to allow contributions from non-experts – perhaps even endusers. Then, Grasch himself plans to add the missing features in Simon to produce a preview version that may be ready as early as this summer.

Even when these goals are reached, Grasch sees “a large, continuous effort” ahead, with no immediate end in sight. However, with the coordination of effort and the successful attracting of casual contributors, the OSI may represent the best chance for free software speech processing to catch the proprietary leaders in the field – and perhaps even show them a thing or two.

Support Our Work

Linux Magazine content is made possible with support from readers like you. Please consider contributing when you’ve found an article to be beneficial.

News

Linux Kernel 6.17 Drops bcachefs

Filesystem , Kernel , Linux

After a clash over some late fixes and disagreements between bcachefs's lead developer and Linus Torvalds, bachefs is out.
ONLYOFFICE v9 Embraces AI

Artificial Inte... , open source , OpenOffice

Like nearly all office suites on the market (except LibreOffice), ONLYOFFICE has decided to go the AI route.
Two Local Privilege Escalation Flaws Discovered in Linux

Kernel , Linux , Security

Qualys researchers have discovered two local privilege escalation vulnerabilities that allow hackers to gain root privileges on major Linux distributions.
New TUXEDO InfinityBook Pro Powered by AMD Ryzen AI 300

Hardware , Linux , Notebook

The TUXEDO InfinityBook Pro 14 Gen10 offers serious power that is ready for your business, development, or entertainment needs.
Danish Ministry of Digital Affairs Transitions to Linux

LibreOffice , Linux , Windows

Another major organization has decided to kick Microsoft Windows and Office to the curb in favor of Linux.
Linux Mint 20 Reaches EOL

With Linux Mint 20 at its end of life, the time has arrived to upgrade to Linux Mint 22.
TuxCare Announces Support for AlmaLinux 9.2

AlmaLinux , Enterprise Linux , Security

Thanks to TuxCare, AlmaLinux 9.2 (and soon version 9.6) now enjoys years of ongoing patching and compliance.
Go-Based Botnet Attacking IoT Devices

IoT , Security , Systemd

Using an SSH credential brute-force attack, the Go-based PumaBot is exploiting IoT devices everywhere.
Plasma 6.5 Promises Better Memory Optimization

Desktop , Linux , Plasma

With the stable Plasma 6.4 on the horizon, KDE has a few new tricks up its sleeve for Plasma 6.5.
KaOS 2025.05 Officially Qt5 Free

KDE , Linux , Operating Systems

If you're a fan of independent Linux distributions, the team behind KaOS is proud to announce the latest iteration that includes kernel 6.14 and KDE's Plasma 6.3.5.

Fast-Tracking Speech Recognition

The Open Speech Initiative Begins

Related content

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters

Support Our Work

News

Linux Kernel 6.17 Drops bcachefs

ONLYOFFICE v9 Embraces AI

Two Local Privilege Escalation Flaws Discovered in Linux

New TUXEDO InfinityBook Pro Powered by AMD Ryzen AI 300

Danish Ministry of Digital Affairs Transitions to Linux

Linux Mint 20 Reaches EOL

TuxCare Announces Support for AlmaLinux 9.2

Go-Based Botnet Attacking IoT Devices

Plasma 6.5 Promises Better Memory Optimization

KaOS 2025.05 Officially Qt5 Free

Fast-Tracking Speech Recognition

The Open Speech Initiative Begins

Related content

Subscribe to our Linux Newsletters Find Linux and Open Source Jobs Subscribe to our ADMIN Newsletters

Support Our Work

News

Tag Cloud

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters