The task sounded simple: let a shopper dictate a search query out loud from any browser and any device. The voice turns into text in the search field, while individual words act as commands — "search", "cancel" — and trigger the matching actions on the site. Below we take such a system apart from the inside: from the microphone in the browser to the recognition service's response and back.
The task
We needed to build speech recognition on the web page of an online store. The shopper presses the record button, dictates a query — and sees it as text in the search field. Part of what is said is handled as commands: "search" runs the search, "cancel" clears the input.
The key requirement was full compatibility with the most common browsers: Chrome, Firefox, Safari, Opera and Edge. In other words, the solution cannot rely on capabilities that only one of them has.
What the system has to do
Three things are needed to make this work:
- capture the audio stream from the browser;
- stream the audio data to a speech recognition service;
- receive results in real time — while the user is still speaking.
Step 1. How to capture the audio stream from the browser
The getUserMedia / Stream API helps here, and most popular browsers support it. Through this browser interface we get the audio stream from the microphone and send it to our speech recognition service.

Next we initialise the audio context and call getUserMedia. Although this configuration is compatible with the main browsers, the parameters passed to the API may differ depending on your needs — check the details against the official documentation.
Configuring the audio context
The configuration of the audio context and the script processor also depends on the task. The tricky part is that some combinations of parameters currently do not work in Safari — that is worth verifying separately.
Pay particular attention to subscribing to the "audioprocess" event: for a stream coming from the microphone it effectively plays the role of an "ondata" event.
Sound is recorded in stereo, so two audio channels come from the microphone. In practice we take only one of them, because the recognition service needs single-channel audio. But depending on your needs you can do whatever you like with the sound.
What to do when recording stops
- Turn off the microphone. Stop using the browser's microphone and reset any variables you have.
- Detach the audio process event listener. Without this the stream keeps firing the event even after the microphone has been deactivated.
Step 2. Where to send the audio
The simplest route is to use the speech recognition API built into the browser. With it a web application recognises speech as a stream right in the browser, without any third-party services.
But, as with every interesting new feature, there are small compatibility problems. And it is exactly those that make this approach entirely unsuitable for production applications.

For a good user experience we need a real-time recognition system that returns results while the user is speaking, rather than recording the sound in full and only then transcribing it. So we need a third-party service with the following properties:
Streaming recognition
Results have to arrive continuously, as the person speaks, and not in a single chunk once the recording is finished.
Compatibility with every browser
The service must not depend on what a particular user's browser can or cannot do.
Why we chose the Google Speech API
There are plenty of services online that let applications recognise speech through an API. We settled on the Google Speech API: it provides a high-quality streaming recognition service that is genuinely efficient, and a real-time response was critical for us.
The SDK library offers both an asynchronous recognition service (via gRPC and the REST API) and a real-time recognition service — the latter only through the gRPC API.
Our goal was real-time recognition, so we chose streaming recognition. Google effectively provides this API only through a gRPC call, using the streaming feature of the gRPC protocol.
Architecture: browser, server, recognition service
This is precisely why audio cannot be sent from the browser over gRPC directly. The browser can indeed make gRPC calls, but not streaming ones.
So we took another route: stream the data to a server, and from there into Google's streaming recognition API, and back again to collect the recognition result.
There are not many options for streaming audio data from the browser to a server. In practice the only systems capable of streaming data out of the browser are the webRTC protocol and WebSocket. For simplicity we chose WebSocket. It makes sense to build the server side on nodeJS: stream handling there is simple and websockets are very straightforward to implement.

- Connection. After the record button is pressed we connect to the WebSocket — we want to avoid a long-lived connection that could slow the server down — and add a listener for the audio stream data.
- Listening. We listen for events on the browser's microphone.
- Sending. We send the data buffer to the backend over the WebSocket.
- Passing it on. We forward the data from the backend to Google's speech recognition service.
- Receiving. We collect the recognition results.
- Returning to the frontend. We send the results to the frontend over the websocket so they can be displayed as they arrive.
Implementing the websocket in the browser
The important thing to remember here is that the audio data stream is a Float32Array, and it has to be turned into a buffer and then unpacked again on the server.
We convert the Float32Array into an Int16Array and only then send it to the server. Without this conversion the server will not receive the data in the required format. The ConversionFactor variable exists purely for conversion purposes.
Implementing the websocket on the server
The server-side implementation is standard; there are only two painful points:
- exactly when to start recognition through Google;
- how to handle the difference between the format of the stream data arriving from the websocket and the audio stream format expected by the speech recognition stream.
First you need to create a speech recognition stream — the details are in the service's documentation.
After that you can stream the audio chunks arriving from the browser to the speech API. And this is where the second painful point surfaces — data formats.
| Stage | Data format | What we do |
|---|---|---|
| Microphone in the browser | Float32Array, stereo | Take only one channel |
| Before sending to the websocket | Int16Array | Convert from Float32Array |
| On the server from the websocket | Standard buffer, 8-bit integers | Accept as is |
| Before writing to the recognition stream | Int16Array, "LINEAR16" encoding | Convert the buffer |
The buffer arriving from the websocket is a standard buffer, an array of 8-bit integers. In the speech recognition settings, however, we specified that the audio encoding is "LINEAR16", that is, an array of 16-bit integers. So the standard buffer has to be turned into an Int16Array before it is written into the recognition stream.
The recognition stream returned by Google's speech recognition library is an ordinary standard stream, and it can be handled in the usual way.
The result
What we end up with is a cross-browser speech recognition system that makes interacting with the site simpler. The shopper no longer has to type the product name by hand — saying it out loud is enough.
If you are planning something similar in your own project, start with two checks: support for the getUserMedia / Stream API in the browsers you need, and the capabilities of the speech-to-text service you choose. These two things determine how complex the rest of the work will be.
