Why Us Services Portfolio Blog Technologies Development process Start a Project →
UA EN RU
← All posts
Insight · 20.05.2019 · 9 min read

Voice Control in E-commerce: Cross-Browser Speech Recognition for Site Search

How to let shoppers dictate a search query in any browser: capturing audio with getUserMedia, streaming it over WebSocket to the Google Speech API, and returning results in real time.
Voice Control in E-commerce: Cross-Browser Speech Recognition for Site Search

The task sounded simple: let a shopper dictate a search query out loud from any browser and any device. The voice turns into text in the search field, while individual words act as commands — "search", "cancel" — and trigger the matching actions on the site. Below we take such a system apart from the inside: from the microphone in the browser to the recognition service's response and back.

The task

We needed to build speech recognition on the web page of an online store. The shopper presses the record button, dictates a query — and sees it as text in the search field. Part of what is said is handled as commands: "search" runs the search, "cancel" clears the input.

The key requirement was full compatibility with the most common browsers: Chrome, Firefox, Safari, Opera and Edge. In other words, the solution cannot rely on capabilities that only one of them has.

What the system has to do

Three things are needed to make this work:

  • capture the audio stream from the browser;
  • stream the audio data to a speech recognition service;
  • receive results in real time — while the user is still speaking.
5
browsers it has to work in: Chrome, Firefox, Safari, Opera, Edge
1
audio channel out of the stereo pair — the recognition service needs no more
65 s
limit on audio length in a single recognition request

Step 1. How to capture the audio stream from the browser

The getUserMedia / Stream API helps here, and most popular browsers support it. Through this browser interface we get the audio stream from the microphone and send it to our speech recognition service.

Support for the getUserMedia / Stream API in popular browsers
Browser support for the getUserMedia / Stream API

Next we initialise the audio context and call getUserMedia. Although this configuration is compatible with the main browsers, the parameters passed to the API may differ depending on your needs — check the details against the official documentation.

Configuring the audio context

The configuration of the audio context and the script processor also depends on the task. The tricky part is that some combinations of parameters currently do not work in Safari — that is worth verifying separately.

Pay particular attention to subscribing to the "audioprocess" event: for a stream coming from the microphone it effectively plays the role of an "ondata" event.

Sound is recorded in stereo, so two audio channels come from the microphone. In practice we take only one of them, because the recognition service needs single-channel audio. But depending on your needs you can do whatever you like with the sound.

Format matters. The data arriving from the audio stream is a typed Float32Array. That means it has to be converted before being sent to the recognition service.

What to do when recording stops

  1. Turn off the microphone. Stop using the browser's microphone and reset any variables you have.
  2. Detach the audio process event listener. Without this the stream keeps firing the event even after the microphone has been deactivated.

Step 2. Where to send the audio

The simplest route is to use the speech recognition API built into the browser. With it a web application recognises speech as a stream right in the browser, without any third-party services.

But, as with every interesting new feature, there are small compatibility problems. And it is exactly those that make this approach entirely unsuitable for production applications.

Compatibility of the browser speech recognition API across different browsers
Browser compatibility of the built-in speech recognition API

For a good user experience we need a real-time recognition system that returns results while the user is speaking, rather than recording the sound in full and only then transcribing it. So we need a third-party service with the following properties:

Streaming recognition

Results have to arrive continuously, as the person speaks, and not in a single chunk once the recording is finished.

Compatibility with every browser

The service must not depend on what a particular user's browser can or cannot do.

Why we chose the Google Speech API

There are plenty of services online that let applications recognise speech through an API. We settled on the Google Speech API: it provides a high-quality streaming recognition service that is genuinely efficient, and a real-time response was critical for us.

The SDK library offers both an asynchronous recognition service (via gRPC and the REST API) and a real-time recognition service — the latter only through the gRPC API.

Our goal was real-time recognition, so we chose streaming recognition. Google effectively provides this API only through a gRPC call, using the streaming feature of the gRPC protocol.

Architecture: browser, server, recognition service

This is precisely why audio cannot be sent from the browser over gRPC directly. The browser can indeed make gRPC calls, but not streaming ones.

So we took another route: stream the data to a server, and from there into Google's streaming recognition API, and back again to collect the recognition result.

There are not many options for streaming audio data from the browser to a server. In practice the only systems capable of streaming data out of the browser are the webRTC protocol and WebSocket. For simplicity we chose WebSocket. It makes sense to build the server side on nodeJS: stream handling there is simple and websockets are very straightforward to implement.

Diagram of audio being passed from the browser through the server to the speech recognition service
The final infrastructure of the solution
  1. Connection. After the record button is pressed we connect to the WebSocket — we want to avoid a long-lived connection that could slow the server down — and add a listener for the audio stream data.
  2. Listening. We listen for events on the browser's microphone.
  3. Sending. We send the data buffer to the backend over the WebSocket.
  4. Passing it on. We forward the data from the backend to Google's speech recognition service.
  5. Receiving. We collect the recognition results.
  6. Returning to the frontend. We send the results to the frontend over the websocket so they can be displayed as they arrive.

Implementing the websocket in the browser

The important thing to remember here is that the audio data stream is a Float32Array, and it has to be turned into a buffer and then unpacked again on the server.

We convert the Float32Array into an Int16Array and only then send it to the server. Without this conversion the server will not receive the data in the required format. The ConversionFactor variable exists purely for conversion purposes.

Implementing the websocket on the server

The server-side implementation is standard; there are only two painful points:

  • exactly when to start recognition through Google;
  • how to handle the difference between the format of the stream data arriving from the websocket and the audio stream format expected by the speech recognition stream.

First you need to create a speech recognition stream — the details are in the service's documentation.

The 65-second limit. As soon as the speech stream is created, the recognition request is already running, and Google caps the length of audio for recognition at 65 seconds. That is why recognition must be created at exactly the moment the user presses the record button on the web page and the server receives the browser's request to connect to the websocket.

After that you can stream the audio chunks arriving from the browser to the speech API. And this is where the second painful point surfaces — data formats.

StageData formatWhat we do
Microphone in the browserFloat32Array, stereoTake only one channel
Before sending to the websocketInt16ArrayConvert from Float32Array
On the server from the websocketStandard buffer, 8-bit integersAccept as is
Before writing to the recognition streamInt16Array, "LINEAR16" encodingConvert the buffer

The buffer arriving from the websocket is a standard buffer, an array of 8-bit integers. In the speech recognition settings, however, we specified that the audio encoding is "LINEAR16", that is, an array of 16-bit integers. So the standard buffer has to be turned into an Int16Array before it is written into the recognition stream.

The recognition stream returned by Google's speech recognition library is an ordinary standard stream, and it can be handled in the usual way.

Do not forget to close the stream. It is very important to close the recognition stream data when the websocket closes. That is how we avoid a Google speech recognition error caused by streaming audio for too long.

The result

What we end up with is a cross-browser speech recognition system that makes interacting with the site simpler. The shopper no longer has to type the product name by hand — saying it out loud is enough.

If you are planning something similar in your own project, start with two checks: support for the getUserMedia / Stream API in the browsers you need, and the capabilities of the speech-to-text service you choose. These two things determine how complex the rest of the work will be.