MS, Cut Delays in Home Appliance AI Conversations... LG Unveils Collaboration on ‘ThinQ On’
TECHWORLD ·
✦ AI Summary
MS and LG Electronics collaborated to improve the conversation quality of AI home voice agents.
The two companies linked speech recognition, response generation, and voice delivery to reduce latency, and made the system stop the existing response and respond when a new request came in.
The collaboration was presented in the manufacturing session of the 'Microsoft Industry Summit' held on September 30 at the Grand Ballroom of COEX in Seoul.
MS and LG Electronics collaborated to improve the conversation quality of AI home voice agents. The two companies linked user speech recognition, response generation, and voice delivery to reduce latency, and when a new request was entered during a conversation, they implemented a system that stopped the existing response and then responded to the new request.
Microsoft Korea held the 'Microsoft Industry Summit' on September 30 at the Grand Ballroom of COEX in Seoul. At the event, the company shared AI use cases in finance, retail, manufacturing, healthcare, and education, and discussed ways to apply AI and agents to work, customer service, research and development, and production sites.
The collaboration was presented as a case study in the manufacturing session at the event. Kim Cheol-ha, team leader at LG Electronics' HS Platform Business Center, and Seong Min-ji, technology strategy manager at Microsoft Korea, introduced the development process for applying Azure Voice Live, speech-to-text (STT), and text-to-speech (TTS) to the voice agent of the smart home hub 'ThinQ On,' and related demo videos were also released.
The collaboration made reducing latency at each stage of voice processing its core task. In the existing structure, large language model (LLM) response generation began after speech recognition was completed, and text-to-speech began after response generation was finished. As a result of this sequential structure, delays accumulated.
To reduce this, Voice Live was equipped with 'overlapped streaming' technology. The overlapped streaming method delivers intermediate processing results to the next stage. Voice Live starts inference even while listening to the user's speech.
In addition, generated responses are output sequentially as speech. Manager Seong explained that the key to overlapped streaming is starting the next stage as soon as partial results are ready. She added that it is a structure that starts inference while listening and begins speaking while inference is still under way.
Service-by-service calls were also consolidated into a single WebSocket session. The purpose of unifying WebSocket sessions is to reduce network round-trip latency.
The criteria for deciding when to respond were also improved. If the end of speech is judged based only on silence time, there is a problem in that a user who pauses briefly while thinking may be seen as having stopped speaking, and there was also the issue of responses being delayed even after a request had ended.
By applying 'semantic VAD' and adjusting the response start timing and the way conversation history is reflected, the two companies improved the naturalness of voice interaction and the handling of interruptions. 'Semantic VAD' is a method that reflects sentence completeness and conversation context.
Accordingly, the system responds when a command is complete, and waits if there is a possibility of follow-up speech, such as conjunctions or enumerations. If the user starts speaking during an answer, voice output is stopped and the new request is processed, while any unplayed audio is discarded.
Only what the user has actually heard is reflected in the conversation history. This approach takes into account situations in which the user starts speaking again before the response has finished playing to the end.
When there is a delay in checking or controlling home appliance status, a short guide is provided first. LG Electronics implemented this so that the LLM generates phrases such as "I'll check" that fit the context, allowing users to recognize progress even while a task is underway.
Regarding this, Manager Seong said the actual tool execution time remains the same. However, she explained that the approach reduces perceived waiting time by outputting the first voice immediately.
It also implemented recognition from product names to LG's distinctive voice. In speech recognition, it supplemented product and feature names that general-purpose models often miss, a measure taken in consideration of the difficulty of accurately processing requests when LG Electronics' proprietary terms are recognized as everyday words.
In speech recognition development, the company used both 'phrase list,' which can reflect specific terms without separate training, and 'custom speech,' which specializes training for language and acoustic models. It also reflected repetitive home appliance control commands and real acoustic environments, and established a system to continuously improve quality based on real-world usage data.
For multilingual expansion, the company used product and feature names accumulated in the ThinQ language pack as development resources. Based on words and sentences extracted from that material, the agent generated sentences for training, translated the generated sentences, and used them to train language-specific speech recognition models.
It said linking existing language assets with the training data generation process also helped speed up development. Team leader Kim said that, compared with the company's other divisions' 10 years of development experience in 29 languages, this time it improved speech recognition performance in 16 languages over 1 year.
Text-to-speech was built by combining Azure Custom HD and LG Electronics' AI persona. After conducting test recordings with voice actor candidates, the company selected a voice through internal evaluation, then trained the selected voice's tone, manner of speaking, and emotions to implement context-appropriate intonation.
LG Electronics improved both speech accuracy and output speed by adding complementary measures before and after speech synthesis. It added text-conversion preprocessing for proper pronunciation of numbers and abbreviations, and also added preprocessing that assigns language tags to mixed-language sentences. It also applied a method of storing and reusing already synthesized audio, which improved output speed.
On top of this preprocessing and reuse structure, Voice Live was configured to link speech recognition, LLM, and text-to-speech, and to connect the existing LG AI agent through tool calls. Through this, the company utilized existing home appliance control and service functions, and continued home appliance control and service execution through voice conversations with the aim of improving voice conversation quality.
In the public demonstration, the system showed continuous execution of control based on prior conversation. When a user asked about the temperature suitable for a child's sleep and then requested that the bedroom air conditioner be set to that temperature, the system processed it according to the context of the voice conversation, and it also handled air purifier status checks, lighting control, and story generation by voice.
Team leader Kim said the filming date of the demo video was September 4, and that response speed was improved afterward and the music playback function was also linked. Future tasks include reducing the influence of surrounding environments and developing stable voice transmission and reception technology like human-to-human conversation.
Team leader Kim said the conditions for a good agent include high understanding, an interaction style that feels like talking with a person, and the embodiment of a brand's voice. He then explained the direction of enriching the voice experience by working with Microsoft and targeting Voice Live, STT, and TTS for further advancement.
Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407593
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.