DeepSeek Voice Chat Gray Test Adds Four Selectable TTS Voices

Date:

DeepSeek appears to be testing spoken replies inside its consumer app. A small speaker control in the upper-right corner of the chat surface lets some users have assistant responses read aloud. A new settings menu reportedly offers four selectable text-to-speech voices. The experiment remains a gray test. The feature is staged, unannounced, and separate from DeepSeek’s September model and API news.

This article explains what the DeepSeek voice chat test includes and who can use it. It also covers the four voices, compares the feature with voice modes from OpenAI, Google, and xAI, and lists what DeepSeek has not confirmed.

What the DeepSeek Voice Chat Test Actually Shows

The test itself dates to September 12, 2026. English-language summaries appeared two days later, and they agree on the basics. Users selected for the experiment see a speaker button in the upper-right corner of the chat window. Tapping it plays a spoken version of the assistant’s written reply without leaving the conversation thread.

The voice picker lives inside the app’s settings page. Testers can choose between four built-in voice profiles there. Beyond that, public details thin out quickly.

DeepSeek has published no product blog post, no release note, and no marketing page for the feature as of mid-September 2026 coverage.

Early coverage disagrees on one key point.

Some reports describe the test as a full two-way voice chat. In that version, users speak to DeepSeek and hear it speak back. Others are more cautious. They frame it as voice playback of written responses for a limited group of users. A safer reading is that spoken output is confirmed for testers. Continuous, hands-free conversation remains unverified.

The Four Voices: Beike, Bailang, Haixing, and Anchao

The most concrete detail to emerge from the test is the voice lineup. Chinese storefront and settings copy lists four profiles, and English summaries use their pinyin names because no official English voice IDs have surfaced in DeepSeek materials.

Voice Chinese Reported description
Beike 贝壳 Versatile and lively, also rendered as lively and changeable
Bailang 白浪 Bright and firm
Haixing 海星 Playful and sweet
Anchao 暗潮 Deep and resonant

The names translate roughly to “seashell,” “white wave,” “starfish,” and “undercurrent.” That fits a wider pattern: Chinese AI assistants often name voices after natural imagery instead of numbering them. For context, rival assistants typically ship larger voice libraries: OpenAI’s ChatGPT voice mode has offered a roster of around nine selectable voices, and Google’s Gemini Live includes multiple voice options of its own. Four voices is a modest start. Still, it shows DeepSeek treats voice as a product surface with character, not just as a basic accessibility toggle.

Four colorful abstract soundwaves representing DeepSeek's four selectable text-to-speech voice profiles
Concept visualization of four selectable text-to-speech voice profiles (Credit: Intelligent Living)

What a Gray Test Means for Availability

“Gray test” is a direct translation of 灰度测试 (huīdù cèshì), the standard Chinese software term for a gray-scale or staged rollout. In practice, it means a feature is pushed to a small, server-selected subset of users before any general availability announcement. Testers are not told why they were chosen, and the feature can be modified or withdrawn without notice.

Grid of smartphone icons with a small glowing subset representing users selected for a staged app test
A gray test reaches a small, server-selected subset of users first (Credit: Intelligent Living)

That framing matters for the flood of autocomplete searches around this story, such as “deepseek voice chat app,” “deepseek voice chat android,” and “deepseek voice chat iphone.” There is no setting to enable, no beta program to join, and no confirmed platform breakdown. If your app does not show a speaker button in the chat view, you are not in the test group.

Public reports do not specify which platforms the test covers. iOS, Android, or both all remain unconfirmed. The same goes for regional limits. Until DeepSeek publishes formal release notes, the speaker button should be treated as a staged user experience experiment, not a contractual feature.

Searchers asking “deepseek voice chat free” also lack a definitive answer. The consumer app is free to download, and core chat has been free. DeepSeek has said nothing about whether voice features will carry a charge.

A Client Feature, Not Another Model Launch

September 2026 has been a busy month for DeepSeek headlines, and it is easy to conflate them. The table below separates what is confirmed about the model stack from what the voice test is and is not.

September DeepSeek item What it is Surface affected
V4.1-Flash release New generally available model on the API, called via the model name deepseek-flash, with native multimodal support API and model serving
Retirement of older aliases deepseek-v4-flash and deepseek-v4-flash-vision-exp names now route to V4.1-Flash for compatibility API billing and routing
Announced Pro-to-Flash cutover A plan to route deepseek-v4-pro traffic to V4.1-Flash economics from 04:00 UTC on September 14, 2026 API pricing
Reversal of that cutover DeepSeek’s own API changelog states that, in response to user demand, V4 Pro API service continues after September 14, 2026, with billing unchanged API pricing
Voice chat gray test A staged in-app experiment adding a speaker control and four TTS voice profiles Consumer app client

That fourth row is worth dwelling on.

Several early summaries described the Pro-to-Flash routing as an operational fact, but DeepSeek’s official API changelog records a change of course after user pushback. Separately, the company’s V4.1-Flash announcement frames Flash as the leader on performance and cost. For now, V4 Pro service continues. The voice test sits one layer above all of this billing churn. It changes how people hear the assistant, not which checkpoint answers them.

DeepSeek models already reach users through third-party hardware. Reported placements include Tesla in-car systems, the Huawei Watch 5, Xiaomi’s XiaoAI assistant, and Cat King smart speakers. In those cases, the device maker built the voice interface.

The gray test is DeepSeek’s first effort to own both halves of the interaction in its own app. DeepSeek also shipped a push-to-talk input function in January 2026. Reports describe it as speech-to-text only, returning written replies. The current test completes the loop by speaking back.

Readers tracking IL’s DeepSeek coverage will recognize the pattern: our recent look at DeepSeek V4.1-Flash versus GLM 5.3 Flash is a serving and pricing story, while this is a client product story. Neither is a new checkpoint launch.

DeepSeek Voice Chat vs. ChatGPT Voice, Gemini Live, and Grok Voice

The single most common question Google surfaces for this topic is blunt: “Which AI allows voice chat?” As of September 2026, the answer includes several major assistants, with DeepSeek joining the field in limited testing rather than general availability.

Assistant Voice interaction status Voice selection Availability
DeepSeek app Gray test: spoken replies confirmed, two-way conversation unverified Four profiles (Beike, Bailang, Haixing, Anchao) Limited, unannounced, server-selected testers
ChatGPT (OpenAI) Mature voice mode with real-time, interruptible conversation Roster of around nine voices General availability on mobile and desktop, with free-tier limits; see OpenAI’s official voice mode documentation
Gemini Live (Google) Full two-way conversation with natural turn-taking Multiple selectable voices General availability in the Gemini app across platforms
Grok (xAI) Voice mode with spoken conversation in the Grok app Multiple voices General availability, strongest on mobile

The competitive gap is not subtle. OpenAI and Google shipped production voice modes in 2024. Both support interruption, emotional inflection, and camera or screen context alongside speech.

DeepSeek is entering two to three years late with a limited test that, on current evidence, may be closer to a text-to-speech readout than a live conversation partner.

Where DeepSeek could still reshape the field is price. Its API prices have repeatedly undercut Western rivals. A future DeepSeek voice API built on those rates could pressure voice app developers the same way the text API did. That is speculation for now, not reporting.

What Developers Can Build Today

Search interest in “deepseek voice chat api” and even hardware-specific queries like “deepseek ai voice chat esp32 s3” reveals a builder audience that does not want to wait for DeepSeek’s client experiment. The good news is that a voice interface can be assembled around DeepSeek’s existing text API today.

The standard pipeline uses three components:

  • Speech-to-text on the input side, converting the user’s microphone audio into a prompt
  • DeepSeek’s text model in the middle, generating the reply
  • Text-to-speech on the output side, reading that reply aloud in whatever voice the developer chooses
Diagram of a speech-to-text, LLM API, and text-to-speech pipeline that developers can build around DeepSeek's text API
How developers can assemble speech-to-text and text-to-speech around a text model API (Credit: Intelligent Living)

Speech-to-text options now include dedicated models such as Google’s Gemini 3.5 Transcribe speech-to-text model, while text-to-speech can be handled by commercial voice APIs or open-source engines.

On the maker side, boards in the ESP32-S3 class can run this pipeline in a small robot or desktop companion. That explains the hardware-specific search queries appearing around this feature.

The caveat is architectural. A hand-built pipeline is not the same as native integration. Wake-word detection, streaming latency, barge-in handling, and context continuity all fall on the developer. None of them inherit DeepSeek’s in-app experience. External voice wrappers should not be confused with what the gray test delivers.

Open Questions DeepSeek Has Not Answered

Question mark formed from soundwaves, symbolizing unanswered questions about AI voice features
Open questions surround DeepSeek’s voice test (Credit: Intelligent Living)

Public coverage of the gray test, including English summaries, stops at the existence of the speaker button and the four named voices. The following gaps remain open, and they are the questions that will determine whether this becomes a real product:

  • Latency: No end-to-end response time targets have been published for spoken replies, even though perceived voice quality depends heavily on speed
  • Platform coverage: It is unconfirmed whether the test spans iOS, Android, or both
  • Two-way conversation: Whether users can speak to DeepSeek and receive continuous spoken responses, or only trigger playback of written answers
  • Model routing: Whether voice sessions run the same model stack as text turns and how they interact with the ongoing V4.1-Flash and V4 Pro serving changes
  • Language and dialect handling: No details on output-side language coverage, accents, or dialect support
  • Privacy: No published policy on whether spoken input or synthesized output is retained, a live concern discussed widely in user forums and research on voice assistant data practices
  • Pricing and rollout timeline: No word on costs, regional expansion, or a general availability date

Until those answers arrive, the enthusiast speculation filling Reddit threads and maker forums should be read as exactly that. Research into greener and more inclusive voice AI architectures suggests that voice features will only grow more compute-intensive, which makes DeepSeek’s efficiency-first approach strategically interesting if the test graduates to a product.

Frequently Asked Questions

Does DeepSeek have voice mode?

Partially. DeepSeek has voice input via a push-to-talk function added in January 2026, and as of September 2026 it is gray-testing spoken replies with four selectable voices. A full, officially announced voice mode available to all users does not yet exist.

Is DeepSeek voice chat two-way?

Unconfirmed. Some reports describe two-way voice conversation, while others describe voice playback of written responses for selected testers. DeepSeek has published no documentation clarifying the interaction model.

Is DeepSeek voice chat free?

Unknown. The DeepSeek app is free to download, and core chat features have been free, but the company has not stated whether the voice experiment will carry any charge if it launches broadly.

Is DeepSeek voice chat available on Android or iPhone?

Public reports do not specify platform coverage. Testers report the feature inside the existing mobile app, but neither iOS nor Android availability has been confirmed, and there is no manual way to enable it.

Is there a DeepSeek voice API?

No dedicated voice API has been announced. DeepSeek’s API serves text and multimodal models; developers wanting voice today must compose speech-to-text and text-to-speech services around the text API themselves.

Which AI assistants allow voice chat?

ChatGPT, Gemini Live, and Grok all offer generally available voice modes with two-way conversation. DeepSeek is running a limited in-app voice test with four TTS voices, and several other assistants offer partial voice features.

What are DeepSeek’s four voice profiles?

Beike (贝壳), described as versatile and lively; Bailang (白浪), bright and firm; Haixing (海星), playful and sweet; and Anchao (暗潮), deep and resonant. These pinyin names are the clearest public handles, as no official English voice IDs have been published.

Conclusion

The newsworthy element of DeepSeek’s September voice experiment is packaging, not parameters: a native speaker control plus four named text-to-speech voices, arriving alongside but distinct from the V4.1-Flash launch and the reversed Pro-to-Flash API cutover. For now, the feature remains a staged gray test with more open questions than answers, from latency and platform coverage to pricing and true two-way capability. If DeepSeek’s efficiency playbook eventually extends to voice, the assistants that treat voice as a premium feature today may find the ground shifting beneath them.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Unisound U2-Flash Takes on Xiaomi MiMo in the Ultra-Cheap LLM Tier

Unisound announced U2-Flash in a voluntary filing to the...

Neuromorphic AI Inference: How China Mobile Cloud Cut Power Use by 40%

China's state telecom giant has paired brain-inspired silicon with...

Kimi K2.8 Preview: 1M Context Behind One Unchanged Model ID

Moonshot AI has quietly changed the model that powers...

DAMO RADAR: Alibaba’s Medical AI Detects Cancer and 146 Conditions

Alibaba's research arm has released something rare in medical...