Two audio releases arrived with the kind of demos that make a product meeting move too fast.
Alibaba’s Qwen-Audio-3.0-TTS presents a production-oriented speech system with Flash and Plus service tiers, multilingual and dialect coverage, long-form synthesis, expressive instructions, and inline performance controls. ByteDance’s Seed Audio 1.0 goes wider: it generates voices, sound effects, and ambience as one scene, and lets a creator place dialogue on a timeline.
Both releases are relevant to a small team building a support agent, localized onboarding, a game character, or short-form product video. They are not interchangeable products, and their launch metrics do not answer the same question. A low word error rate does not prove that a generated door slam lands on the cut. A coherent 30-second sound scene does not prove that the first audio packet arrives fast enough for a conversational agent.
The immediate product change is to stop asking which demo sounds best. Write down the audio job, assemble a small set of difficult fixtures, capture the whole delivery path, and measure the human repair required before a clip is usable.
This field note provides a proposed 12-run production bake-off. We did not call either new hosted service, run a listening panel, or reproduce the vendors’ reported results. Every threshold and score field below is a starting point for your own test, not a Y Build benchmark.
What launched—and why this is not a fair one-column comparison
The Qwen-Audio-3.0-TTS project page describes a speech-synthesis system built around a 12.5 Hz speech tokenizer and a progressive training process. The team reports support for 16 languages, 20 Chinese dialect regions, up to three minutes of one-pass synthesis, 86 fine-grained inline tags, free-form instructions, and 48 kHz output through super-resolution. It also presents results on content consistency, speaker similarity, control, long-form speech, and noisy reference audio.
The service is not simply the open Qwen3-TTS repository with a new label. Alibaba’s current Model Studio speech documentation lists qwen-audio-3.0-tts-flash as a hosted model and shows both non-streaming and server-sent-event output. The older Qwen3-TTS repository remains useful as an auditable open baseline, but its downloadable 0.6B and 1.7B artifacts, language set, runtime, and release date are different. A team should not infer that the new hosted model has published weights merely because the earlier family does.
Seed Audio 1.0 targets a different production unit. ByteDance says it models voice, sound effects, and ambience in a shared scene representation; accepts text and authorized reference audio; supports 20-plus languages; generates roughly two minutes in one pass with continuation; and can schedule dialogue at 100-millisecond intervals. The company says most evaluated scenarios exceeded a 90% usable-output rate, but it does not publish enough task distribution, rater, comparator, or API detail on that page for an outside team to treat 90% as its expected yield.
The bounded conclusion is useful: Qwen focuses on controllable production speech, while Seed expands the object from speech to a composed sound scene. A bake-off should compare them only on overlapping jobs and test each against a suitable existing workflow on non-overlapping jobs.
Six terms that keep audio evaluation honest
Time to first audio is the interval from the client sending an accepted request to the first playable audio frame arriving at the product. It includes network, queue, service, buffering, and client behavior. A vendor’s model-only latency or lab “first packet” number is not your user’s wait.
Real-time factor (RTF) is generation time divided by output duration. An RTF below 1 means generation completes faster than playback, but it says nothing about the first packet or jitter between later chunks.
Intelligibility asks whether a listener can recover the intended words. Word error rate (WER) or character error rate (CER) from an automatic transcription can flag failures, but product-critical names, amounts, dates, negation, and codes deserve their own weighted error count.
Performance adherence asks whether the output follows requested emotion, pace, emphasis, pronunciation, character, and non-verbal cues. A clip can be intelligible and still feel wrong for the moment.
Temporal alignment asks whether an event begins and ends inside its allowed window. For a video, “the door closes around three seconds” must become a tolerance such as 2.85–3.20s, not a subjective impression.
Repair minutes measure the human work between generated output and shippable output: regenerating, editing text, moving cues, cutting breaths, removing artifacts, mixing stems, checking pronunciation, and securing approval. It is the metric most likely to reverse a demo-based decision.
Do not combine these into one unexplained “quality score.” A conversational voice may trade a small naturalness preference for a materially faster response. A branded narrative clip may accept slower generation but reject one mispronounced product name.
Start with two jobs, not two model names
Imagine a three-person SaaS team preparing two audio surfaces for an international launch.
The first is an in-product setup guide. It reads short, dynamic instructions after a user connects a data source. It must start quickly, pronounce account fields correctly, survive English, Mandarin, and Japanese text, and never turn a warning into reassurance by dropping a negation.
The second is a 30-second launch clip. It contains one narrator, a notification sound at 7.5 seconds, a door sound between 14 and 15 seconds, light rain under the final scene, and a localized closing line that must end before the logo appears. The current baseline uses separate voice, effects, ambience, and a human timeline edit.
These jobs create separate questions:
- Can a hosted speech model improve the setup guide without increasing the user’s perceived wait or the team’s pronunciation risk?
- Can a unified audio model reduce assembly and repair work on the launch clip while preserving precise timing and editability?
Qwen-Audio-3.0-TTS belongs in the first comparison and may be tested for the narrator portion of the second. Seed Audio belongs in the second and may be tested on speech where the interface supports it. Neither should receive points for a capability the product does not need.
The 12-run fixture pack
Create four fixture families and generate three independent runs for each. Three runs are not a confidence interval, but they are enough to expose whether the “good demo” depends on a lucky sample. Keep model version, region, API mode, voice, sampling parameters, input, reference asset, and post-processing fixed within a comparison.
| Family | Fixture | Why it is difficult | Pass evidence |
|---|---|---|---|
| A. Conversational start | 20-second setup instruction with one variable account name | First-packet delay, dynamic text, natural pause | Client-side first-audio trace; exact account name; no gap over the jitter limit |
| B. Critical language | Warning containing a date, amount, negation, acronym, and product name in EN/ZH/JA | Ordinary WER hides unequal harm | Zero critical-token errors; full transcript; language reviewer sign-off |
| C. Directed performance | One neutral, one urgent, and one reassuring line using the same voice | Tests control without changing identity | Blind pairwise adherence vote; voice continuity; no exaggerated affect |
| D. Sound scene | 30-second dialogue, notification, door, rain, and logo cutoff timeline | Composition, event placement, mixing, editability | Cue-window ledger; dialogue transcript; repair minutes; export format |
Each family gets runs 01, 02, and 03, producing 12 observations per candidate configuration. If a provider offers Flash and Plus tiers, treat them as separate configurations rather than averaging them. If a unified scene model cannot return separate stems, record that as an operational constraint; do not silently give an editor a mixed file and call the job complete.
Use production-shaped text. Include punctuation, numbers, abbreviations, borrowed words, homographs, and a name that your product genuinely needs. Use only reference voices for which you have documented permission. Never upload a convenient employee, customer, celebrity, or scraped voice merely to make the test realistic.
A run card that another person can audit
Save one run card beside every generated file. The point is not paperwork; it is preventing a winning clip from becoming detached from the configuration that produced it.
audio_fixture_run:
fixture_id: "B-critical-language-ja"
run: "02"
candidate: "provider/model/version-or-tier"
region: "record-exact-service-region"
api_mode: "streaming|non-streaming|scene"
requested_at_utc: ""
first_playable_audio_ms: null
completed_audio_ms: null
output_duration_ms: null
critical_tokens:
expected: ["July 31", "$29", "do not disconnect", "YBuild"]
observed: []
event_windows: []
transcript_url: ""
audio_hash: ""
model_response_id: ""
retries: 0
provider_cost: null
repair_minutes: null
rights_receipt: "voice-rights/fixture-speaker-01.md"
reviewer_decision: "pass|repair|reject"
reviewer_notes: ""
Calculate RTF from the recorded completion and output duration, but retain the raw timestamps. Store the original response, file hash, and any failure or retry. If a provider silently changes a “latest” alias, a future regression can otherwise look like reviewer inconsistency.
The output of this experiment is not one favorite MP3. It is a folder containing inputs, run cards, generated audio, transcripts, cue ledgers, listener decisions, rights receipts, and a short promotion decision.
Listen with tasks and errors, not one overall vibe score
Mean Opinion Score (MOS) is familiar because listeners can rate naturalness or quality on a five-point scale. It is still useful when designed carefully, but it is too blunt to be the only gate. The Practical & Contextual Speech Synthesis Evaluation work argues for time-aligned error annotations with type and use-case severity because decontextualized naturalness scores do not tell a team what failed. A recent voice-reconstruction evaluation framework likewise reports that standard naturalness and similarity measures can be insensitive to the actual task trade-off.
Use two passes.
In the first pass, reviewers see no provider or model name. For overlapping outputs, they make pairwise A/B decisions on one question at a time: which starts the interaction better, which preserves the requested reassurance, which makes the product name clearer, or which scene better matches the timeline. Google’s replication work on MOS versus A/B tests found pairwise tests more reliable for system comparison while warning that standard errors were underestimated in both designs. With a tiny internal panel, report counts and disagreement; do not manufacture statistical certainty.
In the second pass, reviewers see the script and job. They mark timestamped errors using a fixed taxonomy: missing word, wrong word, critical-token error, wrong speaker, identity drift, unnatural pause, clipped onset, unwanted non-verbal sound, timing miss, ambience mask, artifact, emotion miss, or rights concern. Rate severity as cosmetic, repairable, or blocking.
If recruiting remote listeners, the ITU-T P.808 recommendation provides a real protocol for crowdsourced speech-quality evaluation, including test material and procedure. A handful of colleagues wearing unknown headphones is user feedback, not a standards-compliant listening study. It can still inform the launch if you label it honestly.
Measure the delivery path and the sound scene separately
For conversational speech, instrument the client. Capture request accepted, first byte, first decodable frame, playback start, each chunk arrival, playback underrun, completion, cancellation, and retry. Run from the regions and networks your users actually have. The Alibaba documentation shows that the new model can be requested in non-streaming or SSE modes, but interface availability does not prove product latency.
For sound scenes, create a cue ledger:
| Cue | Target window | Observed onset/end | Tolerance | Result | Repair |
|---|---|---|---|---|---|
| Narration line 1 | 0.0–6.8s | ±150ms | |||
| Notification | 7.35–7.65s | ±150ms | |||
| Door close | 14.0–15.0s | ±250ms | |||
| Rain bed | 18.0–29.2s | ±400ms | |||
| Final word ends | before 29.4s | hard gate |
Research on timing-controlled generation evaluates more than perceptual quality. ControlAudio reports event- and clip-level temporal measures, audio-distribution metrics, CLAP relevance, WER, MOS, and RTF. You do not need to recreate its research stack, but the multidimensional design is the right lesson: one pleasant clip can still miss words, cues, or operating constraints.
Add downstream outcome measures when the audio enters a real flow. For the setup guide, measure completion, replay, skip, interruption, and support escalation. For the launch clip, measure reviewer acceptance and final edit time before measuring engagement. A model that produces more audio but sends every clip through 25 minutes of repair has not automated the job.
Failure modes the happy-path demo will not show
Critical tokens fail while average intelligibility looks good. A transcript can be 98% correct and still change $29 to $99, “do reconnect” to “do not reconnect,” or the product name to a competitor’s name. Critical tokens are a zero-tolerance subset.
The first chunk is fast but playback stalls. A service may optimize time to first packet while later chunks arrive unevenly. Log underruns and maximum inter-chunk gap, not only the first response.
Expressiveness becomes instability. An emotion instruction can change pace, pronunciation, apparent age, or identity. Test the same voice across neutral, urgent, and reassuring lines and decide which properties must remain stable.
Timeline control covers dialogue but not the whole mix. Seed’s release page says timing control currently focuses mainly on character dialogue and lists more fine-grained control of other sounds as future work. Test every required cue; do not generalize a dialogue-timing feature to effects or ambience.
A single mixed output destroys editability. A coherent scene may still be unusable if legal asks to remove one line or the localization team needs a longer ending. Record stem availability, regeneration scope, and whether a small change preserves the rest of the scene.
Long output drifts. Speaker identity, room acoustics, loudness, language, or pacing can change across a two-minute clip or after continuation. Compare fixed checkpoints and join boundaries.
The service envelope changes. Region availability, quotas, price, retention, aliases, and rate limits are part of the product. The Qwen hosted endpoint shown in current documentation is region-specific; verify the exact region and data path available to your account.
Voice rights and data handling are release gates
Voice generation changes the artifact, but it does not create permission. Before a fixture can run, attach a rights receipt with the source of the voice, the person or licensor, allowed uses, territories, duration, revocation path, and whether derivative or cross-language performance is allowed.
Keep three voice classes separate:
- a stock or provider voice covered by published service terms;
- a designed synthetic voice with no human reference submitted by your team;
- a cloned or adapted voice built from an authorized recording.
The third class requires the strongest control. Restrict enrollment, review the exact reference, bind the resulting voice asset to an owner and use policy, and make deletion testable. Do not assume that permission to publish one recording includes permission to synthesize new languages, emotions, products, or political messages in the same identity.
Also document retention and training behavior for scripts, reference audio, and outputs; account access; logs; deletion; and incident contacts. If the provider does not answer, mark the field unknown. Unknown is a decision input, not a blank to fill with optimistic inference.
A promotion matrix based on the job
Score each candidate separately for conversational speech and composed scenes. Use hard gates before weighted preferences.
| Dimension | Conversational guide | Composed launch clip | Gate or preference |
|---|---|---|---|
| Critical-token accuracy | Zero errors across 3 runs | Zero in dialogue | Hard gate |
| First playable audio | Product-specific ceiling | Not primary | Hard gate for conversation |
| Playback continuity | No underrun over limit | N/A for offline export | Hard gate |
| Cue alignment | N/A | All required cues within tolerance | Hard gate |
| Rights receipt | Complete | Complete for every voice/reference | Hard gate |
| Repair time | Median minutes per accepted clip | Median minutes per accepted scene | Preference with ceiling |
| Listener preference | Task-specific A/B count | Task-specific A/B count | Preference |
| Cost | Per accepted clip, including retries | Per accepted scene, including repair | Preference |
| Editability | Regenerate one line safely | Stems or bounded local revision | Product requirement |
| Repeatability | Three-run variance visible | Three-run variance visible | Hold if unstable |
Promote only a configuration, not a brand. The decision should name model or tier, region, API mode, voice, parameters, fixture version, accepted use case, fallback, and review date.
Use four decisions:
- Go: all hard gates pass, repair stays under the agreed ceiling, and no unresolved rights or data question remains.
- Limited go: one narrow job passes; other languages, voices, or scene types remain blocked.
- Hold: the promise is relevant, but repeatability, service terms, editability, or metrics are incomplete.
- No-go: a critical token, rights, timing, or reliability gate fails and cannot be contained by the current workflow.
The fallback matters. Conversational speech needs cached or local safe prompts, text display, and cancellation behavior. A creative scene needs the existing multi-track production path. “Retry until one sounds right” is not a fallback; it is unbounded cost and selection bias.
Where this protocol applies—and where it does not
Use the fixture pack when audio affects a user task, brand identity, localization, accessibility, narrative timing, or support load. It is deliberately small enough for a startup to run before integrating a new provider.
It is not a clinical intelligibility study, an accessibility certification, a forensic speaker-verification evaluation, a language-wide fairness study, or proof that one vendor leads the market. Twelve runs cannot estimate rare failures. An internal panel cannot represent every accent, dialect, hearing condition, device, or cultural interpretation. Automated WER inherits the errors and biases of the recognizer used to score it.
The two launches also leave important unknowns. The public Seed page does not provide an open checkpoint, complete technical report, stem contract, or enough evaluation detail to reproduce its “usable” rate. Qwen’s project page reports broad results, but a hosted service still needs account-specific checks for price, quota, region, privacy, and alias behavior. Vendor-reported leaderboards and MOS values are evidence about the vendor’s test, not your launch outcome.
If the product only needs ten fixed accessibility prompts, a reviewed human recording library may remain simpler, safer, and cheaper. If every scene receives professional sound design, use AI as a draft source and evaluate time saved rather than pretending the final artifact is autonomous.
The 48-hour Build Lab plan
Hours 0–4: name the two audio jobs, write critical tokens and cue windows, choose an existing baseline, and complete rights receipts. Stop if the team cannot state who owns a reference voice.
Hours 4–8: prepare the four versioned fixtures, client instrumentation, run-card template, and blind file naming. Freeze text and parameters before anyone listens.
Hours 8–16: generate three runs per family and candidate configuration. Record failures, retries, raw timing, response IDs, cost, and hashes. Do not discard bad outputs.
Hours 16–24: run blind pairwise listening and timestamped error annotation. Have a native reviewer inspect each launch language’s critical tokens and pragmatics.
Hours 24–36: perform only the repairs the real workflow allows. Track minutes and tools. Test one small revision to learn whether the team can change a line without rebuilding the full artifact.
Hours 36–44: calculate pass counts, accepted-output cost, repair time, cue misses, and three-run spread. Separate confirmed observations from provider claims and remaining unknowns.
Hours 44–48: sign a go, limited-go, hold, or no-go receipt for each job. Save the fallback and a retest trigger for model alias, terms, price, region, or product-script changes.
The useful result may be that Qwen fits an interactive voice surface, Seed fits an early creative draft, the existing pipeline still wins final production, or neither passes. That is not an inconclusive bake-off. It is a product boundary discovered before users and editors paid for it.
References
- Qwen Audio, “Qwen-Audio-3.0-TTS”
- Alibaba Cloud Model Studio, “Non-real-time speech synthesis”
- QwenLM, “Qwen3-TTS” repository
- ByteDance Seed, “Seed Audio 1.0 Audio Creation Model”
- Dinkel et al., “ControlAudio”
- Pine et al., “Practical & Contextual Speech Synthesis Evaluation”
- Sanchez et al., “An Evaluation Framework for Text-to-Speech Voice Reconstruction”
- Ribeiro et al., “MOS vs. AB: Evaluating Text-to-Speech Systems Reliably Using Clustered Standard Errors”
- ITU-T Recommendation P.808, “Subjective evaluation of speech quality with a crowdsourcing approach”
- Hu et al., “Automatic Evaluation of Speaker Similarity”