Mercury 2.5 Diffusion LLM: 770 tok/s and Voice Latency

Learn how Mercury 2.5 Diffusion LLM throughput differs from voice response time, reconcile reported speeds, and benchmark your full agent pipeline.
Mercury 2.5 Diffusion LLM: 770 tok/s and Voice Latency
A language model generating 770 tokens per second can still leave a caller waiting for its first spoken word. The Mercury 2.5 Diffusion LLM makes that distinction especially important: faster text generation could shorten voice-agent responses, but throughput alone does not determine how quickly a conversation feels responsive.
As of October 2026, the headline figure has a specific source: a September 24, 2026, summary from Bittide AICompass reports that Mercury 2.5 achieved 770 tokens per second in evaluations attributed to Artificial Analysis. Separately, GIGAZINE’s September 9, 2026, coverage reports that Inception announced Mercury 2.5 on September 8 and claimed generation speeds of 1,107 tokens per second. Those figures should not be treated as interchangeable: the supplied reports do not establish identical workloads, hardware, or measurement conditions.
The architectural shift is what makes the announcement worth examining. Unlike conventional autoregressive language models, which generate text sequentially, Inception’s Mercury 2.5 uses diffusion to generate and refine multiple tokens in parallel, according to GIGAZINE. For voice-agent developers, that raises a practical question: can parallel text generation make an assistant start speaking sooner, or does it mainly help the assistant finish composing its answer faster?
Consider a simplified calculation. At a sustained 770 tokens per second, generating a 100-token answer would take approximately 130 milliseconds, excluding startup time and other processing. That is a throughput illustration—not a measured voice-response benchmark. A caller still waits for speech recognition, turn detection, any retrieval or tool execution, model startup, speech synthesis, and audio delivery.
The difference matters because a voice pipeline is not a single stopwatch. Several stages can overlap, and a streaming speech synthesizer may begin speaking before the language model finishes its answer. Conversely, quick completion offers limited conversational benefit if usable text arrives late or an external tool dominates the delay.
This article will unpack three questions:
- Throughput versus first response: What does 770 tok/s reveal, and what does it leave unmeasured?
- Streaming versus completed text: When can a diffusion model’s output become usable speech?
- Pipeline bottlenecks: Which delays should teams measure before changing their language model?
As of October 2026, platforms such as CallMissed offer real-time voice sessions through an API and SDK, making these pipeline-level questions relevant to teams building conversational AI.
The opportunity is real, but the useful benchmark is not simply “tokens generated per second.” It is time from the caller finishing a turn to hearing a useful, accurate response.
Does 770 tokens per second mean faster voice replies? Only if generation is the bottleneck

770 tokens per second can reduce voice-reply latency when text generation occupies a substantial part of the response’s critical path. If the longest wait comes from detecting the end of a caller’s turn, fetching account information, or preparing audio, faster generation will leave much of that delay untouched.
How much latency can faster LLM generation actually remove?
The useful calculation is not “How fast is the model?” but “How much waiting can this model change?” Bittide AICompass’s September 24, 2026, summary attributes Mercury 2.5’s 770-token-per-second result to Artificial Analysis evaluations; that figure measures text-generation speed, not an end-to-end voice interaction.
Consider an illustrative, non-streaming pipeline that waits for a complete 60-token response before starting speech synthesis. Assume the existing model generates 100 tokens per second and all other sequential work takes 900 milliseconds:
- Existing generation time: 60 ÷ 100 = 600 milliseconds.
- Generation at 770 tok/s: 60 ÷ 770 ≈ 78 milliseconds.
- Existing total wait: 900 + 600 = 1,500 milliseconds.
- New total wait: 900 + 78 ≈ 978 milliseconds.
Under those assumptions, generation becomes 7.7 times faster, but the total wait falls by only about 35%. These are calculated examples, not measured Mercury 2.5 voice-agent results; they also assume unchanged startup time, output length, and response quality.
This is Amdahl’s law in practical terms: accelerating one component cannot eliminate time spent elsewhere. If generation accounts for only 10% of a sequential response delay, even instantaneous generation can remove no more than that 10%.
Which delays should voice-agent developers measure first?
Measure the route to the first useful audible response, rather than adding every component’s duration together. Streaming stages can overlap, so a slow stage matters most when another stage must wait for it.
Instrument these milestones:
- Caller stops speaking: establish the reference point for perceived response latency.
- Turn is committed: record when endpointing and speech recognition make the request actionable.
- Required information arrives: distinguish retrieval and tool execution from model processing.
- Speakable text becomes available: capture the first stable phrase the speech synthesizer can use.
- Audio reaches the caller: include synthesis, buffering, and transport.
For example, a booking assistant awaiting a calendar lookup cannot safely confirm availability merely because its language model writes faster. A greeting without external dependencies is a better candidate for generation-led improvement.
As of October 2026, CallMissed offers custom REST tools and integrations with Cal.com and Google Calendar. Those capabilities illustrate why testing tool-dependent turns separately from simple conversational turns matters.
Why is first speakable text more important than completion speed?
Time to first token is not necessarily time to first speakable phrase. A speech synthesizer may need sufficient stable text and punctuation before producing natural audio; finishing the entire answer quickly is a different milestone.
GIGAZINE’s September 9, 2026, report describes Inception’s Mercury 2.5 as generating and modifying multiple tokens in parallel. That architecture makes it important to verify when output becomes stable and available for downstream synthesis, rather than assume its throughput establishes streaming behavior.
The decision rule is straightforward: replace the model when traces show generation blocking useful speech. Otherwise, optimize the stage actually keeping the caller waiting—and compare median and tail latency, not just average token throughput.
How does a diffusion LLM generate text differently from an autoregressive model?

An autoregressive LLM generates text one token after another; a diffusion LLM refines multiple token positions in parallel across successive passes. The key difference is the dependency structure: sequential generation builds on a committed prefix, while diffusion can revise parts of a candidate answer before releasing them.
How does autoregressive text generation work?
An autoregressive language model predicts the next token from the prompt and the tokens already generated. Tokens are text fragments—not necessarily complete words—and each new token extends the answer.
For an illustrative appointment response, generation might proceed like this:
- Predict “Your” from the conversation context.
- Predict “appointment” using the context plus “Your.”
- Continue toward “Your appointment is confirmed for Tuesday.”
Each output position depends on the preceding sequence. Modern inference systems can accelerate this process, including through techniques such as speculative decoding, but the underlying model still assigns probabilities according to a sequential, next-token formulation.
That structure suits streaming: an API can release a growing prefix while generation continues. However, an emitted token is not automatically a useful speech segment; a text-to-speech system may need several words or a clause before producing natural audio.
How does diffusion generate and refine text?
A diffusion language model, or dLLM, treats generation as iterative refinement rather than only next-token prediction. Depending on the implementation, it can begin with masked or corrupted token positions and progressively resolve them into coherent text.
Think of the distinction as extending a sentence versus editing a draft. An autoregressive model adds to the right-hand edge; a diffusion model can work on several positions during the same refinement pass.
An illustrative—not Mercury-specific—process is:
- Establish candidate output positions containing masks or provisional tokens.
- Predict or update multiple positions using the prompt and the current candidate.
- Repeat refinement until the output meets the system’s completion criteria.
GIGAZINE’s September 9, 2026, report describes Inception’s Mercury 2.5 as generating and modifying multiple tokens in parallel and reports Inception’s claimed speed of 1,107 tokens per second. That description supports the architectural distinction, but the supplied coverage does not establish Mercury 2.5’s exact masking schedule, refinement count, or streaming-release policy.
Parallel generation also does not mean every token is resolved instantly. Multiple refinement passes still require computation, and implementation choices determine how much work can happen together.
Why does the difference matter for spoken responses?
The important architectural question is when candidate text becomes stable enough to speak. Text on a screen can sometimes be revised; audio already heard by a caller cannot be silently replaced.
Consider a hypothetical booking answer whose provisional date changes from “Tuesday” to “Thursday” during refinement. Sending “Tuesday” to speech synthesis too early would turn an internal revision into a customer-facing error.
A voice integration therefore needs to distinguish:
- Provisional text: candidate tokens that may still change.
- Committed text: output the serving system has released for downstream use.
- Speakable text: a sufficiently complete, grounded phrase for speech synthesis.
As of October 2026, CallMissed’s developer AI API supports streaming, but that platform capability alone does not establish how any particular diffusion model commits output.
For a voice-agent pipeline, the architectural opportunity is parallel refinement; the integration requirement is a reliable release boundary. Finishing a draft quickly and safely exposing its first speakable phrase are different engineering problems.
Why do reports cite both 770 and 1,107 tokens per second, and which model was tested?

Both figures refer to Inception’s Mercury 2.5, but they come from different reporting chains—not a documented, like-for-like comparison. As of October 2026, the supplied sources establish a 770 tokens-per-second evaluation summary and a 1,107 tokens-per-second vendor claim; they do not establish why those measurements differ.
Where do the two Mercury 2.5 speed figures come from?
Bittide AICompass’s September 24, 2026, summary attributes Mercury 2.5’s 770 tokens per second to Artificial Analysis evaluations. That is an attributed benchmark result, but the supplied excerpt does not include the underlying test methodology.
GIGAZINE’s September 9, 2026, report attributes Mercury 2.5’s 1,107 tokens per second to Inception’s announcement. GIGAZINE dates that announcement to September 8, 2026. The distinction is between evaluation reporting and a reported vendor claim, rather than between two clearly identified model versions.
| Evidence item | Reported detail | Source and date | What it establishes |
|---|---|---|---|
| Evaluation speed | 770 tok/s | Bittide AICompass, Sept. 24, 2026 | Summary attributed to Artificial Analysis |
| Announcement speed | 1,107 tok/s | GIGAZINE, Sept. 9, 2026 | Speed claim attributed to Inception |
| Model identity | Mercury 2.5 in both reports | Bittide and GIGAZINE, Sept. 2026 | Same named model, not Mercury 2 |
| Release timing | Announced Sept. 8, 2026 | GIGAZINE, Sept. 9, 2026 | Dates the announcement, not the benchmark run |
| Generation method | Multiple tokens generated and refined in parallel | GIGAZINE, Sept. 9, 2026 | Describes diffusion, not a complete test protocol |
| Comparable conditions | Not established in supplied excerpts | Available reporting, as of Oct. 2026 | Hardware, workload and settings remain unresolved |
The 337 tok/s numerical gap, calculated from the two September 2026 reports, makes 1,107 approximately 44% higher than 770. That arithmetic is not evidence of a 44% performance improvement: no controlled before-and-after experiment is supplied.
Which model was actually tested?
The model identified in Bittide AICompass’s September 24, 2026, evaluation summary is Mercury 2.5. Nothing in the supplied evidence supports relabeling the 770 tok/s result as Mercury 2, or treating 1,107 tok/s as a separate Mercury 2.5 variant.
However, a shared product name does not establish an identical deployment. The excerpts leave several important questions unanswered:
- Endpoint and configuration: Were both measurements taken against the same serving setup?
- Workload: Were prompt lengths, output lengths and concurrent requests comparable?
- Measurement boundary: Was startup time excluded, and how was output throughput calculated?
- Output behavior: When did stable, usable text become available to downstream software?
These are possible sources of variation—not confirmed explanations for this particular gap.
Which number should voice-agent developers use?
Use 770 tok/s as the attributed evaluation figure and 1,107 tok/s as the reported Inception claim, with those qualifications intact. Neither should become an unconditional production-performance promise.
For a procurement document or technical design review, a defensible formulation is: “September 2026 reporting places Mercury 2.5 at 770 tok/s in an Artificial Analysis-attributed summary, while Inception’s reported announcement claims 1,107 tok/s under unspecified comparable conditions.”
Then validate the actual endpoint with your own workload:
- Record the model identifier, settings and test date.
- Measure output throughput alongside time to first usable text.
- Measure when the speech synthesizer receives that text and when the caller hears audio.
The practical conclusion is not to choose the larger headline. It is to preserve source provenance and measurement boundaries so a text-generation statistic does not silently become a voice-latency guarantee.
How do Time to First Token and first usable chunk affect end-to-end voice latency?

Time to First Token (TTFT) measures when text generation first becomes visible; time to first usable chunk measures when enough stable text exists to start speech synthesis. For end-to-end voice latency, the second milestone is often more actionable: a caller cannot hear a token counter, and a speech synthesizer may need more than the first fragment.
What is the difference between TTFT and the first usable chunk?
TTFT usually measures the interval from sending an LLM request to receiving its first output token. Benchmark implementations can differ, so confirm whether the measurement excludes connection setup and whether the first event contains actual answer text rather than metadata.
The first usable chunk is the earliest text segment your application can safely send to text-to-speech (TTS). Depending on the synthesizer and application, that might be a short phrase, a clause, or a complete sentence.
Consider a booking assistant beginning with “Your appointment is…”:
- The first token demonstrates that generation has started.
- The opening phrase provides little useful information.
- “Your appointment is confirmed for Tuesday at 3 p.m.” supplies a speakable, meaningful response—but only if the booking has actually been verified.
A pipeline can therefore achieve excellent TTFT while still delaying useful speech.
Does Mercury 2.5’s throughput reveal when speech can start?
No: output throughput does not establish TTFT or the arrival time of a stable, speakable segment. Bittide AICompass’s September 24, 2026, summary reports Mercury 2.5 throughput of 770 tokens per second, attributing the evaluation to Artificial Analysis.
According to GIGAZINE’s September 9, 2026, coverage, Inception’s Mercury 2.5 generates and modifies multiple tokens in parallel through diffusion. That makes the output interface particularly important: internal parallel generation and externally delivered streaming chunks are different things.
As of October 2026, the supplied reports do not establish Mercury 2.5’s chunk-release timing, revision behavior, or measured voice-onset latency. Before choosing a buffering strategy, developers should verify:
- Release timing: Does usable text arrive incrementally or after a larger block completes?
- Stability: Can already-delivered text change, or is each emitted segment committed?
- Boundaries: Can the application identify phrases suitable for synthesis without waiting for the entire answer?
Do not infer those properties from tokens per second alone.
How should teams measure the delay callers actually experience?
Instrument the path from end of caller turn to first audible response, recording intermediate timestamps rather than one aggregate duration.
Capture these milestones:
- Caller-turn endpoint detected.
- LLM request dispatched.
- First answer token received.
- First usable text chunk submitted to TTS.
- First audio frame received.
- Playback started at the caller’s device.
For a hypothetical trace, not a Mercury 2.5 benchmark, suppose the request leaves 200 milliseconds after the caller stops, the first token arrives at 300 milliseconds, usable text reaches TTS at 450 milliseconds, and playback begins at 600 milliseconds. TTFT is 100 milliseconds, but caller-perceived onset latency is 600 milliseconds.
That distinction identifies the optimization target. Smaller text chunks may start playback sooner but weaken prosody or expose incomplete statements; larger chunks improve context while adding buffering delay. Report median and p95 timings, and separately track first useful speech so that a quick “One moment” does not disguise a slow answer.
How much generation time could 770 tokens/s save on a short spoken answer?

At 770 tokens per second, an illustrative 80-token spoken answer needs about 104 milliseconds of generation time—roughly 1.50 seconds less than at an assumed 50 tokens per second. That saving applies to text generation, not automatically to the caller’s wait before hearing speech.
How do you calculate generation time from tokens per second?
Use generation time = output tokens ÷ sustained output rate. Bittide AICompass’s September 24, 2026, summary reports 770 tokens per second for Mercury 2.5, attributing the evaluation to Artificial Analysis; the calculations below use that reported rate as an assumption, not a guaranteed production speed.
These October 2026 illustrative calculations exclude startup time and assume constant throughput. The 50-token/s comparison is a hypothetical baseline, not a benchmark for a named competing model.
| Answer length | At an assumed 50 tok/s | At 770 tok/s | Generation time saved |
|---|---|---|---|
| 40 tokens | 800 ms | 52 ms | 748 ms |
| 80 tokens | 1,600 ms | 104 ms | 1,496 ms |
| 120 tokens | 2,400 ms | 156 ms | 2,244 ms |
| 160 tokens | 3,200 ms | 208 ms | 2,992 ms |
Across these examples, moving from an assumed 50 tok/s to 770 tok/s reduces generation time by approximately 93.5%. That percentage does not describe an equivalent reduction in end-to-end voice latency.
Also, tokens are not spoken words. Tokenization varies with language, punctuation, names, and formatting, so teams should count actual model output rather than estimate answer length from an English word count.
What would that saving mean for a short customer-service reply?
Consider an illustrative appointment response: “Your booking is confirmed for Tuesday afternoon. Please arrive ten minutes early and bring your reference number.” Its exact token count depends on the tokenizer; use an 80-token response budget for this worked example rather than assigning the sentence a universal count.
For an October 2026 hypothetical pipeline that waits for the complete answer before synthesizing speech:
- Hold other processing constant at 600 milliseconds. This is an illustrative assumption, not a measured platform result.
- Generate 80 tokens at 50 tok/s. Text generation takes 1,600 milliseconds, producing a combined wait of 2,200 milliseconds.
- Generate 80 tokens at 770 tok/s. Text generation takes approximately 104 milliseconds, producing a combined wait of 704 milliseconds.
The modeled total falls by about 68%, despite the generation stage shrinking by 93.5%. This is why a large throughput improvement can produce a smaller—but still meaningful—conversational improvement.
When should teams expect less than the calculated saving?
Treat the table as a generation-budget comparison, not a prediction of audible responsiveness. The practical benefit depends on how text becomes speech:
- Complete-answer buffering: Faster completion can translate more directly into an earlier synthesis start.
- Sentence-level streaming: Earlier usable sentences matter more than the final token’s arrival.
- Long spoken delivery: Faster text generation does not make a naturally paced answer finish speaking proportionally sooner.
As of October 2026, CallMissed’s developer AI API supports streaming and usage and request logs. Those capabilities are relevant to testing integrations, but they do not establish Mercury 2.5 availability or a measured latency improvement.
The decision rule is straightforward: calculate the potential saving from your actual answer lengths, then measure whether that saving reaches the first useful spoken phrase.
Mercury 2.5 vs Gemini Diffusion: what should a voice-focused inference API comparison measure?

A voice-focused inference API comparison should measure time to usable speech, response quality, and cost per successful turn—not rank Mercury 2.5 and Gemini Diffusion by token throughput alone. As of October 2026, the supplied research does not include comparable Gemini Diffusion measurements, so a defensible comparison needs a shared test harness rather than a declared winner.
Which metrics should a Mercury 2.5 vs Gemini Diffusion comparison include?
Bittide AICompass’s September 24, 2026, summary reports Mercury 2.5 throughput of 770 tokens per second, attributed to Artificial Analysis. That is a useful reference point, but not a matched voice benchmark against Gemini Diffusion.
Use the following comparison framework as of October 2026. “Not established” means the supplied evidence does not support a directly comparable result—not that a model lacks the capability.
| Measurement | Mercury 2.5 evidence | Gemini Diffusion evidence | Voice-focused test |
|---|---|---|---|
| Output throughput | 770 tok/s reported by Bittide AICompass | Not established | Measure delivered output over identical tasks |
| First speakable text | Not established | Not established | Time the first stable, TTS-ready phrase |
| End-of-turn to audio | Not established | Not established | Measure caller silence to first useful audio |
| Streaming stability | Not established | Not established | Check whether emitted text changes before speech |
| Tool-call reliability | Not established | Not established | Score valid arguments and successful execution |
| Quality-adjusted cost | Not established | Not established | Calculate spend per correctly completed turn |
The distinction between first token and first speakable text is especially important. A token can arrive quickly without providing enough stable meaning for a speech synthesizer to begin an accurate response. Conversely, a short, complete phrase may be useful even before the rest of the answer is available.
How do you make the API comparison fair?
Run both endpoints through the same speech recognition, turn-detection, retrieval, tool, and text-to-speech configuration. Keep prompts and intended answer lengths equivalent, while recording each provider’s actual token usage: token counts need not represent identical amounts of text across tokenizers.
A practical evaluation sequence is:
- Separate model-only and full-pipeline tests. The first isolates inference behavior; the second captures what callers actually experience.
- Replay representative tasks. Include short answers, knowledge-grounded questions, clarification requests, and tool-dependent actions.
- Report distributions, not just averages. Measure median and tail latency, documenting concurrency, endpoint region, context length, and warm versus cold requests.
For illustration, use a scheduling request that requires a calendar lookup. Score whether the model selects the correct tool, supplies valid arguments, and speaks only after availability is confirmed. A fast but incorrect booking response should count as a failure, not a latency win.
As of October 2026, CallMissed’s developer AI API provides usage and request logs, streaming, function calling, and structured outputs—capabilities relevant to instrumenting this kind of evaluation. Its catalogue should be checked separately before assuming either comparison model is available.
What should determine the winner for a voice workload?
Choose the endpoint that meets the application’s latency and correctness requirements together. Publish these outcomes alongside throughput:
- First useful audio latency, excluding empty acknowledgments.
- Task success and tool validity, with failures retained in results.
- Cost per successful turn, including retries and downstream speech processing.
Without those measurements, Mercury 2.5’s reported speed is a reason to test—not evidence that either model delivers the better phone conversation.
How can you benchmark a voice pipeline before switching models?

Benchmark the current and candidate models inside the same voice pipeline, using identical caller audio and measuring time to the first useful audible response. Switch only when the candidate improves real conversational latency without degrading task accuracy, tool execution, or interruption handling.
What should you keep fixed when comparing voice-agent models?
Start with a paired comparison: replay each test conversation against both models while holding everything else constant. Use the same speech recognizer, turn-detection settings, speech synthesizer, deployment region, knowledge base, and tool endpoints.
Bittide AICompass’s September 24, 2026, summary attributes Mercury 2.5’s 770-token-per-second result to Artificial Analysis; that reported text-generation benchmark is a reason to test, not a voice-pipeline acceptance criterion.
Build a representative test set rather than selecting only easy questions:
- Short exchanges: greetings, confirmations, and appointment changes.
- Knowledge retrieval: questions requiring evidence from your actual documents.
- Tool-dependent requests: checking availability or retrieving order status.
- Difficult audio: background noise, pauses, accents, and code-switching.
- Interruptions: callers correcting details while the agent is speaking.
Keep the instructions equivalent, but record any model-specific prompt changes. Otherwise, you may accidentally compare two different agent designs rather than two language models.
Which timestamps reveal whether the caller actually waits less?
Instrument a shared event timeline with synchronized clocks. Capture these milestones for every turn:
- Caller speech ends: the final speech sample, not merely the recognizer’s final transcript.
- Turn is committed: the pipeline decides the caller has finished.
- Model request starts: including whether retrieval or tools ran beforehand.
- Usable text reaches synthesis: text your production adapter actually allows the synthesizer to consume.
- First response audio plays: ideally measured at the client, not just when the server sends an audio packet.
- Useful answer begins: distinguish substantive speech from filler such as “Let me check.”
Also capture tool execution, response completion, cancellations, and errors. For diffusion output, test whether the integration exposes stable, speakable chunks; do not assume the first emitted content is suitable for playback.
Report median, p95, and failure rate, separating tool-free turns from tool-dependent turns. A pooled average can conceal slow booking requests behind many fast greetings.
How do you turn benchmark results into a switching decision?
Use a worked example to set expectations. In an illustrative October 2026 test—not a measured Mercury 2.5 result—suppose median end-of-speech-to-useful-audio latency falls from 900 milliseconds to 720 milliseconds. That represents a 20% reduction, but the improvement is insufficient if the candidate introduces incorrect bookings or worsens tail latency.
Before testing, define acceptance gates for:
- Responsiveness: median and p95 useful-audio latency.
- Correctness: successful task completion and valid tool arguments.
- Conversation quality: unnecessary verbosity, premature responses, and interruption recovery.
- Economics: total cost per successfully resolved conversation, including retries.
Run repeated trials in randomized order under expected concurrency, explicitly recording caching and warm-versus-cold conditions.
As of October 2026, CallMissed provides eval suites, A/B experiments, and call scoring against custom QA rubrics—capabilities relevant to evaluating these trade-offs. That does not establish Mercury 2.5 availability or performance on CallMissed.
Finish with a limited live rollout and a rollback threshold. The winning model is the one that makes your actual conversations faster and reliable, not simply the one with the largest throughput headline.
What should you change in your voice-agent workflow after measuring latency?

Change the stage that delays the first useful spoken response—not automatically the language model. After measuring voice-agent latency, turn the results into targeted experiments, then keep only changes that improve responsiveness without reducing accuracy or making interruptions harder.
For an October 2026 optimization review, treat Mercury 2.5’s throughput as a reason to test, not a deployment verdict. Bittide AICompass’s September 24, 2026, summary attributes a 770-token-per-second result to Artificial Analysis; that figure does not establish an end-to-end voice latency improvement.
Which workflow change matches your measured bottleneck?
Use the following decision table to connect an observed delay with a practical intervention. These are proposed experiments, not measured Mercury 2.5 performance results.
| Measured bottleneck | Workflow change to test | Primary success metric | Guardrail |
|---|---|---|---|
| Turn detection waits too long | Tune end-of-turn settings on representative recordings | Speech-end-to-turn-commit time | Avoid cutting off pauses or unfinished sentences |
| Speech recognition finishes late | Test incremental transcripts and earlier processing of stable text | Time to usable transcript | Preserve names, numbers, and code-mixed speech |
| Retrieval or tools dominate | Cache eligible reads; parallelize independent lookups | Tool-stage p95 latency | Check freshness; never duplicate write actions |
| Model output arrives late | Compare models with identical prompts, tools, and answer requirements | Time to first usable text | Maintain task accuracy and tool-call correctness |
| Speech synthesis starts late | Test smaller, semantically complete text chunks | Text-ready-to-first-audio time | Avoid awkward prosody or speaking unstable text |
| Audio delivery adds delay | Inspect buffering and test network paths | Audio-ready-to-caller-playback time | Monitor dropouts and interruption handling |
The important distinction is usable output. A text fragment that arrives quickly but cannot safely be spoken is not equivalent to a useful first sentence.
How should you test a faster LLM without changing everything?
Run a controlled comparison before redesigning the pipeline. GIGAZINE’s September 9, 2026, coverage reports Inception’s claimed 1,107 tokens per second for Mercury 2.5, but the supplied evidence does not establish measurement conditions identical to the 770-token-per-second result.
- Hold the workflow constant. Use the same caller recordings, system instructions, knowledge sources, tools, and response-length requirements.
- Change one variable. Compare the language model first; test speech chunking or endpointing separately so improvements remain attributable.
- Evaluate complete interactions. Include short questions, interrupted answers, slow tool responses, and noisy or multilingual calls.
Track p50 and p95 speech-end-to-useful-audio latency, alongside task completion, factual accuracy, and successful interruption handling. Faster median replies can conceal a worse slow-call experience; shorter answers can also appear faster while omitting necessary information.
As of October 2026, CallMissed offers eval suites, A/B experiments, call recordings, transcripts, and call scoring against teams’ own QA rubrics. Those capabilities can support comparative workflow reviews, while stage-level timing still needs appropriate instrumentation.
When should you keep, revise, or roll back a change?
Set acceptance criteria before examining the results:
- Keep changes that improve the targeted delay while preserving required answer quality.
- Revise changes that improve typical calls but worsen tail latency or conversational flow.
- Roll back changes that introduce premature speech, incorrect tool actions, or unreliable interruption recovery.
For example, if order-status lookup dominates the wait, prioritize that lookup rather than expecting faster text generation to solve it. If useful model text arrives promptly but audio starts late, test synthesis and buffering next. The workflow should follow the measured critical path—not the most impressive benchmark headline.
Frequently Asked Questions

Is the Mercury 2.5 LLM’s 770 tokens per second enough for a responsive voice agent?
Is time to first token the same as voice-agent latency?
Can the Mercury 2.5 LLM stream text directly into speech synthesis?
Why do Mercury 2.5 benchmarks report both 770 and 1,107 tokens per second?
Which latency metrics should developers benchmark before changing a voice agent’s LLM?
Does CallMissed offer the Mercury 2.5 LLM as of October 2026?
Conclusion
Mercury 2.5’s reported 770 tokens per second could shorten voice-agent replies, but only when text generation is the bottleneck. The meaningful outcome is not how quickly an answer is completed—it is how soon a caller hears a useful, accurate response.
Bittide AICompass’s September 24, 2026, summary attributes Mercury 2.5’s 770-token-per-second result to Artificial Analysis evaluations. GIGAZINE’s September 9, 2026, coverage separately reports Inception’s 1,107-token-per-second claim. These figures signal the potential of diffusion-based generation, but the supplied reporting does not establish matching workloads, hardware, or measurement conditions. Neither figure, by itself, establishes end-to-end voice latency.
Four takeaways should guide the next round of voice-agent experimentation:
- Throughput and first-response latency answer different questions. At a sustained 770 tokens per second, a 100-token answer would take approximately 130 milliseconds to generate, excluding startup and other processing. That calculation illustrates text-generation speed; it does not measure the interval between a caller finishing a sentence and hearing the assistant begin a useful reply.
- Parallel generation matters only when its output becomes usable. According to GIGAZINE’s September 2026 coverage, Inception’s Mercury 2.5 generates and refines multiple tokens in parallel rather than producing them sequentially. For voice applications, the critical question is when that text can reach speech synthesis—not simply when the entire answer finishes. Faster completed text and earlier spoken audio are related, but distinct, outcomes.
- The whole pipeline still determines conversational responsiveness. Speech recognition, turn detection, retrieval, tool execution, model startup, speech synthesis, and audio delivery all contribute to waiting time. Some stages can overlap, particularly when speech synthesis consumes streaming text. Others can dominate the delay, leaving even a substantial increase in generation throughput with little effect on the caller’s experience.
- Measure before changing models. Teams should identify which stage delays the first useful spoken response, then compare whether a different model actually reduces that delay. A faster generator is promising when generation dominates; it is less decisive when an external tool or late-arriving usable text holds up the conversation. Response accuracy must remain part of that assessment.
What should voice-agent developers watch next?
Watch for comparable end-to-end voice measurements, clearer streaming behavior, and evidence that diffusion-generated text can reach speech synthesis sooner—not just finish faster. Those results will determine whether the throughput breakthrough becomes a conversational breakthrough.
As of October 2026, CallMissed offers real-time voice sessions through an API and SDK, making it a relevant platform to explore these pipeline-level questions. Before choosing your next model, ask: will callers hear a useful answer sooner, or will the system merely finish writing it faster?
Related Reading
- Best LLM for Voice Agents in 2026: GPT-6 vs Claude
- Best LLM for Voice Agents in 2026: GPT-6 Astra vs Claude Fable 5.1
- Agentic AI Governance: India Voice Support Safety Guide
Sources
Discussion
Related Posts
Ready to automate customer conversations?
Launch AI voice agents and WhatsApp bots with CallMissed — one API, 22+ Indian languages.



