For now, I tried testing a similar reference-audio issue using publicly available materials:
Short answers to your eight questions
I would treat your goals as two separate conditioning problems, not as a reason to discard your WebUI or workflow: (a) preserve a recognizable speaker while generating new, correctly timed dialogue without an extra voice, and (b) preserve the person while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but not your exact overlapping-voices output. The distinction matters throughout.
1. How closely should Ref2VA reproduce the speaker?
Using an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The official full-reference format distinguishes reference (use characteristics without directly copying the signal) from fully_copy (reuse the source audio). It does not promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge voice resemblance, correct words, intelligibility, speech timing, and unwanted additional voices separately. Success on one is not proof of success on the others.
2. Is the apparent “two speakers at once” behavior known?
There are neighboring reports: ComfyUI #16155 describes unexpected speech and timing errors with an audio reference; MiniMax-H3 discussions #64 and #76 include reports of reference material intruding into dialogue or clipped/gibberish vocalization; MiniMax-H3 #17 concerns voice binding between characters. Those are related symptoms, not proof of the same bug. Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which.
In a separate test using #16155’s public assets, extra words were recognized in 3/3 seeds with audible reference audio, versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block. This points toward the encoded reference values, not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below.
3. Could three LoRAs interfere with the audio?
Yes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be sequential patches before the final sample, not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred without LoRAs or Heretic, so those are not prerequisites for that behavior, but they could still contribute in your pipeline.
4. Is strength 1.0 on each LoRA too much; should it be 0.5–0.7?
Three 1.0 settings are worth testing, but they do not add up to a meaningful global “3.0 strength.” Effects depend on target layers and adapter interactions. Nor is 0.5–0.7 a universal safe range. Try all off → one at a time → necessary combinations, then sweep strengths only where a difference appears; this preserves your visual goal.
5. Could the pruned INT8 ConvRot checkpoint be responsible?
Still an open variable. My runs used minimax_h3_ref2va_pruned_int8_convrot.safetensors and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not uniformly break dialogue, but might affect this failure. No matched official-BF16 test exists. With 12 GB VRAM, check the graph and one-seed contrasts first. The model card and native ComfyUI guide document the distinct deployment artifacts; successful loading alone does not establish output parity.
6. Must the audio be attached to both Ref2VA nodes (400 and 587)?
Not necessarily. MiniMaxH3ReferenceToVideo builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to final node 587 may be sound. Check the JSON actually submitted to /prompt after Gradio rewrites it: trace both final BasicGuider conditioning and sampler latent, and establish what 400 still supplies. Adding audio to 400 without that trace could introduce a new constraint. The native implementation and node docs are useful references.
7. Why does the old background persist despite requesting a basement?
A full reference picture encodes both the person and the old scene. If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish identity reference from frame preservation: compare the full picture with a subject-focused crop or background-removed reference, keeping the target basement/action/seed fixed. You have already softened preservation wording, added a single-subject constraint, and switched max to match; their outcome is not established. I did not reproduce this background issue.
8. Is ref_image_size="max" …?
The public post ends at this heading, so I will not invent the missing question. ref_image_size controls image preprocessing, not a linear reference-strength setting. The official node docs explain that match scales references toward the output pixel area, while max keeps greater detail (up to the 2048px short-edge reference convention), at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your match change is unknown.
Suggested default route: preserve the features, isolate the boundary
- Freeze what actually runs. Save one failing run’s final API JSON after file uploads, prompt edits, LoRA chaining, and the WebUI’s input rewrites. Follow the final guider’s conditioning and sampler’s latent back through
587, 400, and the media loaders. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.
- Use one paired voice test, not a rebuild. At the same seed and settings, compare the current audible reference with all custom LoRAs off. If the bad voice remains, prioritize reference roles, native Audio VAE conditioning, and model/encoder differences. If it clears, reintroduce adapters one by one. Continue to use the audible reference for the target result: silence/zero latent are diagnostic controls, not voice cloning.
- If extra utterances remain, verify the reference contract. Bind
<Audio N> to (S1), use reference rather than fully_copy for new dialogue, keep the sample clean, and specify the exact line. For timing, connect speech to an on-screen action as well as a shot interval; absolute timestamps alone do not guarantee onset. A shorter clip is optional, not a universal fix: cropping still left extras in 2/3 seeds.
- Handle the background independently. Start from your already-updated
match setting; compare full-scene versus subject-focused image. Check whether anything actually anchors the first frame or reuses an old visual guide. Score person identity and new location separately, even if the LoRA stack remains unchanged.
- Escalate when a cheaper comparison identifies a boundary. Standard Qwen versus Heretic, alternative checkpoints, matched official BF16, or ComfyUI versus Diffusers can be useful, but each adds downloads, VRAM demand, and possible unrelated numerical differences. They are not prerequisites for the first two checks.
The goal remains new speech in the target voice and the same person in a new room.
First checks: what to observe before another expensive render
Useful optional evidence: one failing submitted /prompt JSON (private paths removed), a short generated-audio excerpt, the intended line, seed, reference duration, and checkpoint/adapter names. These are not prerequisites for the suggestions below.
No-render audio check: If the voice sounds doubled, compare ComfyUI’s raw decoded audio (if retained) with the final Gradio/WebUI-exported or muxed audio from the same run. Check left and right channels separately before downmixing. Only after export? Inspect muxing, duplicated tracks, and post-processing first. Already in the raw audio? Focus on generation, reference binding, or decoding. If the raw track was not saved, mark this comparison as unavailable instead of inferring an export bug; when both tracks exist, no new inference is needed.
Gradio workflow template
-> file substitutions / prompt edits / adapter chain
-> actual payload sent to POST /prompt
-> LoadAudio -> 587.ref_audios.ref_audio_0
-> 587.positive -> final BasicGuider (?)
-> 587.LATENT -> final sampler input (?)
-> 400 outputs -> which actual consumers, if any?
-> final H3 model / guider / sampler -> decode / output
The question marks are edges to check, not accusations that they are wrong. ref_audios.ref_audio_0 is a zero-based API field, while model text uses the one-based label <Audio 1>; that apparent mismatch is normal. Other enabled audio-bearing reference media can affect label ordering, so inspect the packed/reference list rather than treating numbering as a universal equivalence. To classify “double voice,” distinguish: overlapping independent speakers, extra speech before/after the line, timbre drift, and duplicated/reverberant audio. These suggest different causes; ASR alone cannot distinguish them.
A related public reproduction, in brief
I used the public image and voice WAV from someone else’s ComfyUI issue #16155, where the requested line was This park sucks. around 2 seconds. This is not a reproduction of your particular reference, output, multi-node graph, or two-simultaneous-voices claim. My Colab setup used native ComfyUI, pruned INT8 H3 Ref2VA, a standard (non-Heretic) Qwen, no LoRAs, 512×288, 124 frames at 24 fps (~5.17 s), 20 sampling steps, and three deliberately reused seeds (4242, 5151, 6262).
The decisive panel kept the reference block in place while changing the encoded values:
| Reference input |
Extra words recognized beyond requested line |
Interpretation |
| Original audible 7.05 s WAV |
3/3 seeds, both Whisper models |
Related anomaly reproduced in this setup |
| Same-duration time-reversed WAV |
2/3 agreed; 1/3 ASR disagreement |
Ordinary forward-intelligible English is not clearly required |
| Same-duration PCM silence, Audio-VAE encoded |
0/3, both ASR models |
Reference block still exists; encoded silence is not a zero latent |
| Original reference latent zeroed after encoding |
0/3, both ASR models |
Same reference block shape, but an artificial out-of-distribution control |
Cropping the reference still left extras in 2/3 seeds. Two Whisper versions agreed on most observations but are not human listening tests; neither words alone nor waveform activity establish concurrent speakers, literal copying, or restored voice identity.
Interpretation: audio-derived conditioning values matter in this small panel, while the responsible feature and root cause remain unknown. The expanded controls follow.
A. Reproduction protocol, per-condition results, and evidence limits
Scope and controls
The #16155 author reported both anomalies and successful seeds. I used their public image and WAV, not your audio or WebUI graph. Each panel reuses seeds 4242, 5151, 6262 for paired comparisons; those are not population samples for an H3 failure-rate estimate.
Reproducibility capsule (local Colab tests, not a public benchmark):
ComfyUI: a4b5a045e56fc334903db8457b728b64e006119c
HF model repository revision recorded: e5eb578a89295337b8ff433a035929ce0279e0b6
Checkpoint: minimax_h3_ref2va_pruned_int8_convrot.safetensors
Conditioner: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (non-Heretic)
Audio VAE: minimax_h3_audio_vae_fp32.safetensors
No LoRA; 512x288; 124 frames at 24 fps; 20 steps
Seeds: 4242 / 5151 / 6262
Input WAV: 7.05 s, stereo, 48 kHz
WAV SHA-256: eb0e6e869cebe54325388affda68aa9f6bcf892365b3d169f60fee2257946f49
The repository revision is recorded in the downloaded artifact metadata; it is not a published hash of every loaded tensor. Output FLACs, per-run logs, and ASR tables are retained locally, not hosted at a public results URL. The test also uses a newer pinned ComfyUI core-node reconstruction rather than the #16155 author’s exact environment.
| Panel |
Intervention |
Both Whisper models’ extra-text findings |
Confound to remember |
| E1: 6 outputs |
Reference absent vs. original present |
0/3 vs. 3/3 |
The Qwen label and packed audio layout also change |
| E2: 9 outputs |
Original 7.05 s / label only / cropped ~5.17 s |
3/3 / 0/3 / 2/3 |
Label-only removes the audio block; crop changes its length |
| E3: 12 outputs |
Original / time-reversed / PCM silence / post-VAE zero latent |
3/3 / 2/3 + one disputed / 0/3 / 0/3 |
Fixed block shape, but several acoustic properties still differ |
E2’s “label only” retained <Audio 1> but removed the audio reference block; attention/time positions could change. E3 resolved that principal confound by retaining a [1, 32, 2, 282] reference block, original duration, prompt, and seeds in all 12 runs. Same shape is not proof that every internal tensor operation is identical.
Why silence and zero latent must not be conflated
PCM silence was a zero waveform of the original sample duration encoded normally by Audio VAE; the resulting latent was not zero (mean absolute magnitude ~0.461, versus ~0.394 for the original). Post-VAE zero latent started with the encoded original, then used an experimental, local ComfyUI patch to apply zeros_like to its reference representation. This kept a block with the same dimensions but supplied an artificial value. Neither supplies a usable example of the intended person’s voice. Indeed, the silence latent’s larger mean absolute magnitude argues against “lower latent norm” as a sufficient explanation.
The reversed WAV kept sample count and broad amplitude properties but reversed phoneme progression and temporal structure. Both Whisper decoders detected extras at 4242 and 5151; for 6262, base.en found extra text and small.en did not. This does not prove that the source’s linguistic meaning is irrelevant—only that intact, forward-spoken English does not appear necessary for the two agreed anomalous outputs.
Measurement boundaries
Whisper base.en and small.en are two versions of one ASR family, not independent listeners. Some reported extra suffix words differed between them; neither can prove exactly what a person would hear, let alone simultaneous distinct speakers or verbatim copying from the reference. Early RMS/audio activity likewise measures waveform energy, not linguistic onset. For E1, the first substantial activity was approximately 0.1–0.3 seconds with the original reference versus 1.6–2.1 seconds without it, despite an instruction to speak near 2 seconds. E3 silent/zero controls also showed different volume peaks, so “no extra ASR words” is not equivalent to proven sound quality or speaker fidelity.
Established here: changing values inside a same-shape reference block changes extra-speech-like outcomes. Not established: which acoustic/linguistic feature drives this, whether BF16/Diffusers agrees, or whether your two overlapping voices share the mechanism. One image/WAV pair and three selected seeds are not a general benchmark. The capsule above identifies the relevant artifact revisions and WAV hash; exact API graph, prompt, outputs, and listening/ASR procedure also matter.
B. How reference labels, sound latents, and timing interact
Two routes for one reference
The inspected ComfyUI path conceptually separates the reference’s description/label from its waveform encoding:
standalone reference audio
+--> Qwen-side presentation: <Audio 1>, speaker (S1), role
+--> PCM -> 32 kHz stereo -> H3 AudioVAE.encode
-> packed reference-audio latent and time positions
-> shared H3 audiovisual transformer -> new audio/video
Qwen does not receive raw PCM samples as though they were text. The Audio VAE supplies numerical reference data; the Qwen-side presentation tells the model how to use them. A prompt saying “voice timbre, not copied audio” expresses a desired relationship; it is not a guaranteed filter that removes lexical or timing information from the sound latents.
The official full-reference guide distinguishes <Subject N>, <Picture N>, <Video N>, <Audio N> and reference from fully_copy. Speaker binding should be explicit when appropriate; numbering of standalone audio and enabled audio tracks extracted from videos needs attention. The model’s <Audio 1> is one-based, whereas ref_audios.ref_audio_0 is a zero-based API input socket.
For example, a role fragment, not a guaranteed complete prompt:
subject_definitions:
<Subject 1> (S1) is the person referenced by <Picture 1>.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
summary:
[reference generation + audio reference] A new scene with <Subject 1>
speaking newly scripted words in the referenced voice style.
retention_analysis:
<Subject 1>: fully_preserved - preserve identity, not the old location.
<Audio 1>: reference - follow speaker timbre without copying the signal.
... in the target shot ...
<Subject 1> (S1) says: <d>[English] The new sentence.</d>
This fragment is intentionally incomplete. In a full shot, anchor the line to a visible event (e.g., when the person reaches the doorway) as well as describing the shot’s timing and quiet intervals. The ComfyUI prompt guide recommends visible-event anchors in multi-speaker reference cases, rather than expecting absolute timecodes to bind speech reliably. This may help, but does not guarantee the intended onset or eliminate reference-induced vocalization.
The audio length difference is real, but not a full explanation
The H3 Audio VAE operates at 32 kHz with 800 samples per latent step (40 steps/s). The 7.05-second reference yields 282 latent steps, compared with ~207 for the 5.17-second output. In inspected Diffusers normalization, input reference sound is truncated to generated duration before resampling; our pinned ComfyUI path retained the longer source. This implementation difference deserves inspection, but after cutting to output duration, E2 still showed extras at 2/3 seeds. Thus duration truncation is neither a proven cause nor a universal fix. The separate Audio VAE crop issue #15970 concerns an edge case; our original resampled WAV had exactly 225,600 = 282 × 800 samples, so fractional-hop loss is a weak explanation for multi-second extra dialogue here.
If the voice source is dense or noisy, a clean single-speaker excerpt, explicit (S1) binding, and constrained new dialogue are reasonable practical trials. People describe mixed success with such measures in HF #64 and #76. They should not be represented as established repairs. To identify a deeper mechanism, an optional equal-duration test can vary amplitude, matched-energy nonspeech, or where within the same reference length speech occurs. PCM amplitude and latent magnitude must both be logged because the VAE is not a simple linear gain control.
C. Two Ref2VA nodes: trace the final API graph, not the node count
The native Ref2VA node assembles conditioning plus an initial audio-video latent. A separate guider/sampler performs the diffusion steps. Therefore 400 and 587 can be two builders, alternative branches, or inputs into a genuine multi-stage graph; their names alone do not define the execution stages.
Inspect the rewritten payload: class_type and inputs of 400/587; the final BasicGuider’s positive source; the final sampler’s LATENT source; the patched model; 587’s clip, video VAE, Audio VAE, and ref_audios.ref_audio_0; plus any first-frame/guide path. Even if BasicGuider uses 587, the final sampler may get its latent elsewhere. That can be intentional, but should be explicit.
"587": {
"class_type": "MiniMaxH3ReferenceToVideo",
"inputs": {"ref_audios.ref_audio_0": ["LOAD_AUDIO_ID", 0]}
}
This is a fragment illustrating the key, not a runnable full node. In ComfyUI’s API representation, that dotted input name is normal. The important questions are whether the linked producer exists, which output it returns, which media are packed before it, and what downstream nodes consume it. A template may be correct while the WebUI’s later JSON rewrite silently points a final edge elsewhere.
If 587 alone supplies both final conditioning and its intended latent: final-only audio can be perfectly sensible; test reference semantics rather than adding it to 400. If 400 supplies a visual guide or latent that survives into the final sampler: inspect whether it anchors the original scene; this may matter for your background issue without making 400 an audio problem. If both branches really have separate samplers: decide explicitly which stage should use the voice reference; that is a stage-design choice, not a universal requirement. If an edge points to stale inputs after WebUI rewriting: repair that specific link before judging H3’s model behavior. This static graph audit has a very high information-to-compute ratio.
D. LoRAs, Heretic Qwen, the INT8 checkpoint, and 12 GB VRAM
A “visual” LoRA may touch shared audiovisual layers; presence alone is not evidence of causation. Three 1.0 settings are loader scales, not a summed global strength. Use none → each alone → necessary combinations → targeted strength sweep, stopping when the result is informative. Fix WAV, reference ordering, dialogue, seed, sampler, and final API graph. Score intended visual quality as well as speech to avoid “fixing” the voice by losing the desired look.
Your Heretic Qwen3-VL variant differs from the standard Qwen used in my reproduction. Since the latter showed extra reference-dependent speech too, Heretic is not necessary for the related failure. Still, an encoder can influence prompt/reference interpretation. Compare it separately only if an exact checkpoint/processor-compatible substitute is available; successful loading is not proof that two multimodal presentation paths are semantically identical.
The pruned INT8 ConvRot model generated both control and anomalous outputs; this says nothing about BF16 parity. Comparing precisions meaningfully would require matching revision, references, schedule, seed/noise, and backend as far as practical; differences are otherwise not uniquely due to quantization. With 12 GB, leave this resource-intensive branch until cheaper comparisons are exhausted.
You also mentioned trying to manually unpatch/unload the Audio VAE after reference encoding and then removing that change. Treat that as already reverted. Extra model-manager intervention can make state handling harder to reason about, but it is not evidence of the current symptom.
E. Keeping the person while changing the environment
The reference can encode who the person is and where the original photo was taken. They are not automatically separated by the word “subject.” For your basement scene, the key distinction is between a character/identity reference and a frame anchor that must resemble the supplied image. A first-frame anchor (if present) can reinforce that bias.
A low-cost scene comparison holds the person, action, new basement description, seed, and model constant, while swapping full original picture for subject-focused crop/background-removed picture. Then ask separately: is the face recognizably the same; does the basement appear; is the pose or old lighting copied; does quality degrade? The crop may remove useful hair/clothing context, so select it to retain the identity cues that matter. If a visual guide or retained first frame is actually connected, test that constraint rather than endlessly rewriting the text.
Describe the new basement positively: walls, floor, lighting, objects, camera, position, and action; preserve facial identity without asking for the old composition. Your relaxed wording and match setting are already the baseline, not untried remedies. max may later improve facial detail, but does not substitute for subject-only reference selection. This branch is inferred from the conditioning design, not reproduced on your image.
F. Symptom decision tree, further research, and useful links
What exactly is wrong in the output audio?
|
+-- Two clearly simultaneous voices
| -> compare raw decoded stereo with final WebUI/exported audio;
| then inspect speaker binding if overlap is already in raw audio.
+-- One extra utterance, especially early or after the scripted line
| -> closest to public #16155; examine reference audio values,
| label order, dialogue timing, and a fixed-seed reference A/B.
+-- The speaker identity changes during otherwise correct words
| -> compare voice resemblance separately from transcript accuracy;
| then test adapter/encoder/reference variants one at a time.
+-- Echo, doubled channel, or post-export duplication
| -> separate waveform generation from mixing/decoding/export.
+-- Unintelligible sounds but no clearly separate speaker
-> listen; log activity/peaks; do not infer exact words from ASR.
For a deeper optional investigation, high-information experiments are: fixed-duration gain sweep with PCM and latent RMS logged; matched-energy nonspeech versus voiced reference; speech followed by silence padding without changing total reference length; same-speaker clips with different words; and only then a qualified Diffusers or official-BF16 comparison. Changes in latent interpolation or post-encoder zeros are diagnostic interventions, not production fixes. If two implementations differ, compare reference normalization, label presentation, block packing/time positions, weights, and backend before assigning cause to the framework itself.
For future readers, keep reference/fully_copy, label/latent, builder/sampler, and subject/frame distinct. The official reference guide, native ComfyUI guide, Ref2VA node documentation, Diffusers implementation, and public #16155 reproduction source cover those boundaries. The experiences in #64 and #76 are useful comparisons, but not established fixes.
Bottom line: the related public test gives a concrete reason to examine audible-reference conditioning values before blaming the complete WebUI or the three-LoRA stack. Your exact doubled-voice mechanism and retained-background mechanism remain unverified. Both goals are still worth pursuing with small, reversible, separately scored comparisons.