I am having two major issues with MiniMax-H3 Ref2VA

I am not running the official MiniMax-H3 Diffusers demo directly.

I built a custom Gradio WebUI around the ComfyUI MiniMax-H3 implementation and communicate with ComfyUI through its API.

The WebUI loads a ComfyUI workflow JSON, modifies the API prompt programmatically, uploads the input/reference files to ComfyUI, and then submits the resulting prompt to the ComfyUI /prompt API.

The workflow is based on a MiniMax-H3 Ref2VA / Live Wallpaper style workflow and contains two MiniMaxH3ReferenceToVideo passes.

In my current graph:

  • first/intermediate MiniMaxH3ReferenceToVideo: node 400
  • final MiniMaxH3ReferenceToVideo: node 587

The final BasicGuider / sampler path consumes the conditioning from the second/final Ref2VA node.

I intentionally attach the reference audio only to the final Ref2VA pass.

The intermediate pass does not receive reference audio.

The relevant API structure generated by my WebUI is effectively:

LoadAudio(reference_audio)
    -> MiniMaxH3ReferenceToVideo.ref_audios.ref_audio_0

For the final Ref2VA node, the API prompt contains the equivalent of:

"ref_audios.ref_audio_0": ["<LoadAudio node id>", 0]

The first/intermediate MiniMaxH3ReferenceToVideo node does not have ref_audios.ref_audio_0.

This uses ComfyUI’s native MiniMax-H3 reference-audio implementation rather than manually generating audio embeddings in my own WebUI.

ComfyUI internally performs approximately:

waveform = audio["waveform"]
sr = audio["sample_rate"]

if sr != audio_vae.audio_sample_rate:
    waveform = comfy.audio.resample(
        waveform,
        sr,
        audio_vae.audio_sample_rate
    )

z = audio_vae.encode(
    waveform[:1].movedim(1, -1)
)

and then puts the resulting audio latent into MiniMax-H3 conditioning as a reference block.

So the reference audio is not merely mentioned in the text prompt. It is actually passed through ref_audios and encoded by the MiniMax-H3 Audio VAE.


Models used

Main Ref2VA transformer

minimax_h3_ref2va_pruned_int8_convrot.safetensors

Loaded as the MiniMax-H3 Ref2VA diffusion model in ComfyUI.

Text / multimodal encoder

qwen3vl_32b_heretic_minimax_h3_nvfp4.safetensors

Audio VAE

minimax_h3_audio_vae_fp32.safetensors

This is loaded through ComfyUI’s VAE loader and is used for both MiniMax-H3 audio handling and reference-audio encoding.

Video VAE

minimax_h3_video_vae_fp16.safetensors

The Audio VAE is the ComfyUI-formatted MiniMax-H3 Audio VAE, not a separate model from models/audio_encoders.

My models/audio_encoders directory is essentially empty except for the placeholder file. MiniMax-H3 reference audio is going through the Audio VAE path.


LoRAs used in the problematic generation

I used three custom LoRAs simultaneously:

AfterMidnight_ref2va_h3_sexytime_rank64_v1.1.safetensors
AfterMidnight_ref2va_h3_softer_rank64_v1.safetensors
Astro%20nsfw.safetensors

Strength:

1.0
1.0
1.0

The WebUI inserts these sequentially into the model chain before sampling rather than generating separate outputs and merging them afterward.

In other words, I am not doing an ensemble of three separately rendered videos.

The model is effectively passed through the LoRA chain and then sampled once.

The original Live Wallpaper workflow also contains its own internal workflow/model modification path, so this is not a completely minimal official Ref2VA graph.


Custom WebUI behavior

The WebUI performs the following operations before submitting the ComfyUI API prompt:

  1. Loads the original workflow.
  2. Converts it into an API prompt.
  3. Replaces input image/video/audio nodes with uploaded files.
  4. Appends the user’s requested edit/action to the MiniMax-H3 prompt.
  5. Applies selected custom LoRAs sequentially.
  6. Disables the WebUI’s automatic prompt rewriting where possible.
  7. Locates the final MiniMaxH3ReferenceToVideo node.
  8. Removes reference-audio inputs from intermediate Ref2VA nodes.
  9. Connects the uploaded reference audio only to the final Ref2VA node using:
ref_audios.ref_audio_0
  1. Submits the graph to ComfyUI through its normal API.

The WebUI itself does not synthesize the voice.

Actual audio generation is still performed by MiniMax-H3 / ComfyUI.


Problem 1: reference voice is very abnormal

The reference audio contains one speaker.

However, the generated result often sounds as if two voices / two speakers are speaking simultaneously.

It is difficult to describe exactly, but it sounds like overlapping vocal identities rather than one clean speaker.

The generated voice is also substantially different from the voice in the supplied reference audio.

This is not a small pitch difference.

The speaker identity itself feels different.

At times the generated audio is also difficult to understand and resembles speech-like vocalization rather than clean intelligible dialogue.

The strange part is that the Audio VAE clearly appears to be functioning:

  • audio is generated,
  • it is not completely silent,
  • it is not simply white noise,
  • it often has speech-like prosody,

but speaker identity and intelligibility can be severely wrong.

This makes me suspect one of the following:

  • reference-audio conditioning problem,
  • interaction between Ref2VA and the INT8 ConvRot transformer,
  • incorrect audio-latent alignment,
  • Audio VAE preprocessing issue,
  • interference from stacked LoRAs,
  • or interference from the two-pass workflow.

Problem 2: scene/environment instructions are mostly ignored

I explicitly requested a different environment.

For example, I specified:

basement

as the intended setting/background.

However, the generated result did not meaningfully create a basement environment.

Instead, the model mostly preserved the background of the source/reference image.

The result looked much closer to:

animate the existing photograph for approximately 10 seconds

than:

keep the person identity but regenerate the environment according to the requested scene.

The character may move, but the environment prompt is effectively ignored.

This also happened with other visual editing instructions.

Some simple pose/action instructions can work, but larger scene or appearance changes are much less reliable.


Prompt preservation logic

Earlier versions of my WebUI used a strong preservation instruction similar to:

fully preserve identity, face, body appearance, clothing, and visual continuity

I suspected this could directly conflict with scene/edit instructions.

I have since changed this behavior.

The current version is intended to preserve primarily:

  • face,
  • identity,
  • subject continuity,

while allowing the prompt to modify:

  • clothing,
  • shoes,
  • pose,
  • body position,
  • environment,
  • background.

I also added an explicit instruction to keep exactly one instance of the subject to reduce duplicate-person generation.

Despite this, large environmental changes still appear weak compared with preservation of the source image.


Reference-image strength

The original workflow used a very strong reference-image setting.

I previously used:

ref_image_size = max

I have changed the final Ref2VA pass to:

ref_image_size = match

because I suspected max was causing excessive visual anchoring to the source image.

The scene-following problem still needs further testing after this change.


Audio VAE memory modification I tested

Because my GPU has 12 GB VRAM, I previously experimented with manually unloading the Audio VAE immediately after reference-audio encoding.

I tried calling the ComfyUI model patcher’s unload path after:

audio_vae.encode(...)

This proved too aggressive, so I removed that modification.

The current version does not manually call:

patcher.unpatch_model(...)

after reference-audio encoding.

I now leave model/offload management to ComfyUI.

So the abnormal audio described above should not be interpreted as being intentionally decoded through a separately modified Audio VAE implementation.


Hardware

GPU VRAM:

12 GB

The system has substantially more normal RAM available.

Because of the 12 GB VRAM limit, I am using the pruned INT8 ConvRot Ref2VA transformer instead of the full original BF16 transformer.

This is also why I am interested in whether the quantized/pruned Ref2VA model may behave differently from the original official transformer for reference audio.


Questions

I would appreciate clarification on the following.

1. How closely should Ref2VA preserve speaker identity?

When ref_audios.ref_audio_0 is supplied, should the generated speech normally sound recognizably like the supplied speaker?

Or is the reference audio intended mainly for prosody/style rather than speaker identity?

2. Has anyone seen a “two speakers at once” effect?

Is there a known failure mode where reference-audio conditioning produces something that sounds like two overlapping speakers or mixed voice identities?

3. Could stacked LoRAs cause this?

Could these three LoRAs at strength 1.0 each:

AfterMidnight_ref2va_h3_sexytime_rank64_v1.1
AfterMidnight_ref2va_h3_softer_rank64_v1
Astro

interfere with MiniMax-H3 audio/reference conditioning?

In particular, can a visual/video LoRA modify enough of the transformer to damage speaker conditioning or audio generation?

4. Is strength 1.0 + 1.0 + 1.0 excessive for Ref2VA?

Would a lower value such as approximately 0.5-0.7 per LoRA generally be safer when stacking several MiniMax-H3 LoRAs?

5. Could the pruned INT8 ConvRot Ref2VA checkpoint be responsible?

I am currently using:

minimax_h3_ref2va_pruned_int8_convrot.safetensors

rather than the original large official Ref2VA transformer shards.

Could pruning or INT8 ConvRot quantization noticeably degrade the audio_patch_proj / reference-audio path or audio-video conditioning?

6. Does MiniMax-H3 require reference audio on both Ref2VA passes?

My workflow contains two MiniMaxH3ReferenceToVideo nodes.

Currently:

intermediate pass: no reference audio
final pass: reference audio enabled

Is this valid?

Or should reference audio be supplied to every Ref2VA conditioning stage in a two-pass workflow?

7. Why might environment instructions be ignored?

Is Ref2VA expected to strongly preserve the entire source-image background?

If I want:

same person / same identity
different environment

is there a recommended reference strength, ref_image_size, conditioning strategy, or workflow structure?

8. Is ref_image_size="max"

For now, I tried testing a similar reference-audio issue using publicly available materials:


Short answers to your eight questions

I would treat your goals as two separate conditioning problems, not as a reason to discard your WebUI or workflow: (a) preserve a recognizable speaker while generating new, correctly timed dialogue without an extra voice, and (b) preserve the person while replacing the original scene. I reproduced a related reference-audio anomaly using another person’s public example, but not your exact overlapping-voices output. The distinction matters throughout.

1. How closely should Ref2VA reproduce the speaker?

Using an audio clip to guide the speaker’s timbre and delivery is a documented Ref2VA use case. The official full-reference format distinguishes reference (use characteristics without directly copying the signal) from fully_copy (reuse the source audio). It does not promise a particular numerical voice-similarity score or perfect identity for every newly scripted line. A recognizable voice is a reasonable goal; it is useful to judge voice resemblance, correct words, intelligibility, speech timing, and unwanted additional voices separately. Success on one is not proof of success on the others.

2. Is the apparent “two speakers at once” behavior known?

There are neighboring reports: ComfyUI #16155 describes unexpected speech and timing errors with an audio reference; MiniMax-H3 discussions #64 and #76 include reports of reference material intruding into dialogue or clipped/gibberish vocalization; MiniMax-H3 #17 concerns voice binding between characters. Those are related symptoms, not proof of the same bug. Your “two voices” might mean genuinely simultaneous speakers, an extra utterance, speaker drift, or an echo/mixed track. Without your generated audio I cannot decide which.

In a separate test using #16155’s public assets, extra words were recognized in 3/3 seeds with audible reference audio, versus 0/3 with equally long silence or zeroed audio latents while retaining the reference block. This points toward the encoded reference values, not merely the presence of a slot. It neither reproduces your overlapping voices nor identifies a root cause; the controls and limitations are below.

3. Could three LoRAs interfere with the audio?

Yes, potentially: a “visual” adapter may alter shared audiovisual layers, depending on the modules it targets. Your three LoRAs appear to be sequential patches before the final sample, not three separate renders. Compare one fixed-seed run with all custom LoRAs off; if the problem clears, restore them individually. My related anomaly occurred without LoRAs or Heretic, so those are not prerequisites for that behavior, but they could still contribute in your pipeline.

4. Is strength 1.0 on each LoRA too much; should it be 0.5–0.7?

Three 1.0 settings are worth testing, but they do not add up to a meaningful global “3.0 strength.” Effects depend on target layers and adapter interactions. Nor is 0.5–0.7 a universal safe range. Try all off → one at a time → necessary combinations, then sweep strengths only where a difference appears; this preserves your visual goal.

5. Could the pruned INT8 ConvRot checkpoint be responsible?

Still an open variable. My runs used minimax_h3_ref2va_pruned_int8_convrot.safetensors and produced both clean ASR-control transcripts and reference-dependent extras: INT8 does not uniformly break dialogue, but might affect this failure. No matched official-BF16 test exists. With 12 GB VRAM, check the graph and one-seed contrasts first. The model card and native ComfyUI guide document the distinct deployment artifacts; successful loading alone does not establish output parity.

6. Must the audio be attached to both Ref2VA nodes (400 and 587)?

Not necessarily. MiniMaxH3ReferenceToVideo builds conditioning and an initial audio-video latent; it is not a sampler. Your plan to attach audio to final node 587 may be sound. Check the JSON actually submitted to /prompt after Gradio rewrites it: trace both final BasicGuider conditioning and sampler latent, and establish what 400 still supplies. Adding audio to 400 without that trace could introduce a new constraint. The native implementation and node docs are useful references.

7. Why does the old background persist despite requesting a basement?

A full reference picture encodes both the person and the old scene. If a first-frame/guide anchor survives in the actual graph, it may impose still more scene continuity. Distinguish identity reference from frame preservation: compare the full picture with a subject-focused crop or background-removed reference, keeping the target basement/action/seed fixed. You have already softened preservation wording, added a single-subject constraint, and switched max to match; their outcome is not established. I did not reproduce this background issue.

8. Is ref_image_size="max" …?

The public post ends at this heading, so I will not invent the missing question. ref_image_size controls image preprocessing, not a linear reference-strength setting. The official node docs explain that match scales references toward the output pixel area, while max keeps greater detail (up to the 2048px short-edge reference convention), at potentially much higher token/compute cost. Neither is a dedicated background switch; the result after your match change is unknown.

Suggested default route: preserve the features, isolate the boundary

  1. Freeze what actually runs. Save one failing run’s final API JSON after file uploads, prompt edits, LoRA chaining, and the WebUI’s input rewrites. Follow the final guider’s conditioning and sampler’s latent back through 587, 400, and the media loaders. This is mostly a static check and costs no render. Do not assume the original workflow screenshot equals the submitted graph.
  2. Use one paired voice test, not a rebuild. At the same seed and settings, compare the current audible reference with all custom LoRAs off. If the bad voice remains, prioritize reference roles, native Audio VAE conditioning, and model/encoder differences. If it clears, reintroduce adapters one by one. Continue to use the audible reference for the target result: silence/zero latent are diagnostic controls, not voice cloning.
  3. If extra utterances remain, verify the reference contract. Bind <Audio N> to (S1), use reference rather than fully_copy for new dialogue, keep the sample clean, and specify the exact line. For timing, connect speech to an on-screen action as well as a shot interval; absolute timestamps alone do not guarantee onset. A shorter clip is optional, not a universal fix: cropping still left extras in 2/3 seeds.
  4. Handle the background independently. Start from your already-updated match setting; compare full-scene versus subject-focused image. Check whether anything actually anchors the first frame or reuses an old visual guide. Score person identity and new location separately, even if the LoRA stack remains unchanged.
  5. Escalate when a cheaper comparison identifies a boundary. Standard Qwen versus Heretic, alternative checkpoints, matched official BF16, or ComfyUI versus Diffusers can be useful, but each adds downloads, VRAM demand, and possible unrelated numerical differences. They are not prerequisites for the first two checks.

The goal remains new speech in the target voice and the same person in a new room.

First checks: what to observe before another expensive render

Useful optional evidence: one failing submitted /prompt JSON (private paths removed), a short generated-audio excerpt, the intended line, seed, reference duration, and checkpoint/adapter names. These are not prerequisites for the suggestions below.

No-render audio check: If the voice sounds doubled, compare ComfyUI’s raw decoded audio (if retained) with the final Gradio/WebUI-exported or muxed audio from the same run. Check left and right channels separately before downmixing. Only after export? Inspect muxing, duplicated tracks, and post-processing first. Already in the raw audio? Focus on generation, reference binding, or decoding. If the raw track was not saved, mark this comparison as unavailable instead of inferring an export bug; when both tracks exist, no new inference is needed.

Gradio workflow template
  -> file substitutions / prompt edits / adapter chain
  -> actual payload sent to POST /prompt
     -> LoadAudio -> 587.ref_audios.ref_audio_0
     -> 587.positive -> final BasicGuider (?)
     -> 587.LATENT   -> final sampler input (?)
     -> 400 outputs -> which actual consumers, if any?
     -> final H3 model / guider / sampler -> decode / output

The question marks are edges to check, not accusations that they are wrong. ref_audios.ref_audio_0 is a zero-based API field, while model text uses the one-based label <Audio 1>; that apparent mismatch is normal. Other enabled audio-bearing reference media can affect label ordering, so inspect the packed/reference list rather than treating numbering as a universal equivalence. To classify “double voice,” distinguish: overlapping independent speakers, extra speech before/after the line, timbre drift, and duplicated/reverberant audio. These suggest different causes; ASR alone cannot distinguish them.

A related public reproduction, in brief

I used the public image and voice WAV from someone else’s ComfyUI issue #16155, where the requested line was This park sucks. around 2 seconds. This is not a reproduction of your particular reference, output, multi-node graph, or two-simultaneous-voices claim. My Colab setup used native ComfyUI, pruned INT8 H3 Ref2VA, a standard (non-Heretic) Qwen, no LoRAs, 512×288, 124 frames at 24 fps (~5.17 s), 20 sampling steps, and three deliberately reused seeds (4242, 5151, 6262).

The decisive panel kept the reference block in place while changing the encoded values:

Reference input Extra words recognized beyond requested line Interpretation
Original audible 7.05 s WAV 3/3 seeds, both Whisper models Related anomaly reproduced in this setup
Same-duration time-reversed WAV 2/3 agreed; 1/3 ASR disagreement Ordinary forward-intelligible English is not clearly required
Same-duration PCM silence, Audio-VAE encoded 0/3, both ASR models Reference block still exists; encoded silence is not a zero latent
Original reference latent zeroed after encoding 0/3, both ASR models Same reference block shape, but an artificial out-of-distribution control

Cropping the reference still left extras in 2/3 seeds. Two Whisper versions agreed on most observations but are not human listening tests; neither words alone nor waveform activity establish concurrent speakers, literal copying, or restored voice identity.

Interpretation: audio-derived conditioning values matter in this small panel, while the responsible feature and root cause remain unknown. The expanded controls follow.

A. Reproduction protocol, per-condition results, and evidence limits

Scope and controls

The #16155 author reported both anomalies and successful seeds. I used their public image and WAV, not your audio or WebUI graph. Each panel reuses seeds 4242, 5151, 6262 for paired comparisons; those are not population samples for an H3 failure-rate estimate.

Reproducibility capsule (local Colab tests, not a public benchmark):

ComfyUI: a4b5a045e56fc334903db8457b728b64e006119c
HF model repository revision recorded: e5eb578a89295337b8ff433a035929ce0279e0b6
Checkpoint: minimax_h3_ref2va_pruned_int8_convrot.safetensors
Conditioner: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (non-Heretic)
Audio VAE: minimax_h3_audio_vae_fp32.safetensors
No LoRA; 512x288; 124 frames at 24 fps; 20 steps
Seeds: 4242 / 5151 / 6262
Input WAV: 7.05 s, stereo, 48 kHz
WAV SHA-256: eb0e6e869cebe54325388affda68aa9f6bcf892365b3d169f60fee2257946f49

The repository revision is recorded in the downloaded artifact metadata; it is not a published hash of every loaded tensor. Output FLACs, per-run logs, and ASR tables are retained locally, not hosted at a public results URL. The test also uses a newer pinned ComfyUI core-node reconstruction rather than the #16155 author’s exact environment.

Panel Intervention Both Whisper models’ extra-text findings Confound to remember
E1: 6 outputs Reference absent vs. original present 0/3 vs. 3/3 The Qwen label and packed audio layout also change
E2: 9 outputs Original 7.05 s / label only / cropped ~5.17 s 3/3 / 0/3 / 2/3 Label-only removes the audio block; crop changes its length
E3: 12 outputs Original / time-reversed / PCM silence / post-VAE zero latent 3/3 / 2/3 + one disputed / 0/3 / 0/3 Fixed block shape, but several acoustic properties still differ

E2’s “label only” retained <Audio 1> but removed the audio reference block; attention/time positions could change. E3 resolved that principal confound by retaining a [1, 32, 2, 282] reference block, original duration, prompt, and seeds in all 12 runs. Same shape is not proof that every internal tensor operation is identical.

Why silence and zero latent must not be conflated

PCM silence was a zero waveform of the original sample duration encoded normally by Audio VAE; the resulting latent was not zero (mean absolute magnitude ~0.461, versus ~0.394 for the original). Post-VAE zero latent started with the encoded original, then used an experimental, local ComfyUI patch to apply zeros_like to its reference representation. This kept a block with the same dimensions but supplied an artificial value. Neither supplies a usable example of the intended person’s voice. Indeed, the silence latent’s larger mean absolute magnitude argues against “lower latent norm” as a sufficient explanation.

The reversed WAV kept sample count and broad amplitude properties but reversed phoneme progression and temporal structure. Both Whisper decoders detected extras at 4242 and 5151; for 6262, base.en found extra text and small.en did not. This does not prove that the source’s linguistic meaning is irrelevant—only that intact, forward-spoken English does not appear necessary for the two agreed anomalous outputs.

Measurement boundaries

Whisper base.en and small.en are two versions of one ASR family, not independent listeners. Some reported extra suffix words differed between them; neither can prove exactly what a person would hear, let alone simultaneous distinct speakers or verbatim copying from the reference. Early RMS/audio activity likewise measures waveform energy, not linguistic onset. For E1, the first substantial activity was approximately 0.1–0.3 seconds with the original reference versus 1.6–2.1 seconds without it, despite an instruction to speak near 2 seconds. E3 silent/zero controls also showed different volume peaks, so “no extra ASR words” is not equivalent to proven sound quality or speaker fidelity.

Established here: changing values inside a same-shape reference block changes extra-speech-like outcomes. Not established: which acoustic/linguistic feature drives this, whether BF16/Diffusers agrees, or whether your two overlapping voices share the mechanism. One image/WAV pair and three selected seeds are not a general benchmark. The capsule above identifies the relevant artifact revisions and WAV hash; exact API graph, prompt, outputs, and listening/ASR procedure also matter.

B. How reference labels, sound latents, and timing interact

Two routes for one reference

The inspected ComfyUI path conceptually separates the reference’s description/label from its waveform encoding:

standalone reference audio
   +--> Qwen-side presentation: <Audio 1>, speaker (S1), role
   +--> PCM -> 32 kHz stereo -> H3 AudioVAE.encode
          -> packed reference-audio latent and time positions
          -> shared H3 audiovisual transformer -> new audio/video

Qwen does not receive raw PCM samples as though they were text. The Audio VAE supplies numerical reference data; the Qwen-side presentation tells the model how to use them. A prompt saying “voice timbre, not copied audio” expresses a desired relationship; it is not a guaranteed filter that removes lexical or timing information from the sound latents.

The official full-reference guide distinguishes <Subject N>, <Picture N>, <Video N>, <Audio N> and reference from fully_copy. Speaker binding should be explicit when appropriate; numbering of standalone audio and enabled audio tracks extracted from videos needs attention. The model’s <Audio 1> is one-based, whereas ref_audios.ref_audio_0 is a zero-based API input socket.

For example, a role fragment, not a guaranteed complete prompt:

subject_definitions:
<Subject 1> (S1) is the person referenced by <Picture 1>.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

summary:
[reference generation + audio reference] A new scene with <Subject 1>
speaking newly scripted words in the referenced voice style.

retention_analysis:
<Subject 1>: fully_preserved - preserve identity, not the old location.
<Audio 1>: reference - follow speaker timbre without copying the signal.

... in the target shot ...
<Subject 1> (S1) says: <d>[English] The new sentence.</d>

This fragment is intentionally incomplete. In a full shot, anchor the line to a visible event (e.g., when the person reaches the doorway) as well as describing the shot’s timing and quiet intervals. The ComfyUI prompt guide recommends visible-event anchors in multi-speaker reference cases, rather than expecting absolute timecodes to bind speech reliably. This may help, but does not guarantee the intended onset or eliminate reference-induced vocalization.

The audio length difference is real, but not a full explanation

The H3 Audio VAE operates at 32 kHz with 800 samples per latent step (40 steps/s). The 7.05-second reference yields 282 latent steps, compared with ~207 for the 5.17-second output. In inspected Diffusers normalization, input reference sound is truncated to generated duration before resampling; our pinned ComfyUI path retained the longer source. This implementation difference deserves inspection, but after cutting to output duration, E2 still showed extras at 2/3 seeds. Thus duration truncation is neither a proven cause nor a universal fix. The separate Audio VAE crop issue #15970 concerns an edge case; our original resampled WAV had exactly 225,600 = 282 × 800 samples, so fractional-hop loss is a weak explanation for multi-second extra dialogue here.

If the voice source is dense or noisy, a clean single-speaker excerpt, explicit (S1) binding, and constrained new dialogue are reasonable practical trials. People describe mixed success with such measures in HF #64 and #76. They should not be represented as established repairs. To identify a deeper mechanism, an optional equal-duration test can vary amplitude, matched-energy nonspeech, or where within the same reference length speech occurs. PCM amplitude and latent magnitude must both be logged because the VAE is not a simple linear gain control.

C. Two Ref2VA nodes: trace the final API graph, not the node count

The native Ref2VA node assembles conditioning plus an initial audio-video latent. A separate guider/sampler performs the diffusion steps. Therefore 400 and 587 can be two builders, alternative branches, or inputs into a genuine multi-stage graph; their names alone do not define the execution stages.

Inspect the rewritten payload: class_type and inputs of 400/587; the final BasicGuider’s positive source; the final sampler’s LATENT source; the patched model; 587’s clip, video VAE, Audio VAE, and ref_audios.ref_audio_0; plus any first-frame/guide path. Even if BasicGuider uses 587, the final sampler may get its latent elsewhere. That can be intentional, but should be explicit.

"587": {
  "class_type": "MiniMaxH3ReferenceToVideo",
  "inputs": {"ref_audios.ref_audio_0": ["LOAD_AUDIO_ID", 0]}
}

This is a fragment illustrating the key, not a runnable full node. In ComfyUI’s API representation, that dotted input name is normal. The important questions are whether the linked producer exists, which output it returns, which media are packed before it, and what downstream nodes consume it. A template may be correct while the WebUI’s later JSON rewrite silently points a final edge elsewhere.

If 587 alone supplies both final conditioning and its intended latent: final-only audio can be perfectly sensible; test reference semantics rather than adding it to 400. If 400 supplies a visual guide or latent that survives into the final sampler: inspect whether it anchors the original scene; this may matter for your background issue without making 400 an audio problem. If both branches really have separate samplers: decide explicitly which stage should use the voice reference; that is a stage-design choice, not a universal requirement. If an edge points to stale inputs after WebUI rewriting: repair that specific link before judging H3’s model behavior. This static graph audit has a very high information-to-compute ratio.

D. LoRAs, Heretic Qwen, the INT8 checkpoint, and 12 GB VRAM

A “visual” LoRA may touch shared audiovisual layers; presence alone is not evidence of causation. Three 1.0 settings are loader scales, not a summed global strength. Use none → each alone → necessary combinations → targeted strength sweep, stopping when the result is informative. Fix WAV, reference ordering, dialogue, seed, sampler, and final API graph. Score intended visual quality as well as speech to avoid “fixing” the voice by losing the desired look.

Your Heretic Qwen3-VL variant differs from the standard Qwen used in my reproduction. Since the latter showed extra reference-dependent speech too, Heretic is not necessary for the related failure. Still, an encoder can influence prompt/reference interpretation. Compare it separately only if an exact checkpoint/processor-compatible substitute is available; successful loading is not proof that two multimodal presentation paths are semantically identical.

The pruned INT8 ConvRot model generated both control and anomalous outputs; this says nothing about BF16 parity. Comparing precisions meaningfully would require matching revision, references, schedule, seed/noise, and backend as far as practical; differences are otherwise not uniquely due to quantization. With 12 GB, leave this resource-intensive branch until cheaper comparisons are exhausted.

You also mentioned trying to manually unpatch/unload the Audio VAE after reference encoding and then removing that change. Treat that as already reverted. Extra model-manager intervention can make state handling harder to reason about, but it is not evidence of the current symptom.

E. Keeping the person while changing the environment

The reference can encode who the person is and where the original photo was taken. They are not automatically separated by the word “subject.” For your basement scene, the key distinction is between a character/identity reference and a frame anchor that must resemble the supplied image. A first-frame anchor (if present) can reinforce that bias.

A low-cost scene comparison holds the person, action, new basement description, seed, and model constant, while swapping full original picture for subject-focused crop/background-removed picture. Then ask separately: is the face recognizably the same; does the basement appear; is the pose or old lighting copied; does quality degrade? The crop may remove useful hair/clothing context, so select it to retain the identity cues that matter. If a visual guide or retained first frame is actually connected, test that constraint rather than endlessly rewriting the text.

Describe the new basement positively: walls, floor, lighting, objects, camera, position, and action; preserve facial identity without asking for the old composition. Your relaxed wording and match setting are already the baseline, not untried remedies. max may later improve facial detail, but does not substitute for subject-only reference selection. This branch is inferred from the conditioning design, not reproduced on your image.

F. Symptom decision tree, further research, and useful links
What exactly is wrong in the output audio?
|
+-- Two clearly simultaneous voices
|     -> compare raw decoded stereo with final WebUI/exported audio;
|        then inspect speaker binding if overlap is already in raw audio.
+-- One extra utterance, especially early or after the scripted line
|     -> closest to public #16155; examine reference audio values,
|        label order, dialogue timing, and a fixed-seed reference A/B.
+-- The speaker identity changes during otherwise correct words
|     -> compare voice resemblance separately from transcript accuracy;
|        then test adapter/encoder/reference variants one at a time.
+-- Echo, doubled channel, or post-export duplication
|     -> separate waveform generation from mixing/decoding/export.
+-- Unintelligible sounds but no clearly separate speaker
      -> listen; log activity/peaks; do not infer exact words from ASR.

For a deeper optional investigation, high-information experiments are: fixed-duration gain sweep with PCM and latent RMS logged; matched-energy nonspeech versus voiced reference; speech followed by silence padding without changing total reference length; same-speaker clips with different words; and only then a qualified Diffusers or official-BF16 comparison. Changes in latent interpolation or post-encoder zeros are diagnostic interventions, not production fixes. If two implementations differ, compare reference normalization, label presentation, block packing/time positions, weights, and backend before assigning cause to the framework itself.

For future readers, keep reference/fully_copy, label/latent, builder/sampler, and subject/frame distinct. The official reference guide, native ComfyUI guide, Ref2VA node documentation, Diffusers implementation, and public #16155 reproduction source cover those boundaries. The experiences in #64 and #76 are useful comparisons, but not established fixes.

Bottom line: the related public test gives a concrete reason to examine audible-reference conditioning values before blaming the complete WebUI or the three-LoRA stack. Your exact doubled-voice mechanism and retained-background mechanism remain unverified. Both goals are still worth pursuing with small, reversible, separately scored comparisons.