OncoLLM: adapting MedGemma 4B and 27B to oncology tasks with Unsloth

We fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology with Unsloth on a single DGX Spark. Report with DOI, models and dataset are public.

One finding you may want to look at (section 5): with Unsloth 2026.9.12, the 4B adapter (r=64, alpha=16) behaves as if scaled by ~0.5 instead of alpha/r = 0.25. Loaded with Transformers + PEFT at nominal scale, 0/4 identity probes pass; at 2x, 4/4. Merging with save_pretrained_merged at nominal alpha lost most of the tuning. We did not identify the cause in the code. Is this a known behavior?

Unsloth is credited in the model cards. Feel free to share it if useful. Thanks for the tooling.

See the repo : Grujowmi (ANTOINE Pierre)

See the report : OncoLLM: adapting MedGemma 4B and 27B to oncology tasks | Zenodo

Have fun to give us a feedback

For now, I tried to narrow this down using only publicly available materials:


Short answer: I could not reproduce a general 2Ă— LoRA scaling behavior here.

For ordinary LoRA, PEFT uses the usual lora_alpha / r scaling (PEFT LoRA docs, implementation). So with r=64, alpha=16, I would expect 0.25.

I tried a public surrogate using Unsloth’s MedGemma 4B mirror, with the same r=64, alpha=16, including a small real LoRA train → save → reload cycle. On an actually trained language-model q_proj, I measured an effective scale of about 0.250009:

  • live trained adapter: ~0.250009
  • stock PEFT reload: ~0.250009
  • Unsloth reload: ~0.250009

The saved LoRA tensors were also identical across the reload paths. On a fixed input, nominal x1 scaling was much closer to the live trained model than x2; for the Unsloth reload, x1 matched the live result exactly in this probe, while x2 moved the logits substantially.

So at least from this public reproduction, I would not treat ~0.5 as a general Unsloth/plain-LoRA convention.

I would still take your 0/4 at nominal → 4/4 at 2× result seriously as a behavioral observation. I just would not yet read it as proof that the LoRA layer itself was numerically applying 0.5. There are a few boundaries above and below the LoRA arithmetic that can produce that kind of behavioral threshold.

The cheapest check I can think of is to inspect one genuinely trained LoRA layer in the original environment.

same final input_ids / attention_mask?
|
+-- no
|   -> normalize the prompt/tokenizer path first
|
+-- yes
    |
    +-- one trained layer has effective scale ~0.50?
    |   |
    |   +-- yes
    |   |   -> strong lead for a real runtime scaling/arithmetic anomaly
    |   |
    |   +-- no, it is ~0.25
    |       |
    |       +-- raw and loaded LoRA A/B identical?
    |           |
    |           +-- no
    |           |   -> loading/materialization/state-dict branch
    |           |
    |           +-- yes
    |               -> compare exact base revision, logits,
    |                  dtype/backend and wrapper/model path
    |
    +-- only the merged checkpoint degrades?
        |
        +-- yes
            -> investigate merge precision/export separately

I would especially separate the live-scale question from the merge degradation question. In the public bf16 surrogate, I could reproduce substantial elementwise error when the LoRA delta was materialized into bf16 base weights even though the live LoRA layer itself was using the correct 0.25 scale. So those two observations do not necessarily need one common cause.

Public reproduction details

I used:

base:
  unsloth/medgemma-4b-it

base revision:
  9118034dee0620324cce8f2f6116aabd9b594134

GPU:
  NVIDIA L4

base dtype:
  bf16

LoRA:
  r = 64
  alpha = 16
  use_rslora = False
  nominal scale = 0.25

representative trained layer:
  base_model.model.model.language_model.layers.0.self_attn.q_proj

This was a small surrogate training run rather than just inserting a random adapter and checking its initial forward.

Both LoRA A and B changed during training.

For the selected trained language layer:

configured scale:
  0.25

live effective scale:
  0.25000867

effective / configured:
  1.00003468

After saving the adapter and reloading it through both stock PEFT and an Unsloth/FastModel path, the same layer gave essentially the same measurement.

The complete saved adapter tensor hash before merge was also identical across:

live trained adapter
stock PEFT reload
Unsloth reload

SHA-256:
82cdcd8fa49f89218d47f0445d7e0099435e87d832522fe5b40eedee7acca8f4

So in this surrogate I did not see save/reload changing A/B, changing alpha/r, or creating an Unsloth-only 2Ă— scale.

For a fixed text input, relative differences against the live x1 logits were:

stock PEFT reload, x1:
  0.0169823

stock PEFT reload, x2:
  0.408694

Unsloth reload, x1:
  0.0

Unsloth reload, x2:
  0.410495

stock x1 vs Unsloth x1:
  0.0169661

So there was a small whole-model difference between stock PEFT and the Unsloth execution path, despite identical adapter tensors and layer scaling, but it was much smaller than the effect of actually doubling the LoRA scale.

This is only a public surrogate. It does not reproduce the actual OncoLLM adapter or the exact original DGX Spark binary/runtime stack, so I would use it mainly to rule down a general Unsloth/FastModel x2 rule rather than to identify the original root cause.

One thing I would normalize before interpreting the identity probe as a scale measurement

I noticed that the Unsloth-side and stock-PEFT-side probes do not appear to construct the model input through exactly the same route.

For this kind of comparison I would try to make the final model inputs identical first:

input_ids
attention_mask

and, if they differ, inspect:

rendered prompt
tokenizer revision
chat template
BOS / EOS placement
add_generation_prompt
add_special_tokens

This is worth checking with Gemma-family tokenizers in particular.

There is a concrete Transformers example where:

tokenizer.apply_chat_template(..., tokenize=True)

and:

text = tokenizer.apply_chat_template(..., tokenize=False)
tokenizer.encode(text)

did not produce the same token sequence because the second route added another BOS token:

I am not saying that is the cause here. It just means that before treating:

0/4 at nominal scale
4/4 at 2Ă— scale

as a literal measurement of the LoRA coefficient, I would first make sure both sides consumed exactly the same token tensors.

Generation is nonlinear enough that a small upstream difference can turn into a very sharp pass/fail difference.

I would also log the exact resolved base repository and immutable revision on both paths. A user-facing model ID alone is not quite enough for a strict parity check when different loaders or model-ID remapping may be involved.

A small one-layer check that would distinguish the branches

The main idea is just to compare the adapter contribution produced by the real layer with the unscaled B(A(x)).

It does not need a benchmark or another full generation experiment.

For one representative layer that was genuinely trained:

import torch

model.eval()

adapter = "default"

name, layer = next(
    (name, module)
    for name, module in model.named_modules()
    if "language_model" in name
    and name.endswith("q_proj")
    and hasattr(module, "lora_A")
    and adapter in module.lora_A
)

A = layer.lora_A[adapter]
B = layer.lora_B[adapter]
drop = layer.lora_dropout[adapter]

configured = float(layer.scaling[adapter])

print("layer:", name)
print("r:", layer.r[adapter])
print("alpha:", layer.lora_alpha[adapter])
print("configured scaling:", configured)
print("A dtype:", A.weight.dtype)
print("B dtype:", B.weight.dtype)
print("base dtype:", layer.base_layer.weight.dtype)

# A deterministic arbitrary layer input is enough to test
# the arithmetic of this particular LoRA layer.
g = torch.Generator(device="cpu").manual_seed(1234)

x = torch.randn(
    2,
    3,
    A.in_features,
    generator=g,
    dtype=torch.float32,
).to(
    device=A.weight.device,
    dtype=layer.base_layer.weight.dtype,
)

with torch.no_grad():
    manual_unscaled = B(
        A(
            drop(x.to(A.weight.dtype))
        )
    ).float()

# Make the outer LoRA layer return only its adapter branch
# by zeroing the base-layer output.
calls = {"n": 0}

def zero_base(_module, _inputs, output):
    calls["n"] += 1
    return torch.zeros_like(output)

hook = layer.base_layer.register_forward_hook(zero_base)

try:
    with torch.no_grad():
        observed = layer(x).float()
finally:
    hook.remove()

den = (manual_unscaled * manual_unscaled).sum()

effective = (
    (observed * manual_unscaled).sum() / den
)

print("base hook calls:", calls["n"])
print("effective scale:", effective.item())
print(
    "effective / configured:",
    (effective / configured).item(),
)

For r=64, alpha=16, the interesting distinction would be approximately:

effective scale ~0.25
effective/configured ~1

versus:

effective scale ~0.50
effective/configured ~2

If the original Unsloth environment really gives the second result on a trained layer while stock PEFT gives the first, that would be much stronger bug evidence than a generation-level identity check.

If the layer itself says ~0.25, then the 2Ă—-looking behavior has to enter somewhere outside the basic LoRA arithmetic.

One caveat: if a patched/compiled path bypasses base_layer.forward, the hook count could be zero. In that case I would not interpret this exact snippet’s result; I would instrument the compiled path or compare the equivalent intermediate tensors instead.

Merge / bf16 precision looks like a separate problem

The merge side looks quite different to me.

PEFT’s normal standalone-model path is merge_and_unload(). There are also existing reports of live LoRA and bf16 merged behavior diverging, e.g. Unsloth #3091.

In the public surrogate I compared the expected LoRA delta for the same trained language q_proj with the delta actually visible after adding it into the bf16 base weight.

I got:

expected LoRA delta norm:
  0.1707390

observed bf16 weight-difference norm:
  0.1764309

norm ratio:
  1.03334

elementwise relative error:
  0.31521

fraction of entries where:
  expected delta != 0
  but stored bf16 weight difference == 0

  0.34016

The stock PEFT and Unsloth paths gave the same in-memory merge metrics in this experiment.

The important bit to me is that the total delta norm staying close does not mean the individual small corrections survive faithfully. In this layer, roughly 34% of positions where the expected LoRA update was nonzero became indistinguishable from zero after the bf16 weight update.

That looks like a plausible mechanism for a merged model to lose some tuning signal while the live adapter remains fine.

I would therefore keep these as two independent questions:

A. Is the live adapter arithmetic using the wrong scaling?

B. Does materializing that adapter into low-precision
   base weights lose enough small updates to change behavior?

PEFT’s own GPU merge tests are also useful context here. For quantized paths, they compare behavior in probability space and explicitly allow nontrivial numerical tolerance rather than assuming token-level or bit-exact equality:

So I would not use merged-model degradation by itself as evidence that the live runtime must have been using 2Ă— scaling.

Nearby upstream issues that make these boundaries worth checking

I found a few nearby issues. I do not think any of them proves the same root cause, but they are useful precedents for where configuration/export differences can enter.

FastModel configuration propagation

There was a real FastModel issue where use_rslora=True did not make it into the resulting PEFT configuration:

That is not the same as plain lora_alpha being doubled.

It does, however, make me think that printing the resulting runtime PEFT config / layer state, rather than only the arguments originally passed to get_peft_model, is a worthwhile sanity check.

FastModel + standard PEFT wrapper + export

There is also a separate report where the flow:

FastModel.from_pretrained(base)
    ->
PeftModel.from_pretrained(base, adapter)
    ->
save_pretrained_merged / save_pretrained_gguf

could route the Unsloth save method through a function still bound to the base model:

That is relevant to the broad symptom “live adapter works, exported model loses tuning”, but it does not explain a live 2× scale.

I would only pursue this branch if the actual merge/export loader sequence looks similar.

One small version/reproducibility note

There is one environment detail I would record, but I would not make it the main diagnosis.

The public OncoLLM dependency lock I inspected contains roughly:

unsloth       2026.9.12
unsloth_zoo   2026.9.8
transformers  5.18.0
trl           0.26.1

The package metadata I obtained for those Unsloth versions declared tighter ranges:

unsloth 2026.9.12:
  transformers <= 5.5.0
  trl <= 0.24.0

unsloth_zoo 2026.9.8:
  transformers <= 5.17.0

So the recorded environment is outside the dependency range that those packages declare.

I do not think that is enough to say the version combination caused this issue: I deliberately forced the public surrogate into the recorded combination and still did not reproduce 2Ă— scaling.

It is mostly a reproducibility breadcrumb.

There is also a small release-label nuance: PyPI’s unsloth 2026.9.12 page shows that wheel as uploaded later in September, so the version string alone may not identify exactly which source state was present on a particular calendar date.

If an exact reproduction becomes useful, I would record:

pip freeze
torch version/build
Unsloth package hashes or source commit
Unsloth Zoo version
Transformers / PEFT / TRL
CUDA stack
base-model commit
tokenizer commit

rather than relying on the Unsloth version label alone.

So my current read is:

  1. The reported 0/4 → 4/4 at 2× behavior is interesting and worth preserving as an observation.
  2. I could not reproduce a general Unsloth/MedGemma plain-LoRA 2Ă— scaling rule.
  3. A public real train/save/reload surrogate stayed at the expected ~0.25 scale.
  4. The merge degradation has a plausible independent bf16 precision mechanism, which I could reproduce at the weight-delta level.
  5. The highest-information next check in the original run is just one trained layer plus identical final input tensors — not another large training run.

If that one layer really reports ~0.5 in the original environment, I think that would be a very useful minimal reproduction to take upstream. If it reports ~0.25, then the search space gets much smaller: input/template/base identity first, then whole-model runtime/backend differences, while keeping merge precision on its own branch.

Thank you, this is a careful review, and several points are fair.

Our ~2× claim rests on generation-level probes, not on a layer-level measurement. The Unsloth and PEFT probes also build their inputs differently (Unsloth’s gemma-3 template rendered then tokenized without special tokens, versus the base tokenizer’s own template with tokenize=True), so 0/4 vs 4/4 is confounded. I will state this in the report.
The version note is correct: Unsloth was installed with --no-deps, so the environment sits outside the declared ranges (transformers 5.18.0, TRL 0.26.1). I will add it to the limitations. I also realized at the end that the NVIDIA DGX SPARK wasn’t up to date and that the performance monitoring utility was offering an update.
Merge degradation is a separate question from live scaling. Our weight check found the same bf16 effect you reproduced.

One observation your surrogate does not cover is our merge-scale sweep. The merges apply exactly the configured scale (checked weight by weight), yet the merge at 0.5 is the closest to the Unsloth adapter (identity restored, no significant prediction difference), while 0.25 loses most of the tuning and 0.75 and 1.0 drift again. If the live adapter ran at 0.25, I would expect 0.25 to match best. Your setup also differs from ours (L4 vs DGX Spark GB10, aarch64).

I no longer have access to the original machine, so I cannot run the direct check you propose (one trained layer, identical input_ids) in the original environment right now. Until it is done, I will present the result in the report as an observation consistent with an effective scale of about twice α/r, with the mechanism unidentified and not reproduced on other hardware, rather than as an established property of Unsloth.

Thank you sincerely for this feedback; this is my first fine-tuning job, and the first AI project into which I’ve poured so much heart.