For now, I ran some measurements on an L4:
Since the current measurement archive is still T4-only, I focused mostly on your “does the estimate match measured GPU usage?” question and tried a small second-GPU panel on a Colab NVIDIA L4.
I pinned FitCheck to this commit and tested BF16 with:
- Qwen2.5-Coder-1.5B and TinyLlama-1.1B
- NF4 QLoRA and ordinary non-quantized LoRA
- eager attention and real
flash_attention_2
- batch/sequence variations, including a same-session 2×2 grid
Across three sessions I got 17 successful measurement rows covering 12 unique configurations. I used repeats only as repeatability checks, not as extra independent samples.
The short version is:
| subset |
tensor-tier error |
full-process error |
| Qwen anchor cases |
-0.44% to -0.08% |
-2.72% to +2.20% |
| TinyLlama eager, 2×2 grid |
-0.95% to +2.83% |
-7.83% to +9.06% |
| TinyLlama real FA2, 2×2 grid |
-0.91% to -0.75% |
-8.61% to +1.87% |
Here, signed error is (predicted - measured) / measured, so negative means under-prediction.
So, at least in this small L4/BF16 panel, the tensor accounting transferred quite well, including the real FlashAttention-2 path. The larger residual showed up mainly in the full-process tier instead.
That seems consistent with the separation already present in FitCheck: the physical tensor terms and the card/runtime overhead are different problems. At the tested commit, L4 has no fitted overhead profile, so it falls back to the unmeasured default profile. In these measurements, that residual was shape-dependent and went in both directions, rather than looking like one simple constant offset.
If I were choosing one next step, I would probably use these as second-card calibration/holdout candidates before adding another random model. I would still keep some rows out of the fit rather than fitting and grading on the same L4 points.
Also, on the presentation side: the component breakdown was clear enough for me to follow. In practice, the tensor/process split was especially useful because it made the location of the larger miss much easier to see.
If useful, I can also provide the raw measure.py --json rows, environment capture, and the external process-memory traces.
Test setup and full measurements
Environment
The effective environment was the same across the three probe sessions:
- GPU: NVIDIA L4 (
sm_89)
- reported device memory: 23,034 MiB
- NVIDIA driver: 580.82.07
- Python: 3.13.15
- PyTorch: 2.11.0+cu128
- CUDA reported by PyTorch: 12.8
- Transformers: 5.17.0
- PEFT: 0.20.0
- bitsandbytes: 0.50.2
- Accelerate: 1.14.0
huggingface_hub: 1.29.0
- FitCheck source:
4da83a58f95864003734f5e9fd92af1756c0d685
The real FA2 runs used flash-attn 2.8.3 and actually selected flash_attention_2, rather than the T4 SDPA memory-efficient stand-in.
That path is within the upstream FlashAttention-2 CUDA support envelope: the official implementation lists Ampere/Ada/Hopper and fp16/bf16 as supported.
Qwen anchor
Model: Qwen/Qwen2.5-Coder-1.5B-Instruct
| configuration |
tensor error |
process error |
| NF4, BF16, eager, bs=2, seq=1024 |
-0.294% |
+0.319% |
| NF4, BF16, eager, bs=2, seq=2048 |
-0.442% |
-2.723% |
| NF4, BF16, real FA2, bs=2, seq=1024 |
-0.177% |
+2.197% |
| non-quantized LoRA, BF16, eager, bs=2, seq=1024 |
-0.080% |
+0.112% |
The QLoRA rows used rank 32, the standard [q,k,v,o] target set, AdamW, FP32 optimizer states, and gradient checkpointing.
One representative equivalent command is:
python scripts/measure.py Qwen/Qwen2.5-Coder-1.5B-Instruct \
--gpu l4 \
--qlora \
--precision bf16 \
--lora-r 32 \
--batch-size 2 \
--seq-len 1024 \
--optimizer adamw \
--json
and the FA2 version adds:
--flash-attn
TinyLlama diagnostic grid
Model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
I used TinyLlama because its smaller vocabulary makes the eager attention/layer hump much easier to expose as sequence length grows. The final probe ran the complete batch ∈ {1,2} × seq ∈ {1024,2048} grid in one session, under both eager and real FA2.
Common settings:
- NF4 QLoRA
- BF16 compute
- LoRA rank 32
- standard
[q,k,v,o] targets
- AdamW / FP32 optimizer state
- gradient checkpointing
- no double quantization
Eager
| batch |
seq |
tensors predicted |
tensors measured |
tensor error |
process predicted |
process measured |
process error |
| 1 |
1024 |
2003.1 MiB |
2019.4 MiB |
-0.81% |
2596.4 MiB |
2636 MiB |
-1.50% |
| 2 |
1024 |
2848.6 MiB |
2875.8 MiB |
-0.95% |
3484.2 MiB |
3780 MiB |
-7.83% |
| 1 |
2048 |
4000.6 MiB |
3908.2 MiB |
+2.36% |
4693.8 MiB |
4304 MiB |
+9.06% |
| 2 |
2048 |
6843.6 MiB |
6655.0 MiB |
+2.83% |
7678.9 MiB |
7888 MiB |
-2.65% |
Real FlashAttention-2
| batch |
seq |
tensors predicted |
tensors measured |
tensor error |
process predicted |
process measured |
process error |
| 1 |
1024 |
1833.6 MiB |
1847.6 MiB |
-0.76% |
2418.4 MiB |
2374 MiB |
+1.87% |
| 2 |
1024 |
2509.6 MiB |
2528.5 MiB |
-0.75% |
3128.2 MiB |
3140 MiB |
-0.38% |
| 1 |
2048 |
2509.6 MiB |
2529.0 MiB |
-0.77% |
3128.2 MiB |
3140 MiB |
-0.38% |
| 2 |
2048 |
3861.6 MiB |
3897.0 MiB |
-0.91% |
4547.8 MiB |
4976 MiB |
-8.61% |
The largest absolute tensor-tier error I saw anywhere in the successful L4 rows was about 2.8%.
I would not turn that into an “L4 accuracy = X%” result, though: this was a small, deliberately diagnostic panel, not a random independent benchmark.
Where the remaining error seems to live
I kept FitCheck’s own distinction between the physical tensor terms and the full-process total, because the two behaved quite differently.
For the final TinyLlama grid, the tensor prediction tracks the batch/sequence interaction fairly closely.
I used the ordinary 2×2 interaction here — effectively, how much the batch-size effect changes when sequence length changes.
Eager interaction
Tensor tier:
- predicted interaction: +1997.5 MiB
- measured interaction: +1890.4 MiB
Process-overhead residual:
- predicted interaction: +99.9 MiB
- observed diagnostic residual: +549.6 MiB
FA2 interaction
Tensor tier:
- predicted interaction: +676.0 MiB
- measured interaction: +687.1 MiB
Process-overhead residual:
- predicted interaction: +33.8 MiB
- observed diagnostic residual: +382.9 MiB
So the larger shape-dependent discrepancy appears after the tensor tier.
I am deliberately calling that a process/runtime-overhead residual rather than saying “fragmentation is the cause.” The measurements do not uniquely identify the mechanism.
For localization, I used:
diagnostic process-overhead residual
= harness process total
- measured peak tensor allocation
That is useful diagnostically, but it should not be interpreted as two values that necessarily peaked at exactly the same instant.
This distinction is also consistent with PyTorch’s own CUDA memory model. PyTorch separates memory occupied by tensors (memory_allocated / max_memory_allocated) from memory managed by its caching allocator (memory_reserved / max_memory_reserved). See CUDA memory management.
External process-memory cross-check
I also sampled process GPU memory independently through NVML every 50 ms.
Across the successful rows, the sampled NVML peak was consistently 8 MiB below the FitCheck harness’s process total.
I would not treat a 50 ms sampler as an oracle — it can miss a very short transient — but that stable relationship makes a large measurement-definition artifact less likely in this particular panel.
PyTorch’s CUDA memory debugging documentation recommends comparing against raw device memory when allocations outside the PyTorch allocator are relevant, which is why I added this cross-check.
Cross-session repeatability
Four TinyLlama configurations were independently repeated in a later session and reproduced the same headline process totals:
- eager, bs=2, seq=1024: 3780 MiB
- eager, bs=1, seq=2048: 4304 MiB
- FA2, bs=2, seq=1024: 3140 MiB
- FA2, bs=1, seq=2048: 3140 MiB
The harness-reported CUDA context was also 248 MiB in every cell of the final eight-run grid.
That does not prove the behavior is universal or deterministic under another driver/runtime, but it makes the observed shape pattern look less like a one-off allocation-noise event.
How I read this
At the tested commit, OVERHEAD_DB has measured profiles for T4 × kernel × quantization, while an L4 falls back to the unmeasured default overhead profile.
So my current reading would be:
- the tensor formulas transferred quite well to this L4/BF16 panel;
- real FA2 also behaved well at the tensor tier;
- the uncalibrated full-process overhead is where the larger L4 residual remains;
- that residual depends on batch/sequence shape, so it is not obviously repairable by only changing one fixed context constant.
This seems more like useful input for the existing card/kernel-specific calibration layer than evidence against the overall decomposition.
What I would and would not conclude / possible next steps
What I think the measurements support
For this exact L4/runtime/model/config panel:
- the tensor-tier estimates were close;
- NF4 and ordinary LoRA both worked well in the tested Qwen points;
- the real
flash_attention_2 tensor estimates were close;
- eager attention’s larger sequence-dependent memory behavior was broadly captured;
- the main remaining error was concentrated in the full-process tier;
- that process residual could be either positive or negative depending on shape.
This also lines up fairly well with the project’s existing architecture: the measurement archive already treats the GPU/kernel/quantization overhead calibration separately, and explicitly calls a second card the remaining gap.
What I would not conclude
I would not use these measurements to claim:
- a universal FitCheck accuracy number for L4;
- an Ada-wide accuracy number;
- that allocator fragmentation is proven to be the root cause;
- that an L4 overhead profile is already fully characterized;
- that the same result holds under other PyTorch / Transformers / PEFT / driver versions;
- that inference/serving has been validated on L4;
- that INT8, checkpointing-off, full fine-tuning, MoE, multimodal, compiled, or multi-GPU cases behave the same way;
- that L4 fit/no-fit boundary accuracy has been validated.
I also would not count the repeated configurations as additional independent accuracy samples.
What I would do next
If the goal is still “find where the estimator is wrong with reproducible measurements,” I probably would not jump immediately to another unrelated model.
A higher-information path seems to be:
- treat some of these L4 rows as candidate second-card calibration data;
- keep other L4 rows as a genuine holdout;
- see how much of the process residual an L4-specific
(GPU, kernel, quantization) profile absorbs;
- only if the residual remains interesting, collect allocator-level diagnostics.
For the last branch, PyTorch exposes useful counters through torch.cuda.memory_stats(), including allocated/reserved bytes and segment counts. The CUDA memory snapshot tooling can then give a more detailed allocation history.
I would treat that as a diagnostic option rather than assuming in advance that fragmentation is the explanation.
A separate L4 boundary/OOM holdout would also be useful eventually, because sizing tools care more about a small under-prediction near the fit boundary than about the same percentage error when there is plenty of headroom. But I think the current rows already answer the immediate cross-GPU question well enough that this does not need to block anything.
Raw reproducibility data
If these measurements are useful for the project, I can provide the raw scripts/measure.py --json outputs rather than only the summarized tables.
I also kept:
- exact FitCheck commit
- full Python/package environment
- raw stderr/stdout
- NVML process-memory samples
- model config embedded in each measurement row
- checksums for the result artifacts
so the rows should be reasonably straightforward to inspect or re-score without relying on the summary above.