Marigold V2 at four bits: what quantizing a 20-billion-parameter depth model costs
Marigold V2 ships as a 17 GB inference job, but the released checkpoint is a rank-128 LoRA over a frozen 20.4 B backbone, so the backbone is swappable. Four configurations on one rented A40: going 4-bit costs 10% of the time and saves 65% of the memory, picking GGUF over bitsandbytes costs another 1.87x and saves nothing, and streaming the weights from host RAM costs 3.33x.
Marigold V2 is published as a 17 GB inference job, and the number that gets quoted is 1.9 seconds per image on “a single 32 GB GPU”. The released checkpoint is not a 20-billion-parameter model though. It is 926 million parameters, 92% of which are a rank-128 LoRA sitting on a frozen Qwen-Image-Edit-2509 backbone that you download separately. A frozen backbone can be replaced, so I replaced it with six quantized versions and measured them on rented A40s. Going to 4 bits costs 10% of the runtime and saves 65% of the memory. Which 4-bit you pick costs more than that decision did: the same bit width spans 0.87 s to 1.62 s, and one of the three produces a depth map that looks fine and is wrong. Streaming the weights from host RAM, which is what you do when the model will not fit, costs 3.33x on the same card with the same weights. Two GPU hours, about sixty cents.
Marigold V2 [1] released its code and weights this month, ahead of the paper appearing in December, and it is a genuinely good monocular depth model: single-step, and sharp enough to resolve fur and hair-thin edges that earlier models smeared. The interesting part for anyone who wants to run it is structural rather than visual, and it is visible before you download a single weight.
What is in the box
The depth checkpoint is one file, trainables.safetensors, 1.85 GB. Read its
header and it splits three ways.
That ratio is the whole article. LoRA freezes the base weights and trains a low-rank update beside them [3], which means the base weights are interchangeable as long as the shapes match. Nothing in Marigold’s 926 M parameters cares whether the 20.4 B underneath is stored in bfloat16, in 4-bit NormalFloat, or in a GGUF block format.
Two more details make this cheap to test. The text encoder never loads: Marigold precomputes the prompt embedding once and ships it as a 19 MB tensor, so the 8.3 B Qwen2.5-VL half of the base model is not part of the runtime at all. And inference is a single step at a fixed timestep of 0.499, not a schedule, so the 20.4 B parameters are read exactly once per image.
The measurement
Everything below is one image at 768 px on one rented A40 with 48 GB, median of five runs after two discarded warmups. The image is the kitten the authors ship as an example, and it is a deliberate choice: fur is the thing Marigold V2 claims to resolve, so if a quantized backbone degrades, it should show up here first.

The bf16 backbone is the reference. Every other configuration is scored against it, not against ground truth, because the question is what quantization costs, not what Marigold gets wrong.
| Backend | Bits | Placement | Seconds | Peak VRAM | RMSE / range | Pearson r |
|---|---|---|---|---|---|---|
| bf16 | 16 | resident | 0.79 | 44.72 GB | reference | reference |
| bitsandbytes NF4 | 4 | resident | 0.87 | 15.44 GB | 1.92% | 0.9972 |
| torchao INT8 | 8 | resident | 0.96 | 24.32 GB | 0.49% | 0.9998 |
| torchao INT4 | 4 | resident | 1.22 | 24.31 GB | 21.67% | 0.5373 |
| GGUF Q4_K_M | 4 | resident | 1.62 | 17.04 GB | 1.56% | 0.9982 |
| bitsandbytes INT8 | 8 | resident | 2.35 | 24.39 GB | 0.43% | 0.9999 |
| GGUF Q4_K_M | 4 | offloaded | 5.40 | 2.15 GB | 1.56% | 0.9982 |
Read the Pearson column before the seconds column. One row has an r of 0.5373 where every other quantized row is above 0.997, and that row is the reason this post scores against a reference at all.
Four things fall out of the table, and only one of them is about precision.
Four bits is nearly free
NF4 [4] runs at 0.87 s against bf16’s 0.79 s, a 1.10x cost, while peak memory drops from 44.72 GB to 15.44 GB. The depth maps correlate at r = 0.9972.
That correlation is the part worth pausing on. Storing a weight in fewer bits is a rounding error with a known bound, not a corruption, and the arithmetic of how those errors accumulate through a network is standard numerical analysis [6]. NF4 tightens the bound further by spacing its sixteen levels to match a normal distribution, which is roughly how trained weights are distributed [4]. The measured result is what that theory predicts: a 20.4 B backbone loses about two percent of its output range and none of its structure.
Put the two maps side by side and the honest description is that you cannot see the difference, with a caveat the next section measures: the agreement is not uniform, it is near-total across flat regions and looser at edges.


Which four bits you pick matters more than whether you quantize
Three of the rows are 4-bit, and they are not interchangeable. NF4 takes 0.87 s in 15.44 GB. GGUF Q4_K_M takes 1.62 s in 17.04 GB, which is more time and more memory. torchao’s INT4 takes 1.22 s in 24.31 GB, which is more memory than the 8-bit rows, and it is wrong.
The same spread appears at 8 bits. torchao INT8 and bitsandbytes INT8 sit within 70 MB of each other on memory and agree with bf16 to four decimal places, and one takes 0.96 s while the other takes 2.35 s. Same width, same memory, 2.45x apart.
A backbone can fail without failing
torchao’s INT4 row loaded without a warning, ran at a believable 1.22 s, and produced an image that has a cat in it, nearer at the front and further at the back. Scored against bf16 it correlates at 0.5373, its median pixel is off by 16% of the output range, and 82% of its pixels are off by more than 5%. It also uses 24.31 GB, which is more than the 8-bit rows, so whatever it is doing it is not packing four-bit weights.
Two routes to torchao INT4 exist and neither works here. The documented one,
Int4WeightOnlyConfig, needs the mslk kernel package, and the version on PyPI
is a 0.0.0 placeholder against a >=1.0.0 requirement. The fallback,
IntxWeightOnlyConfig at torch.int4, installs and runs and gives the row
above. Grouped quantization would be the next thing to try, except that the
img_in layer takes the VAE’s 64 latent channels, so any group size above 64
fails an assertion on that one layer.
The reason to publish a dead end is that it is the failure mode this whole class of work is most exposed to. A detector that breaks returns no boxes. A depth model that breaks returns a depth map. You find out by scoring against something you trust, and the cost of not doing it is a post full of numbers that are confidently wrong.
Quantization disagrees at edges, not everywhere
The widget above amplifies the signed difference between a quantized backbone and bf16. At 1x every working backbone is flat grey, which is the claim in the table made visible. Turn the gain up and two things appear. Edges light up sharply, and whole regions drift slowly one way or the other.
Both are worth stating numerically, because a single RMSE hides the shape of what it averages.
The cliff is between fitting and not fitting
The last row is the same GGUF file on the same A40, with the backbone streamed from host RAM one block at a time instead of sitting on the card. Peak VRAM falls from 17.04 GB to 2.15 GB and the time goes from 1.62 s to 5.40 s.
That 3.33x is larger than every precision effect in the table combined. Going from bf16 to the cheapest resident 4-bit path costs 1.10x. Not having room for the model costs 3.33x. When people say a quantized model is slow, this is usually the effect they have measured.
The two GGUF rows are worth one more sentence, because they are a check rather than a finding. Their depth maps are bit-identical, maximum absolute difference exactly zero. Offloading moves weights between host and device and changes no arithmetic, which is what you want to confirm before attributing any difference to it.
Where AbsRel lies to you
The fidelity columns above are RMSE and correlation, and that is a deliberate departure from what depth papers report. AbsRel is the standard metric, and on this data it reads 0.1437 for the GGUF backbone, which would suggest the depth map is 14% wrong.
It is not. AbsRel divides by the target:
Marigold’s output is affine-invariant relative depth, normalised into , so it is centred near zero by construction. In this image 0.46% of pixels sit within of zero, and each of those divides a small numerator by a smaller denominator. A metric built for metric depth, which is strictly positive and bounded away from zero, gets dominated by a handful of pixels when you point it at relative depth.
RMSE as a fraction of the output range gives 1.56%, and Pearson correlation gives 0.9982. Those describe the same two arrays, and they are the ones that survive a sanity check: if the maps really differed by 14% in any meaningful sense, you would see it in the images above, and you do not. The same care applies whenever a metric is transplanted between problems, which is a theme in the object-detection metrics post and in segmentation metrics.
Reproducibility
| GPU | NVIDIA A40, 48 GB (47.7 GB usable), compute capability sm_86. Two sessions on different hosts, drivers 570.195.03 and 580.159.04 |
| Cross-check | bf16 re-measured in the second session at 0.80 s against the first session’s 0.79 s, a 1.3% spread across hosts and drivers |
| Host | RunPod secure cloud, 9 vCPU, 0.32 for both sessions |
| OS | Ubuntu 24.04, container image runpod/pytorch:1.3.2-rc.168-cu1281-torch2130-ubuntu2404 |
| Python | 3.12 |
| Libraries | torch 2.13.0+cu129, diffusers 0.40.0 [5], peft 0.21.0, transformers 5.17.0, bitsandbytes 0.50.2, torchao 0.18.0 |
| Caveat | those versions were captured in the second session. The first session installed the same constraints hours earlier the same day, so the bf16, NF4 and GGUF rows were almost certainly built against them, but I did not record it at the time and am not going to claim it |
| Model | Marigold V2 depth/Log-stage2, sha256 3edec694…56892, Apache-2.0 [1] |
| Backbone | Qwen-Image-Edit-2509, bf16 (40.9 GB), Apache-2.0 [2]; GGUF Q4_K_M from QuantStack (13,065,746,976 bytes) |
| Data | One image, 15_kitten.jpg, from the authors’ own example set |
| Resolution | 768 px long edge, rounded to a multiple of 16 |
| Timing | Median of 5 runs after 2 discarded warmups; covers VAE encode, the transformer step and VAE decode |
| Excluded | Model loading, weight download, image I/O and colourisation |
| Seed | 2026, fixed, because the VAE posterior is sampled |
| Commands | scripts/pod_run.sh and scripts/pod_run2.sh in the companion repository |
Every number in this post comes from a JSON artifact in output/a40/ or
output/a40_session2/ of the companion repository, and the table is regenerated
from them by scripts/build_table.py rather than typed. Rows from the second
session are scored against that session’s own bf16 run, not the first’s.
Limitations
One image, one resolution. Every number here is 768 px on a single photograph. Latency at 768 px is not latency at 2048 px, where the paper’s own figures show the cost rising faster than pixel count. Fidelity on a kitten is not fidelity on a transparent surface or a night scene, which are the cases monocular depth models fail on.
One GGUF level. Q4_K_M is one point on a range that runs from Q2_K at 7.15 GB to Q8_0 at 21.8 GB. I did not sweep it, so this post cannot tell you where GGUF stops being faithful. It only tells you that at Q4_K_M it has not started to fail.
The broken INT4 row is a report, not a diagnosis. I know torchao’s
IntxWeightOnlyConfig at four bits produces a wrong depth map for this model at
per-axis granularity. I did not find out why, and a grouped configuration that
respects the 64-channel img_in layer might well work. Read that row as “this
combination fails”, not as “torchao cannot do four bits”.
One quantizer family is missing entirely. ComfyUI ships an INT8 backbone
built specifically for Marigold V2, qwen_image_edit_2509_int8_convrot, which
its own metadata describes as int8_tensorwise with a rotation at group size
256. Rotation-based quantization is a different technique from the weight-only
schemes measured here and it is the one a lot of people will run.
Loading it outside ComfyUI means porting the inverse rotation, and a subtly
wrong dequantization produces the kind of plausible, wrong depth map the
INT4 row above is about, so I left it alone rather than guess.
No FP8. torchao’s float8 schemes want compute capability 8.9 and an A40 is 8.6. Measuring FP8 means changing the GPU, which would make every row here incomparable with the others.
The offloaded row is one configuration of offloading. It streams one transformer block at a time with no pinned-memory staging. Both of those are dials, and a machine with host RAM to spare can overlap the transfers and pay less than 3.33x. Treat that number as the cost of a conservative setting, not a constant of nature.
Fidelity is measured against bf16, not ground truth. A backbone could track bf16 closely and both could be wrong about the scene. For absolute accuracy the right reference is the paper’s own benchmarks [1], not this table.
Nothing here is TensorRT or OpenVINO. bitsandbytes NF4 and the GGUF kernels are tied to their runtimes, so neither configuration exports. OpenVINO has become reachable since Marigold V2 shipped, because optimum-intel now registers the Qwen-Image transformer and both VAE halves as exportable, but getting there means merging the LoRA into bf16 weights and exporting a 20 B model. That is a separate piece of work and I have not done it.
A single run is not an effect. Each row is five runs on one machine on one day. The spreads were tight, but I have not tested whether a different A40 host, a different driver, or a different container gives the same numbers.
Further reading
- Marigold V2, the paper [1] for the training recipe, the ablations on smaller backbones like SD1.5 and FLUX.2 klein, and the benchmark tables this post deliberately does not duplicate.
- QLoRA [4] for why 4-bit NormalFloat is the shape it is. It is an information-theoretic argument about normally distributed weights, and it explains why NF4 costs so little accuracy here.
- Quantization, bit depth and banding for the same idea one level down, where the thing being quantized is a pixel rather than a weight, and the artifact you get is visible.
- RF-DETR against YOLO-NAS on the edge and the OBS plugin post for the same measure-then-decide approach applied to detection and to a real-time video filter.
References
[1] Pavlovic, I., Wandel, T., Obukhov, A., Bartolomei, L., Davydov, A., Tosi, F., Poggi, M., Süsstrunk, S., Dai, D. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation. To appear in ACM Transactions on Graphics 45(6), presented at SIGGRAPH Asia, December 2026. arxiv.org/abs/2609.08084. The authors announce DOI 10.1145/3842528, which was not yet registered when this post was written, so the arXiv version is the one linked.
[2] Qwen Team. Qwen-Image-Edit-2509. Model card, Apache-2.0, 2026. huggingface.co/Qwen/Qwen-Image-Edit-2509.
[3] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. arXiv, 2021. arxiv.org/abs/2106.09685.
[4] Dettmers, T., Pagnoni, A., Holtzman, A., Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv, 2023. arxiv.org/abs/2305.14314.
[5] Hugging Face. Diffusers Documentation: quantization backends (0.40.0). huggingface.co/docs/diffusers/quantization/overview.
[6] Goodfellow, I., Bengio, Y., Courville, A. Deep Learning, ch. 4 (Numerical Computation). MIT Press, 2016. deeplearningbook.org.