Marigold V2 at four bits: what quantizing a 20-billion-parameter depth model costs

Marigold V2 ships as a 17 GB inference job, but the released checkpoint is a rank-128 LoRA over a frozen 20.4 B backbone, so the backbone is swappable. Four configurations on one rented A40: going 4-bit costs 10% of the time and saves 65% of the memory, picking GGUF over bitsandbytes costs another 1.87x and saves nothing, and streaming the weights from host RAM costs 3.33x.

Luis Condados · · Updated September 22, 2026
The bf16 reference depth map every other configuration in this post is scored against: Marigold V2 on a kitten, resolving individual fur strands and whiskers.
The bf16 reference depth map every other configuration in this post is scored against: Marigold V2 on a kitten, resolving individual fur strands and whiskers.

Marigold V2 is published as a 17 GB inference job, and the number that gets quoted is 1.9 seconds per image on “a single 32 GB GPU”. The released checkpoint is not a 20-billion-parameter model though. It is 926 million parameters, 92% of which are a rank-128 LoRA sitting on a frozen Qwen-Image-Edit-2509 backbone that you download separately. A frozen backbone can be replaced, so I replaced it with six quantized versions and measured them on rented A40s. Going to 4 bits costs 10% of the runtime and saves 65% of the memory. Which 4-bit you pick costs more than that decision did: the same bit width spans 0.87 s to 1.62 s, and one of the three produces a depth map that looks fine and is wrong. Streaming the weights from host RAM, which is what you do when the model will not fit, costs 3.33x on the same card with the same weights. Two GPU hours, about sixty cents.

Marigold V2 [1] released its code and weights this month, ahead of the paper appearing in December, and it is a genuinely good monocular depth model: single-step, and sharp enough to resolve fur and hair-thin edges that earlier models smeared. The interesting part for anyone who wants to run it is structural rather than visual, and it is visible before you download a single weight.

What is in the box

The depth checkpoint is one file, trainables.safetensors, 1.85 GB. Read its header and it splits three ways.

That ratio is the whole article. LoRA freezes the base weights and trains a low-rank update beside them [3], which means the base weights are interchangeable as long as the shapes match. Nothing in Marigold’s 926 M parameters cares whether the 20.4 B underneath is stored in bfloat16, in 4-bit NormalFloat, or in a GGUF block format.

Two more details make this cheap to test. The text encoder never loads: Marigold precomputes the prompt embedding once and ships it as a 19 MB tensor, so the 8.3 B Qwen2.5-VL half of the base model is not part of the runtime at all. And inference is a single step at a fixed timestep of 0.499, not a schedule, so the 20.4 B parameters are read exactly once per image.

RGB768 pxVAE encodeseeded sampleMMDiT, one step20.4 B frozen+ 852 M LoRAt = 0.499VAE decodefine-tuneddepthmean of 3frozen prompt embeds, 19 MB on diskorange is what Marigold ships · the rest is Qwen-Image-Edit-2509, downloaded separatelythe 8.3 B text encoder is never loaded
The whole inference graph. One function evaluation, no schedule, no text encoder at runtime.

The measurement

Everything below is one image at 768 px on one rented A40 with 48 GB, median of five runs after two discarded warmups. The image is the kitten the authors ship as an example, and it is a deliberate choice: fur is the thing Marigold V2 claims to resolve, so if a quantized backbone degrades, it should show up here first.

The input image: a fluffy kitten in a garden

The bf16 backbone is the reference. Every other configuration is scored against it, not against ground truth, because the question is what quantization costs, not what Marigold gets wrong.

BackendBitsPlacementSecondsPeak VRAMRMSE / rangePearson r
bf1616resident0.7944.72 GBreferencereference
bitsandbytes NF44resident0.8715.44 GB1.92%0.9972
torchao INT88resident0.9624.32 GB0.49%0.9998
torchao INT44resident1.2224.31 GB21.67%0.5373
GGUF Q4_K_M4resident1.6217.04 GB1.56%0.9982
bitsandbytes INT88resident2.3524.39 GB0.43%0.9999
GGUF Q4_K_M4offloaded5.402.15 GB1.56%0.9982

Read the Pearson column before the seconds column. One row has an r of 0.5373 where every other quantized row is above 0.997, and that row is the reason this post scores against a reference at all.

Four things fall out of the table, and only one of them is about precision.

Four bits is nearly free

NF4 [4] runs at 0.87 s against bf16’s 0.79 s, a 1.10x cost, while peak memory drops from 44.72 GB to 15.44 GB. The depth maps correlate at r = 0.9972.

That correlation is the part worth pausing on. Storing a weight in fewer bits is a rounding error with a known bound, not a corruption, and the arithmetic of how those errors accumulate through a network is standard numerical analysis [6]. NF4 tightens the bound further by spacing its sixteen levels to match a normal distribution, which is roughly how trained weights are distributed [4]. The measured result is what that theory predicts: a 20.4 B backbone loses about two percent of its output range and none of its structure.

Put the two maps side by side and the honest description is that you cannot see the difference, with a caveat the next section measures: the agreement is not uniform, it is near-total across flat regions and looser at edges.

Depth map from the bf16 backbone

Depth map from the GGUF Q4_K_M backbone, visually indistinguishable from bf16

Which four bits you pick matters more than whether you quantize

Three of the rows are 4-bit, and they are not interchangeable. NF4 takes 0.87 s in 15.44 GB. GGUF Q4_K_M takes 1.62 s in 17.04 GB, which is more time and more memory. torchao’s INT4 takes 1.22 s in 24.31 GB, which is more memory than the 8-bit rows, and it is wrong.

The same spread appears at 8 bits. torchao INT8 and bitsandbytes INT8 sit within 70 MB of each other on memory and agree with bf16 to four decimal places, and one takes 0.96 s while the other takes 2.35 s. Same width, same memory, 2.45x apart.

A backbone can fail without failing

torchao’s INT4 row loaded without a warning, ran at a believable 1.22 s, and produced an image that has a cat in it, nearer at the front and further at the back. Scored against bf16 it correlates at 0.5373, its median pixel is off by 16% of the output range, and 82% of its pixels are off by more than 5%. It also uses 24.31 GB, which is more than the 8-bit rows, so whatever it is doing it is not packing four-bit weights.

Two routes to torchao INT4 exist and neither works here. The documented one, Int4WeightOnlyConfig, needs the mslk kernel package, and the version on PyPI is a 0.0.0 placeholder against a >=1.0.0 requirement. The fallback, IntxWeightOnlyConfig at torch.int4, installs and runs and gives the row above. Grouped quantization would be the next thing to try, except that the img_in layer takes the VAE’s 64 latent channels, so any group size above 64 fails an assertion on that one layer.

The reason to publish a dead end is that it is the failure mode this whole class of work is most exposed to. A detector that breaks returns no boxes. A depth model that breaks returns a depth map. You find out by scoring against something you trust, and the cost of not doing it is a post full of numbers that are confidently wrong.

Try it: where quantization actually disagrees

Quantization disagrees at edges, not everywhere

The widget above amplifies the signed difference between a quantized backbone and bf16. At 1x every working backbone is flat grey, which is the claim in the table made visible. Turn the gain up and two things appear. Edges light up sharply, and whole regions drift slowly one way or the other.

Both are worth stating numerically, because a single RMSE hides the shape of what it averages.

The cliff is between fitting and not fitting

The last row is the same GGUF file on the same A40, with the backbone streamed from host RAM one block at a time instead of sitting on the card. Peak VRAM falls from 17.04 GB to 2.15 GB and the time goes from 1.62 s to 5.40 s.

5.40 s1.62 s=3.33\frac{5.40\,\text{s}}{1.62\,\text{s}} = 3.33

That 3.33x is larger than every precision effect in the table combined. Going from bf16 to the cheapest resident 4-bit path costs 1.10x. Not having room for the model costs 3.33x. When people say a quantized model is slow, this is usually the effect they have measured.

The two GGUF rows are worth one more sentence, because they are a check rather than a finding. Their depth maps are bit-identical, maximum absolute difference exactly zero. Offloading moves weights between host and device and changes no arithmetic, which is what you want to confirm before attributing any difference to it.

Try it: what fits on your card

Where AbsRel lies to you

The fidelity columns above are RMSE and correlation, and that is a deliberate departure from what depth papers report. AbsRel is the standard metric, and on this data it reads 0.1437 for the GGUF backbone, which would suggest the depth map is 14% wrong.

It is not. AbsRel divides by the target:

AbsRel=1N∑i∣pi−ti∣∣ti∣\text{AbsRel} = \frac{1}{N}\sum_i \frac{|p_i - t_i|}{|t_i|}

Marigold’s output is affine-invariant relative depth, normalised into [−1,1][-1, 1], so it is centred near zero by construction. In this image 0.46% of pixels sit within ∣v∣<0.01|v| < 0.01 of zero, and each of those divides a small numerator by a smaller denominator. A metric built for metric depth, which is strictly positive and bounded away from zero, gets dominated by a handful of pixels when you point it at relative depth.

RMSE as a fraction of the output range gives 1.56%, and Pearson correlation gives 0.9982. Those describe the same two arrays, and they are the ones that survive a sanity check: if the maps really differed by 14% in any meaningful sense, you would see it in the images above, and you do not. The same care applies whenever a metric is transplanted between problems, which is a theme in the object-detection metrics post and in segmentation metrics.

Reproducibility

GPUNVIDIA A40, 48 GB (47.7 GB usable), compute capability sm_86. Two sessions on different hosts, drivers 570.195.03 and 580.159.04
Cross-checkbf16 re-measured in the second session at 0.80 s against the first session’s 0.79 s, a 1.3% spread across hosts and drivers
HostRunPod secure cloud, 9 vCPU, 0.49/hr,roughly0.49/hr, roughly 0.32 for both sessions
OSUbuntu 24.04, container image runpod/pytorch:1.3.2-rc.168-cu1281-torch2130-ubuntu2404
Python3.12
Librariestorch 2.13.0+cu129, diffusers 0.40.0 [5], peft 0.21.0, transformers 5.17.0, bitsandbytes 0.50.2, torchao 0.18.0
Caveatthose versions were captured in the second session. The first session installed the same constraints hours earlier the same day, so the bf16, NF4 and GGUF rows were almost certainly built against them, but I did not record it at the time and am not going to claim it
ModelMarigold V2 depth/Log-stage2, sha256 3edec694…56892, Apache-2.0 [1]
BackboneQwen-Image-Edit-2509, bf16 (40.9 GB), Apache-2.0 [2]; GGUF Q4_K_M from QuantStack (13,065,746,976 bytes)
DataOne image, 15_kitten.jpg, from the authors’ own example set
Resolution768 px long edge, rounded to a multiple of 16
TimingMedian of 5 runs after 2 discarded warmups; covers VAE encode, the transformer step and VAE decode
ExcludedModel loading, weight download, image I/O and colourisation
Seed2026, fixed, because the VAE posterior is sampled
Commandsscripts/pod_run.sh and scripts/pod_run2.sh in the companion repository

Every number in this post comes from a JSON artifact in output/a40/ or output/a40_session2/ of the companion repository, and the table is regenerated from them by scripts/build_table.py rather than typed. Rows from the second session are scored against that session’s own bf16 run, not the first’s.

Limitations

One image, one resolution. Every number here is 768 px on a single photograph. Latency at 768 px is not latency at 2048 px, where the paper’s own figures show the cost rising faster than pixel count. Fidelity on a kitten is not fidelity on a transparent surface or a night scene, which are the cases monocular depth models fail on.

One GGUF level. Q4_K_M is one point on a range that runs from Q2_K at 7.15 GB to Q8_0 at 21.8 GB. I did not sweep it, so this post cannot tell you where GGUF stops being faithful. It only tells you that at Q4_K_M it has not started to fail.

The broken INT4 row is a report, not a diagnosis. I know torchao’s IntxWeightOnlyConfig at four bits produces a wrong depth map for this model at per-axis granularity. I did not find out why, and a grouped configuration that respects the 64-channel img_in layer might well work. Read that row as “this combination fails”, not as “torchao cannot do four bits”.

One quantizer family is missing entirely. ComfyUI ships an INT8 backbone built specifically for Marigold V2, qwen_image_edit_2509_int8_convrot, which its own metadata describes as int8_tensorwise with a rotation at group size 256. Rotation-based quantization is a different technique from the weight-only schemes measured here and it is the one a lot of people will run. Loading it outside ComfyUI means porting the inverse rotation, and a subtly wrong dequantization produces the kind of plausible, wrong depth map the INT4 row above is about, so I left it alone rather than guess.

No FP8. torchao’s float8 schemes want compute capability 8.9 and an A40 is 8.6. Measuring FP8 means changing the GPU, which would make every row here incomparable with the others.

The offloaded row is one configuration of offloading. It streams one transformer block at a time with no pinned-memory staging. Both of those are dials, and a machine with host RAM to spare can overlap the transfers and pay less than 3.33x. Treat that number as the cost of a conservative setting, not a constant of nature.

Fidelity is measured against bf16, not ground truth. A backbone could track bf16 closely and both could be wrong about the scene. For absolute accuracy the right reference is the paper’s own benchmarks [1], not this table.

Nothing here is TensorRT or OpenVINO. bitsandbytes NF4 and the GGUF kernels are tied to their runtimes, so neither configuration exports. OpenVINO has become reachable since Marigold V2 shipped, because optimum-intel now registers the Qwen-Image transformer and both VAE halves as exportable, but getting there means merging the LoRA into bf16 weights and exporting a 20 B model. That is a separate piece of work and I have not done it.

A single run is not an effect. Each row is five runs on one machine on one day. The spreads were tight, but I have not tested whether a different A40 host, a different driver, or a different container gives the same numbers.

Further reading

  • Marigold V2, the paper [1] for the training recipe, the ablations on smaller backbones like SD1.5 and FLUX.2 klein, and the benchmark tables this post deliberately does not duplicate.
  • QLoRA [4] for why 4-bit NormalFloat is the shape it is. It is an information-theoretic argument about normally distributed weights, and it explains why NF4 costs so little accuracy here.
  • Quantization, bit depth and banding for the same idea one level down, where the thing being quantized is a pixel rather than a weight, and the artifact you get is visible.
  • RF-DETR against YOLO-NAS on the edge and the OBS plugin post for the same measure-then-decide approach applied to detection and to a real-time video filter.

References

[1] Pavlovic, I., Wandel, T., Obukhov, A., Bartolomei, L., Davydov, A., Tosi, F., Poggi, M., Süsstrunk, S., Dai, D. Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation. To appear in ACM Transactions on Graphics 45(6), presented at SIGGRAPH Asia, December 2026. arxiv.org/abs/2609.08084. The authors announce DOI 10.1145/3842528, which was not yet registered when this post was written, so the arXiv version is the one linked.

[2] Qwen Team. Qwen-Image-Edit-2509. Model card, Apache-2.0, 2026. huggingface.co/Qwen/Qwen-Image-Edit-2509.

[3] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. arXiv, 2021. arxiv.org/abs/2106.09685.

[4] Dettmers, T., Pagnoni, A., Holtzman, A., Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv, 2023. arxiv.org/abs/2305.14314.

[5] Hugging Face. Diffusers Documentation: quantization backends (0.40.0). huggingface.co/docs/diffusers/quantization/overview.

[6] Goodfellow, I., Bengio, Y., Courville, A. Deep Learning, ch. 4 (Numerical Computation). MIT Press, 2016. deeplearningbook.org.