I deployed gemma-4-13b-it and gemma-4-31b-it-nvfp4 on an H200 with the RHAII 3.5.1 image to compare the actual results to the estimates from ConfigIQ. Below is a comparison of the configiq estimates alongside the observed results:
| gemma-4-31b-it |
|
|
| |
ConfigIQ Estimates |
Observed |
| KV Cache Memory (GiB) |
74.95 |
64.68 |
| KV Cache Tokens |
972,134 |
626,949 |
| KV Cache Size per Token (KiB) |
880 |
N/A |
| KV Cache Memory / KV Cache Tokens |
80.8 |
108.2 |
| KV Cache Memory / KV Cache Size per Token |
89307.7 |
|
| gemma-4-31b-it-nvfp4 |
|
|
| |
ConfigIQ Estimates |
Observed |
| KV Cache Memory (GiB) |
102.81 |
103.92 |
| KV Cache Tokens |
2,684,804 |
1,007,303 |
| KV Cache Size per Token (KiB) |
440 |
N/A |
| KV Cache Memory / KV Cache Tokens |
40.2 |
108.2 |
| KV Cache Memory / KV Cache Size per Token |
245009.3 |
|
Overall, the available KV Cache Memory estimate was pretty decent for both, however, the tokens estimate and the per token estimate seemed pretty dramatically wrong.
While the estimates did have issues, the biggest concern is that they are not consistent within the estimate.
For example, with the fp16 version, the "KV Cache Memory / KV Cache Tokens (KiB)" value is simply the 74.95 GiB divided by the estimated 972k tokens which comes out to only 80.8 KiB instead of the tools estimated 880. Conversely, if I take the 74.95 GiB and divide that by the 880 KiB, I would only expect to have a KV Cache size of 89307 tokens instead of the estimated 972k. The issue was similar for the nvfp4 model and a lot of other models I did some quick checks on.
In reality the KV Cache per token is 108 KiB per token for both model versions.
I suspect the tool assumes a different KV Cache dtype for the nvfp4 model. The nvfp4 model still registered as bfloat16 to vLLM, so it uses that as the default KV Cache Dtype, so both the fp16 and the nvfp4 version have the same KV Cache size per token (108 KiB).
With that in mind, the 108 KiB per token is very different from the estimate ConfigIQ came up with for both cases.
I deployed
gemma-4-13b-itandgemma-4-31b-it-nvfp4on an H200 with the RHAII 3.5.1 image to compare the actual results to the estimates from ConfigIQ. Below is a comparison of the configiq estimates alongside the observed results:Overall, the available KV Cache Memory estimate was pretty decent for both, however, the tokens estimate and the per token estimate seemed pretty dramatically wrong.
While the estimates did have issues, the biggest concern is that they are not consistent within the estimate.
For example, with the fp16 version, the "KV Cache Memory / KV Cache Tokens (KiB)" value is simply the 74.95 GiB divided by the estimated 972k tokens which comes out to only 80.8 KiB instead of the tools estimated 880. Conversely, if I take the 74.95 GiB and divide that by the 880 KiB, I would only expect to have a KV Cache size of 89307 tokens instead of the estimated 972k. The issue was similar for the nvfp4 model and a lot of other models I did some quick checks on.
In reality the KV Cache per token is 108 KiB per token for both model versions.
I suspect the tool assumes a different KV Cache dtype for the nvfp4 model. The nvfp4 model still registered as bfloat16 to vLLM, so it uses that as the default KV Cache Dtype, so both the fp16 and the nvfp4 version have the same KV Cache size per token (108 KiB).
With that in mind, the 108 KiB per token is very different from the estimate ConfigIQ came up with for both cases.