GPU utilization is not compute
You seat the card and load the model. nvidia-smi reads GPU-Util 80%. It looks busy. You close the window. Put the datasheet TFLOP/s next to the tokens that actually come out and the numbers disagree. A 312 TFLOP/s card delivers tokens at a few hundredths of that peak. The question is what the 80% was … 더 읽기