K3 model run on 80 Nvidia RTX 5090 GPUs reaches 20 tokens per second
A Twitter user shared that they are running the K3 model on a large GPU cluster. The setup consists of eighty Nvidia RTX 5090 graphics cards. The configuration
A Twitter user shared that they are running the K3 model on a large GPU cluster.
The setup consists of eighty Nvidia RTX 5090 graphics cards. The configuration
achieves a throughput of twenty tokens per second. This performance metric
provides insight into K3's scaling capabilities. The post highlights the
hardware intensity required for such runs. It serves as a reference point for
others testing similar models. The community may use this data to benchmark
future optimizations. No further technical details were provided beyond the
headline numbers.