K3 model run on 80 Nvidia RTX 5090 GPUs reaches 20 tokens per second

A Twitter user shared that they are running the K3 model on a large GPU cluster. The setup consists of eighty Nvidia RTX 5090 graphics cards. The configuration

A Twitter user shared that they are running the K3 model on a large GPU cluster. The setup consists of eighty Nvidia RTX 5090 graphics cards. The configuration achieves a throughput of twenty tokens per second. This performance metric provides insight into K3's scaling capabilities. The post highlights the hardware intensity required for such runs. It serves as a reference point for others testing similar models. The community may use this data to benchmark future optimizations. No further technical details were provided beyond the headline numbers.