dev-python/tokenspeed-mla-bin
TokenSpeed multi-head latent attention CUDA kernels
ChangeLog
commit c563f555bfd7a07238f9ba0b7be911fbb9b6be2f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Aug 15 17:53:23 2026 +0200
dev-python/tokenspeed-mla-bin: add 0.2.5
Add the current upstream kernels and follow their exact apache-tvm-ffi pin and
TokenSpeed Triton floor.
commit c1a398c9fd10a70d6708081bfd1631d99624ff0d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 21:56:10 2026 +0200
dev-python/tokenspeed-mla-bin: add 0.1.8
vLLM pins TokenSpeed MLA for speculative-decoding attention kernels. Declare the
exact upstream runtime closure and preserve its Blackwell-native wheel payload.
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Aug 15 17:53:23 2026 +0200
dev-python/tokenspeed-mla-bin: add 0.2.5
Add the current upstream kernels and follow their exact apache-tvm-ffi pin and
TokenSpeed Triton floor.
commit c1a398c9fd10a70d6708081bfd1631d99624ff0d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 21:56:10 2026 +0200
dev-python/tokenspeed-mla-bin: add 0.1.8
vLLM pins TokenSpeed MLA for speculative-decoding attention kernels. Declare the
exact upstream runtime closure and preserve its Blackwell-native wheel payload.


View
Download
Browse