gpo.zugaina.org

Search Portage & Overlays:

dev-python/tokenspeed-mla-bin

TokenSpeed multi-head latent attention CUDA kernels

Screenshots

  • tokenspeed-mla-bin-0.2.5
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • tokenspeed-mla-bin-0.1.8
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff

ChangeLog

commit c563f555bfd7a07238f9ba0b7be911fbb9b6be2f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Aug 15 17:53:23 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.5

Add the current upstream kernels and follow their exact apache-tvm-ffi pin and
TokenSpeed Triton floor.

commit c1a398c9fd10a70d6708081bfd1631d99624ff0d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 21:56:10 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.1.8

vLLM pins TokenSpeed MLA for speculative-decoding attention kernels. Declare the
exact upstream runtime closure and preserve its Blackwell-native wheel payload.