gpo.zugaina.org

Search Portage & Overlays:

dev-python/tokenspeed-mla-bin

TokenSpeed multi-head latent attention CUDA kernels

Screenshots

  • tokenspeed-mla-bin-0.2.16
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • tokenspeed-mla-bin-0.2.15
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • tokenspeed-mla-bin-0.2.14
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • tokenspeed-mla-bin-0.1.8
    ~amd64
    python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14

    View      Download      Browse     License: MIT   
    Overlay: stuff

ChangeLog

commit 94c9b08eb48d812d4dc7b320d3f58b1cca170037
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Oct 3 23:19:32 2026 +0200

*/*: indent the remaining metadata.xml files with spaces

Twenty-six files imported from ::gentoo in August and September kept
its tab indentation, while the other 705 use two spaces, as Gentoo's
skel.metadata.xml does. Only leading whitespace changes; each file
parses to the same content.

commit e1f5543d4e59fe41a8aa37ccff11d99f64f0e5ef
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Oct 2 10:01:07 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.13

Retention keeps the last three releases; 0.1.8 stays for the consumers
that pin it exactly.

commit 98414deeb8285f52c215b860b9ec3cb51d4adc01
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Oct 2 10:01:07 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.16

Requires-Dist is unchanged from 0.2.15; only the hashed PyPI path
differs.

commit a269e7e403de3769c8ff6a6a09bf2f1f262959a2
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Tue Sep 29 15:16:57 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.12

Retention keeps the last three 0.2.x releases plus 0.1.8, which both
vllm ebuilds pin.

commit 13381381f4ebf480e382e2f61db584556d157174
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Tue Sep 29 15:16:31 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.15

The wheel's package files are identical to 0.2.14; only the metadata
changed. Upstream relaxed apache-tvm-ffi from ==0.1.13.post3 to
>=0.1.11,<0.2, so the dependency becomes a range, which admits every
apache-tvm-ffi this overlay ships.

Both vllm ebuilds pin ~0.1.8, so nothing in the tree takes this yet.
It is added ahead of the vllm bump that will pin it.

commit c6bd38449890ccc177bc90dd8da043912e3f8a4f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Mon Sep 28 13:41:41 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.11

Retention keeps the last three releases; 0.1.8 stays because vllm pins
it exactly.

commit 1e0ba6445e20d3a450d1704f406aef33c629c49c
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Mon Sep 28 13:41:27 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.14

Requires-Dist matches 0.2.13, exact apache-tvm-ffi 0.1.13.post3 pin
included, so only the wheel URL moves. The release reworks the FP8
split-KV reducer and adds an opt-in fp16 partials mode.

commit 024ef7615f0fb2965da9430910655ce48848143f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 24 20:45:44 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.10

Retention keeps the last three releases; 0.1.8 stays because vllm pins
it exactly.

commit 48b265a42bcf36a06a005a98888387b17191b46b
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 24 20:45:43 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.13

Requires-Dist matches 0.2.12, exact apache-tvm-ffi 0.1.13.post3 pin
included, so only the wheel URL moves.

commit 1415e8ad48ed0c342182eff0a4b15e9c2397020a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 24 12:33:19 2026 +0200

dev-python/tokenspeed-mla-bin: restore 0.1.8

vllm 0.29.0 and 0.30.0 pin ~dev-python/tokenspeed-mla-bin-0.1.8 under
USE=cuda, so its removal left their CUDA dependencies unsatisfiable.
Restored unchanged; it stays until vllm moves the pin.

commit 53992b55f2007cb1884e10c3d1bfb63625fcc898
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 24 13:03:31 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.9

Retention keeps the last three releases.

commit 48c18605f27acc2f3261e28982c79cf4a1287967
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 24 10:11:14 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.12

The wheel now requires nvidia-cutlass-dsl>=4.8.0 (0.2.11 left it
unversioned). The exact apache-tvm-ffi 0.1.13.post3 pin and the
tokenspeed-triton floor are unchanged.

commit 39baa4f4929adfeab925e34e1cba6b3ec48963fe
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 23 11:45:47 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.1.8, 0.2.8

Keeps the last three. 0.1.8 is not a rollback anchor for an earlier
major series -- the package has only ever been 0.x.

commit 5611775dac8e95751382a5235478917c36c30121
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 23 11:45:47 2026 +0200

dev-util/codex: drop 0.153.2, 0.153.4, 0.154.0

Keeps the last three. Each release carries its own ~150 MiB crate
tarball plus the source archive, so the distfile tail is the reason to
trim rather than the ebuild count.

commit aec6cceecad94527af66c6d03bbc4a086de05779
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 23 10:11:46 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.11

Declared requirements are identical to 0.2.10, so only the wheel's
hashed PyPI path changes.

commit 81db13711a270346be4a1a7ee23f94887f4aae1d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Sep 20 14:55:50 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.10

The wheel raises its own floor to tokenspeed-triton >=3.8.10.post20260920,
so the RDEPEND floor moves with it. The apache-tvm-ffi ==0.1.13.post3 pin,
nvidia-cutlass-dsl and torch are unchanged from 0.2.9.

commit 612748802d796027ed5d8bfcca622d04fd13b4bc
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 16 13:00:34 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.6 and 0.2.7

Superseded by 0.2.8 and 0.2.9. 0.1.8 stays as the last of the 0.1
series.

commit 48f98bd9c061fe0467df1ed3150f572f677ee834
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 16 09:47:44 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.9

Upstream dropped the optional ahead-of-time prefill backend, so the
wheel no longer carries compiled objects and is plain Python. Follow
that: take the pure wheel and drop the extension, prebuilt-QA and
strip settings that only applied to those objects. All prefill now
goes through the CuTe DSL path, which was already the default.

commit 6121863cea3fce11dbb0609bcc0c085f44842ec1
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Sep 12 13:53:59 2026 +0200

dev-python/tokenspeed-mla-bin: bump to 0.2.8

Upstream 0.2.8 keeps the exact apache-tvm-ffi pin and minimum TokenSpeed
Triton version, so its dependency contract is unchanged.

The wheel has extension modules, which are marked and byte-compiled.
`~amd64` remains the only keyword because dependencies are amd64-only.

A full staged install passed for the enabled Python target.

commit 0367bbe9d96ad1e1e1edb67402b9a39f20c6b9f4
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 9 10:34:07 2026 +0200

dev-python/tokenspeed-mla-bin: drop 0.2.5

Last two of the 0.2 line are 0.2.6 and 0.2.7. 0.1.8 stays as the 0.1
anchor -- dev-python/vllm pins it exactly, and it is what holds
tokenspeed-triton-bin 3.7.10_p20260531 in the tree.

commit 06d803c0ba6e3459d301a8f8df93b84d89164203
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Sep 9 09:48:12 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.7

Upstream raised the tokenspeed-triton floor to >=3.8.10.post20260906; the
atom follows the version added in the preceding commit. apache-tvm-ffi is
still pinned at exactly 0.1.13.post3, so that pair of bounds is unchanged.

commit 4b49d69d779b8d05dca3bb1651813730b23ac2e0
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 22:02:12 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.6

Wheel repackage, so this is not a bare $ bump: SRC_URI hardcodes a
content-addressed pythonhosted path, and the new wheel's hash directory had
to be pulled from the PyPI JSON. Fetched size matches PyPI's declared 755954
bytes exactly.

RDEPEND change, read from the wheel's own METADATA rather than the PyPI
summary: upstream moved its exact tvm-ffi pin from apache-tvm-ffi==0.1.13 to
==0.1.13.post3, so the >=/<= exact-pin pair moves to 0.1.13_p3. That version
is already in the overlay, so the pin resolves. The other three requirements
(nvidia-cutlass-dsl, tokenspeed-triton>=3.8.10.post20260721, torch) are
unchanged.

Adds-only, nothing dropped: every vllm ebuild in the tree exact-pins
~dev-python/tokenspeed-mla-bin-0.1.8, so 0.1.8 is the only version any
consumer can reach and dropping it would break all three. 0.2.5 and 0.2.6
sit unconsumed until a vllm bump re-pins them -- the same chain-driven
retention shape this package has had since it was added.

Build-verified as root; installs the two objs/*.so kernels and the
dist-info. The unresolved-soname QA notice on libcute_dsl_runtime.so is
host-derived, not a missing dep: nvidia-cutlass-dsl is declared in RDEPEND
but is not installed here, and these are sm_100a/sm_103a Blackwell CUDA
blobs on an AMD-only machine. 0.2.5 carries the identical QA_PREBUILT shape.

commit c563f555bfd7a07238f9ba0b7be911fbb9b6be2f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Aug 15 17:53:23 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.2.5

Add the current upstream kernels and follow their exact apache-tvm-ffi pin and
TokenSpeed Triton floor.

commit c1a398c9fd10a70d6708081bfd1631d99624ff0d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 21:56:10 2026 +0200

dev-python/tokenspeed-mla-bin: add 0.1.8

vLLM pins TokenSpeed MLA for speculative-decoding attention kernels. Declare the
exact upstream runtime closure and preserve its Blackwell-native wheel payload.