gpo.zugaina.org

Search Portage & Overlays:

sci-libs/composable-kernel

High Performance Composable Kernel for AMD GPUs

Screenshots

  • composable-kernel-10.0.0
    ~amd64
    debug hiptensor profiler test python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1102 +amdgpu_targets_gfx1103 +amdgpu_targets_gfx1150 +amdgpu_targets_gfx1151 +amdgpu_targets_gfx1152 +amdgpu_targets_gfx1153 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • composable-kernel-7.2.4
    ~amd64
    debug profiler test python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151

    View      Download      Browse     License: MIT   
    Overlay: stuff
  • composable-kernel-7.2.0
    ~amd64
    debug profiler test python_targets_python3_13t python_targets_python3_11 python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151

    View      Download      Browse     License: MIT   
    Overlay: gentoo
  • composable-kernel-7.1.0
    ~amd64
    debug profiler test python_targets_python3_13t python_targets_python3_11 python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151

    View      Download      Browse     License: MIT   
    Overlay: gentoo
  • composable-kernel-7.0.2
    ~amd64
    debug profiler test python_targets_python3_13t python_targets_python3_11 python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151

    View      Download      Browse     License: MIT   
    Overlay: gentoo
  • composable-kernel-6.4.3
    ~amd64
    debug profiler test python_targets_python3_13t python_targets_python3_11 python_targets_python3_12 python_targets_python3_13 python_targets_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151

    View      Download      Browse     License: MIT   
    Overlay: gentoo
  • composable-kernel-6.3.0
    ~amd64
    debug profiler test python_targets_python3_13t python_targets_python3_11 python_targets_python3_12 python_targets_python3_13 +amdgpu_targets_gfx906 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1101 amdgpu_targets_gfx1102

    View      Download      Browse     License: MIT   
    Overlay: gentoo

ChangeLog

commit a40132c48daf0a8addeff306fa918eead0dcf367
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 04:23:40 2026 +0200

sci-libs/composable-kernel: trim 7.2.4 comment prose

commit c31bc9d1e20452bc2b93c72d2ee2a7d5b918318d
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:33:34 2026 +0200

sci-libs/composable-kernel: trim 10.0.0 comment prose

commit ab11e33912878a7131ff5dcade300048048374b7
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Mon Aug 31 03:37:50 2026 +0200

sci-libs/composable-kernel: add USE=hiptensor for the contraction instance set

sci-libs/hipTensor does
find_package(composable_kernel 1.0.0 REQUIRED
COMPONENTS device_contraction_operations
device_reduction_operations
device_other_operations)
and our build provides none of the three, because MIOPEN_REQ_LIBS_ONLY=ON
narrows the instance set to what MIOpen needs (conv + utility).

Upstream ships HIPTENSOR_REQ_LIBS_ONLY as the matching narrowing switch, and
the two COMPOSE rather than conflict. In
library/src/tensor_operation_instance/gpu/CMakeLists.txt a single
`required_pattern` is assembled by appending "conv" for the MIOpen switch and
"contract"/"reduce"/"element" for the hipTensor one, so with both ON the filter
is the union; the library-target guards mirror that
(`NOT MIOPEN_REQ_LIBS_ONLY OR HIPTENSOR_REQ_LIBS_ONLY` on the contraction and
other targets, the reverse on the convolution ones).

Verified on the built artifact rather than from the CMake alone:
libdevice_conv_operations.so.1.2.0 comes out at 804.7 MB, byte-identical in
size to the MIOpen-only build, so MIOpen's instance set is provably untouched.
The flag adds device_contraction_operations (129 MB),
device_reduction_operations (151 MB) and device_other_operations (0.6 MB).

Cost is real and the flag is therefore off by default: ~281 MB installed and a
full rebuild of ~1416 ninja steps, roughly two hours at -j16 on this host.

Rebuilt and merged; sci-ml/caffe2's consumers still work afterwards -- torch
2.13.0a0 on gfx1150 runs a 2048x2048 GEMM (max abs err 1.6e-3, normal for fp32)
and an F.conv2d through MIOpen.

commit 69845e9b093c34696b109171c07678b31f67b65f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 23:39:38 2026 +0200

sci-libs/composable-kernel: drop 7.2.3

Superseded by 7.2.4, the last release of the 7.x series and the rollback
target for the ROCm 10.0 migration. Version retention keeps the last two
plus the last of each previous major series; 7.2.4 and 10.0.0 satisfy both,
so the third version is redundant.

commit 3da2afc87c7bdc4167c09a2a46acd7a9a599dd58
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 23:01:26 2026 +0200

sci-libs/composable-kernel: ship the files/ patches xz-compressed

Clears the two size findings this package has carried since the 10.0 patches
landed: composable-kernel-10.0.0-libcxx-includes.patch was 24.7K against
pkgcheck's 20K SizeViolation cap, and files/ totalled 59.7K against the 50K
TotalSizeViolation cap.

xz -9e (single-threaded, per the overlay's compression rule) takes files/ from
61116 to 8128 bytes -- 59.7K to 7.9K -- with the largest single patch now 2064
bytes. Both caps clear with room to spare, so no extra-stuff migration is
needed here; that is the right tool for bundles, not for a flat patch list this
small once compressed.

eapply(1) does NOT decompress -- portage's __eapply_patch feeds the file
straight to patch(1) -- so each ebuild expands files/*.patch.xz into $ and
repoints PATCHES at the plain-text copies before cmake_src_prepare consumes
them. Same shape already used by sci-ml/caffe2 and dev-python/cupy, including
the glob, so an ebuild decompresses all eight and uses whichever subset its
PATCHES names.

All eight patches are compressed, not just the four 10.0.0 uses, because the
shim is glob-driven: it expands files/*.patch.xz and then repoints every
PATCHES entry to $/patches/${b%.xz}. A leftover uncompressed entry would be
repointed to a $ path that was never created, and the build would die. So
the directory has to be all-or-nothing, and 7.2.3/7.2.4 get the same treatment
and the same shim. (Compressing only the 10.0.0 set would have satisfied the
caps on its own -- 26.7K, well under 50K -- so this is about the shim, not
about headroom.)

No rebuild needed, and that is a proof rather than an assumption: every .xz
decompresses byte-identical (sha256) to the patch it replaces, and src_prepare
applies them with zero hunk failures on all three ebuilds -- 4 patches for
10.0.0, 5 each for 7.2.3 and 7.2.4. Identical patch content applied to
identical sources yields an identical prepared tree, so nothing downstream of
it can change. Spot-checked by effect too: the libcxx-includes patch still adds
its include to 3 files under include/ck/.

commit f23ab76167f2e6313b7a2e5f222785cefe362b32
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 22:13:22 2026 +0200

sci-libs/composable-kernel: guard the -Werror and example seds

The -amdgpu-early-inline-all pair was already guarded because a miss there
OOMs the build; these two carry the same no-match-exits-0 hazard and were
left bare. A miss on the first leaves -Werror in place, turning any warning a
newer compiler emits into a hard failure; a miss on the second builds and
installs the entire example tree, which reads as a normal build that simply
took much longer.

Verified by effect after src_prepare: all four targets (both new anchors plus
the two existing -mllvm flags) are gone from the prepared tree. No rebuild
needed -- these add a precondition check ahead of seds whose behaviour is
unchanged, and the package already built green at 10.0.

Note for anyone re-running this: composable-kernel's pkg_pretend sizes RAM at
jobs*3072M, so it refuses -j24 on a 54 GiB host (wants 72 GiB). Use -j16.

commit 707bce79bca998f82096b3d8b5ac752ea614ed07
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 19:05:09 2026 +0200

sci-libs/composable-kernel: require dev-util/hipcc[amd-llvm]

The requirement was real but unencoded. This package cannot be compiled by a
vanilla LLVM -- amd_wmma.hpp passes bhalf16_t, a __bf16 vector, to
__builtin_amdgcn_wmma_f32_16x16x16_bf16_w32, which vanilla clang 23 declares
as taking short __attribute__((ext_vector_type(16))) -- yet nothing in the
ebuild said so. The only record was a paragraph in profiles/package.mask, and
a mask comment is documentation, not a constraint: building against a hipcc
without the flag failed deep in compilation with no resolver-level signal.

Two things were missing, not one. hipcc was not declared at all: this ebuild
calls rocm_use_clang(), which runs `hipconfig --hipclangpath`, and hipconfig
belongs to dev-util/hipcc. It happened to be present only because
dev-util/hip RDEPENDs it -- and a transitive dependency cannot carry a
USE-dep. So the atom is added to BDEPEND directly, where the tool it names is
actually used.

Unconditional on purpose: every failure on record is gfx11xx/gfx12xx, so a
narrower amdgpu_targets_*? form may well be right, but no one has built this
against a vanilla LLVM for a CDNA-only target and claiming unverified support
is the worse error. The comment says what would justify narrowing it.

Verified: amd-llvm is in dev-util/hipcc-10.0.0's IUSE, `emerge -p` resolves
unchanged, and pkgcheck's MissingUseDepDefault -- confirmed live by pointing
the dep at a nonexistent flag -- is silent on the real one.

commit a88c98e4bbc12ad29c7d5c0036ea51499082773c
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 09:55:26 2026 +0200

ROCm 10.0: slot the rocm-cmake deps, drop a stale Tensile sed, widen a guard

Six packages depended on dev-build/rocm-cmake unslotted while their peers used
:$. Emerging an individual masked 10.0 package could therefore satisfy
the build dep with 7.2.x. Now consistent across the cohort:
rocm-device-libs, rocm-opencl-runtime (which also carried a stale
">=dev-build/rocm-cmake-6.0" bound), composable-kernel, hipBLAS-common,
rocPRIM and rocRAND (twice -- DEPEND and BDEPEND).

Tensile: the v_dot4_i32_i8 syntax fix for clang-20 (bug 949817) is upstream at
10.0. Components/MAC_I8X4.py no longer emits the op_sel/op_sel_hi operands the
sed stripped, so it matched nothing. Dropped.

composable-kernel: the OOM-flag guard checked only -amdgpu-early-inline-all
while the sed removes -amdgpu-function-calls too; both are now asserted. A
silent miss on either reintroduces a build that exhausts RAM rather than
failing fast.

commit 390c8d3336687fea4288734ec7c28986a05e1576
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Aug 29 20:17:30 2026 +0200

sci-libs/composable-kernel: add 10.0.0 (requires AMD's LLVM fork)

Landed masked and not yet buildable: the packaging is right and the patch
triage is worth banking, but the build cannot finish against a vanilla LLVM 23.
llvm-core/rocm-llvm, later in this series, is what makes it build -- via
USE=amd-llvm on dev-util/hipcc.

Packaging is complete: tag retarget, and four of five patches resolved --
no-git-no-hash obsolete; upstream rewrote the block to degrade
gracefully without .git (plain find_package(Git), a
COMMIT_ID default of "unknown", RESULT_VARIABLE/
ERROR_QUIET) which is exactly what the patch did.
conditional-ckprofiler obsolete; upstream added its own BUILD_CK_PROFILER
option, so src_configure drives that instead of the
CK_USE_PROFILER flag our patch invented.
expand-isa 20 of 21 hunks are upstream at 10.0 -- AMD took the
Gentoo fix from ROCm/composable_kernel#775 and went
further (gfx1013, gfx10_1_generic). Only the
gridwise_gemm_dpp.hpp guard remains.
libcxx-includes still needed; regenerated across 41 of its 42 headers
(one no longer exists).
Plus a new clang23-buffer-load-types patch: clang 23 changed
__builtin_amdgcn_raw_buffer_load_b to RETURN and ..._store_b128 to
TAKE unsigned vectors, while CK declares signed ones.

WHY A VANILLA LLVM IS NOT ENOUGH: include/ck/utility/amd_wmma.hpp calls
__builtin_amdgcn_wmma_f32_16x16x16_bf16_w32 with bhalf16_t (__bf16 vector), but
vanilla clang 23 declares that builtin as taking short __attribute__((
ext_vector_type(16))). AMD builds ROCm against its OWN LLVM fork, where the
WMMA intrinsics use bf16 types. Fixing this in-tree would mean rewriting
intrinsic calls in numerical kernels, where a wrong bit_cast compiles clean and
silently produces wrong GPU results -- so the fork is packaged instead.

Also shortens the SRC_URI lines via MY_BASE.

commit d26e488d095396a8343a1053601f18f59982a8df
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Mon Jul 27 10:04:12 2026 +0200

*/*: drop dead Python targets from PYTHON_COMPAT

139 ebuilds across 65 packages still listed python3_10, python3_11 or
python3_13t. None of the three is a supported implementation any more:
gentoo's python-utils-r1 sets _PYTHON_ALL_IMPLS to python3_ plus
python3_t, and everything older falls into the branch commented
"implementations deprecated prior to EAPI 9 are fatal".

Inert today, which is why nothing caught it. Targets outside
_PYTHON_ALL_IMPLS are silently dropped rather than rejected, so they
never became USE flags and never produced a dependency atom -- an ebuild
declaring python3_ was only ever offering 12, 13 and 14.

Two reasons to clean them up anyway. The eclass branch they land in
dies for any EAPI other than 7 or 8, so all 139 are waiting to break on
the first EAPI 9 bump. And pkgcheck's OldPythonCompat is a git-scoped
check that only inspects ebuilds touched in a commit range, so it never
reports these on a full-tree scan and would instead fire on whichever
unrelated commit next touches one of them.

No revbump, because nothing observable changes. Verified on a sample
spanning every transformation shape that the python USE flags portage
offers are byte-identical before and after, and that a pkgcheck run
over all 65 packages returns a set-identical list of findings.

The eight python2_7 ebuilds are deliberately untouched: that target is
also absent from _PYTHON_ALL_IMPLS, but they reach it through the
_PYTHON_ALLOW_PY27 hatch in the overlay's forked python-utils-r1_py2.

commit 7c384c8cac70105c06770e0ccb7af0fc295f2471
Author: Raukaan Cogbrother <cogbrother@raukaan.local>
Date: Sun May 31 03:29:35 2026 +0200

metadata: normalize tab → 2-space indent in 42 imported files

42 metadata.xml files inherited tab indentation from their origin
overlays (mostly the ROCm cluster, plus a handful of dev-python and
AMD-tooling imports). The overlay convention is 2-space; this brings
them in line with the other 498 files. Maintainer attributions
preserved verbatim.

commit 5b3c001a7da4e04ed7804e3db133c11444942553
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 16:14:28 2026 +0200

sci-libs, dev-libs, dev-util: fix S for ROCm 7.2.4 tarball relayout

The rocm-7.2.4 per-component release-asset tarballs are re-rolled with a
lowercase top-level wrapper dir (hipblas/, rocblas/, rocr-runtime/, ...)
where the 7.2.3 assets exploded flat into WORKDIR. The bumped ebuilds
kept S="$", so src_prepare died applying PATCHES and running
seds against paths that had shifted one level down. Point S at the
wrapper dir each tarball actually unpacks to, matching the layout
hipBLASLt/hipsparselt already use. roct-thunk-interface's libhsakmt now
lives under the rocr-runtime/ wrapper too.

commit cb135ea508d4ca7fbd96144c9c706b671b8057a5
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 14:12:06 2026 +0200

sci-libs/composable-kernel: add 7.2.4

commit 91860d958300ed39de84aa0d849f326d2f6d0f4f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 8 20:18:00 2026 +0200

sci-libs/composable-kernel: drop stale patches superseded by 7.2.3

commit 1222acbdb747f7beb6c38ced833ece1b8d2bd817
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 6 10:56:09 2026 +0200

sci-libs/composable-kernel: 7.2.3 fix S= for flat upstream tarball

The 7.2.3 release-asset tarballs from ROCm/rocm-systems and ROCm/rocm-libraries
unpack contents directly into ./ instead of into a wrapping <assetname>/ subdir as in
7.2.0. Adjust S= so src_prepare finds the source tree.

commit f4d5f211eab6d48ce376575060d9728c53da8a02
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 6 01:58:14 2026 +0200

sci-libs/composable-kernel: add 7.2.3