sci-ml/caffe2
A deep learning framework
-
caffe2-2.14.0-r90~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1102 +amdgpu_targets_gfx1103 +amdgpu_targets_gfx1150 +amdgpu_targets_gfx1151 +amdgpu_targets_gfx1152 +amdgpu_targets_gfx1153 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031
View
Download
Browse License: BSD Overlay: stuff -
caffe2-2.13.0-r91~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx950 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1102 +amdgpu_targets_gfx1103 +amdgpu_targets_gfx1150 +amdgpu_targets_gfx1151 +amdgpu_targets_gfx1152 +amdgpu_targets_gfx1153 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031
View
Download
Browse License: BSD Overlay: stuff -
caffe2-2.13.0-r90~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151
View
Download
Browse License: BSD Overlay: stuff -
caffe2-2.12.0-r91~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151
View
Download
Browse License: BSD Overlay: stuff -
caffe2-2.12.0-r3~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151
View
Download
Browse License: BSD Overlay: gentoo -
caffe2-2.11.0-r90~amd64 ~arm64cuda cusparselt distributed fbgemm flash gloo kineto memefficient mimalloc mkl mpi nccl nnpack +numpy onednn openblas opencl openmp qnnpack rocm xnnpack python_single_target_python3_12 python_single_target_python3_13 python_single_target_python3_14 +amdgpu_targets_gfx908 +amdgpu_targets_gfx90a +amdgpu_targets_gfx942 +amdgpu_targets_gfx1030 +amdgpu_targets_gfx1100 +amdgpu_targets_gfx1101 +amdgpu_targets_gfx1200 +amdgpu_targets_gfx1201 amdgpu_targets_gfx803 amdgpu_targets_gfx900 amdgpu_targets_gfx906 amdgpu_targets_gfx940 amdgpu_targets_gfx941 amdgpu_targets_gfx1010 amdgpu_targets_gfx1011 amdgpu_targets_gfx1012 amdgpu_targets_gfx1031 amdgpu_targets_gfx1102 amdgpu_targets_gfx1103 amdgpu_targets_gfx1150 amdgpu_targets_gfx1151
View
Download
Browse License: BSD Overlay: stuff
ChangeLog
commit 8e514fc454b8105fc2a4439fc5ef164dd6f73c26
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Sep 12 14:54:30 2026 +0200
sci-ml/caffe2: pin CUTLASS for 2.13.0-r91
PyTorch 2.13.0 pins CUTLASS 4.4.2. CUTLASS 4.6.1 removed the
TileScheduler overload used by AsyncMM.cu, breaking CUDA compilation.
Limit the backport to the dependency pin so the current ebuild retains
its supported 7.5 fallback and existing cleanups.
commit a801173b7f6ed1c34e6b72014164c2de13bd38b6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:47 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.14.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 930ad200b90eb56d3095877cf81ac044294b4569
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:43 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.13.0-r91
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. CUDA configuration and an actual
translation-unit compile passed. A full staged install was not
completed.
Bug: https://github.com/istitov/stuff/issues/287
commit 94746e468d2055ebd5abe3391fe8fcc7e0667e25
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:41 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.13.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit cc42158e004e54a9e588807cbb3a55dc8abe81e2
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:38 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.12.0-r91
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 3cb83177927e7a517c0d611377a39d189d2127dd
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:35 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.11.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 6ffac649160cdbd969745670b5b39d4786cfa68a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit ea6c1c966f61fd3f2a5d4e2647422435e7db6aff
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 4663a4cb8b242c47608a3326263309a196731256
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 44cb4a12e53e57538b156f1464c6681b998f9847
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 60c55f7130475141b3472b068f815db0f1e6e3b6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 52c737f4832dce9705e56bc4d67b388650996940
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:19 2026 +0200
sci-ml/caffe2: resolve stale 2.14.0-r90 backend TODOs
commit 99d4f5fe95485dae5c132045c0e847aaf47dfd1e
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:19 2026 +0200
sci-ml/caffe2: resolve stale 2.13.0-r91 backend TODOs
commit 271d0e357f0f2f06f4576af90c8137d507f9353e
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.13.0-r90 backend TODOs
commit a1fce06756acae4ae57ade8ed053f1d8076f2ce0
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.12.0-r91 backend TODOs
commit 6db1d1b6302d01ca3dcb7f336d10e66902991de7
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.11.0-r90 backend TODOs
commit e5cde6d905877d96652b85e63a4a2443236083a6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:09:25 2026 +0200
sci-ml/caffe2: trim 2.11.0-r90 comment prose
commit 68d1ec78c265cbd71bd63c5e63d275218561a68f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:08:36 2026 +0200
sci-ml/caffe2: trim 2.12.0-r91 comment prose
commit ae69e6e1d71c37066ab385e2005e7e6729ce6c22
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:08:07 2026 +0200
sci-ml/caffe2: trim 2.13.0-r90 comment prose
commit 7247f056ec3faa86fd59329d8c660116b52936e2
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:07:28 2026 +0200
sci-ml/caffe2: trim 2.13.0-r91 comment prose
commit 8410ac8e9fe97454c92c390261ffd3b576ee4ceb
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:05:55 2026 +0200
sci-ml/caffe2: trim 2.14.0-r90 comment prose
commit dbee56dc1f40ab04027af3fdcdf65c2cde67a070
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 4 02:23:37 2026 +0200
sci-ml/caffe2: correct the hipFile rationale for 2.14.0-r90
The ebuild comment and the patch header both said hipFile is packaged
neither in ::gentoo nor here. sys-libs/hipFile has since landed in this
overlay, so the stated reason for dropping REQUIRED from
find_package(hipfile) was wrong even though the conclusion was right.
The actual reason: PyTorch's USE_CUFILE option is still CUDA-only and
the ROCm build consumes no hipfile target, so requiring it would pull
hipFile and rocprofiler-register into every ROCm dependency closure
without buying any GDS support. Optional discovery finds an installed
hipFile and otherwise leaves the already-disabled USE_CUFILE path
untouched.
Prose only -- the patch hunks are byte-identical, so no revbump.
commit 956c384b4b09287afe613d8d00c45c082ff9c0b9
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 3 01:54:15 2026 +0200
sci-ml/caffe2: add 2.14.0-r90
Forked from 2.13.0-r91 onto pytorch 2.14.0. Build-verified on haarmek
with the host USE set (rocm memefficient distributed fbgemm mkl gloo mpi
nnpack onednn xnnpack qnnpack openblas), gfx1150, ROCm 10.0 / HIP 7.15.
Not just a green build: USE_ROCM ON with ROCM_VERSION 10.0.0 and
TORCH_HIP_VERSION 715, USE_ROCM_CK_GEMM ON, "Using Preinstalled AOTriton
at /usr", libtorch_hip.so installed to /usr/lib64 and linked against the
real HIP stack, _C.cpython-313 relocated into site-packages/torch, and a
914.7 MiB installed tree -- so no silent CPU-only degradation.
Nineteen patches: eight $-named rebases, one new, ten shared files
carried unchanged. Of the eight, three did not apply to 2.14.0 at all
(unbundle_pocketfft, gentoo, removekineto) and five applied only with
fuzz, two of them at fuzz 2. eapply tolerates fuzz, so those would have
built, but a guessed hunk is not a verified one; all nineteen now apply
with zero failures, zero fuzz, zero rejects.
The rebase baseline is NOT the raw tarball -- src_prepare's own sed pass
runs before cmake_src_prepare applies PATCHES and rewrites the very
files the patches touch. A first attempt against the unmodified tarball
verified clean and then died in the real prepare phase; the PATCHES
block now says so.
The rebased gentoo patch also REPAIRS a hunk its predecessor lost rather
than carrying the loss forward. caffe2-2.10.0-gentoo.patch has a bare
blank line between its second and third CMakeLists.txt hunks, and GNU
patch stops parsing that file's hunks there -- so the hunk dropping
`append_cxx_flag_if_supported("-Werror=format")` has never applied on
any caffe2 version, silently, because patch still exits 0. Deleting the
blank line makes it apply at 1221. The 2.14.0 patch carries the hunk;
configure now shows no -Werror=format in CXX flags where the first
2.14.0 build had it. The older shared patch is untouched here -- fixing
it changes four shipping ebuilds and wants its own revbumped commit.
Three upstream changes drove the rest:
- composable_kernel moves to 5a74dec0, the commit 2.14.0 pins, and the
gfx101x ISA patch collapses from twelve hunks to one: upstream CK has
taken everything except kernel_gemm_dpp's guard.
- 2.14.0 pins __AOTRITON_VER 0.13b itself, so the aotriton atom no
longer diverges from upstream -- it now simply agrees.
- 2.14.0 adds a REQUIRED find_package(hipfile) for ROCm >= 7.14.
hipFile is unpackaged (its repo is retired into ROCm/rocm-systems) and
the REQUIRED makes upstream's own `if(USE_CUFILE AND NOT
hipfile_FOUND)` fallback unreachable. Dropping only the REQUIRED lets
that fallback run: configure reports "Optional package hipfile not
found" and USE_CUFILE OFF.
src_install hardening, both cases the 2.13.0 forge learned the hard way:
the torch._C relocation now dies instead of silently skipping when the
glob misses (an unmatched glob leaves the literal pattern behind, the
whitelist purge then deletes the extension, and the package merges green
while `import torch` dies at runtime), and the /usr/lib symlink purge is
gated on get_libdir so a libdir=lib profile cannot have its entire
payload deleted.
The FBGEMM <1.5 cap is kept but its justification changed: every
fbgemm::Quantize call site in 2.14.0 uses the 1-arg form, and 1.4
declares `template <typename T, bool LEGACY = true>`, so the API no
longer forces the cap. Only 1.4 is build-verified here, so lifting it
needs a check against >=1.7 rather than the API reading alone.
Both rebased patches that lost their predecessor's explanatory header --
rocm-distributed-link and removekineto-pr178960 -- carry it again. The
first regains its subject and the upstream bug link (pytorch#158725),
plus a note that it fixes a RUNTIME undefined-symbol failure (rsmi_init)
that no build-check can validate. The second was a three-commit
git-format-patch series that the rebase had to squash, so its header now
enumerates the three fixes against the files they touch.
commit 286c1c5ffb7b7707441d55ea0a4031f4f439c55a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 18:57:20 2026 +0200
sci-ml/caffe2: drop the dead AOTRITON_* variables, move the rationale
AOTRITON_PV, AOTRITON_PN, AOTRITON_P and AOTRITON_tar were defined and never
read: they appear in no SRC_URI, no dependency, and no phase function. The
ebuild's own comment conceded they were "documentation only". Left in place
they read like a functional version pin, so a future bump would spend effort
keeping a string in sync that nothing consumes, and anyone diffing r90 against
r91 would look for the aotriton pin in the wrong place.
The pin is the =sci-libs/aotriton-bin-0.13*:= atom in RDEPEND, so the
rationale moved there -- above the RDEPEND= line, since a '#' inside a
dependency string is parsed as a package atom rather than a comment.
No functional change: metadata regenerates and `emerge -p` resolves the
package identically, with the same USE and AMDGPU_TARGETS.
commit a0bc52523d25a05540fe74d81521ce05e8a49498
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 17:14:07 2026 +0200
sci-ml/caffe2: add 2.13.0-r91, built against ROCm 10.0
Relax the eighteen <...-7.3 ROCm caps to <...-11 and move ROCM_VERSION 6.1 ->
10.0. The caps dated from when AMD numbered releases 7.x; ROCm 10.0 is that
same line renumbered (hip-10.0.0 reports HIP 7.15.0), so <7.3 excluded it by
accident of numbering. ROCM_VERSION only selects rocm.eclass's target-list
branch for THIS package's amdgpu_targets_* IUSE and ROCM_REQUIRED_USE -- the
10.* branch is a superset of 6.1's, adding gfx1152/gfx1153 and promoting
gfx1102/1103/1150/1151 to official. Nothing here uses ROCM_USEDEP, so it does
not constrain the ROCm libraries themselves.
memefficient now wants =sci-libs/aotriton-bin-0.13*, and this DIVERGES FROM
UPSTREAM on purpose: pytorch v2.13.0's cmake/External/aotriton.cmake sets
__AOTRITON_VER "0.12b" with __AOTRITON_ROCM_LIST = rocm6.4/7.0/7.1/7.2. But
0.12b ships no rocm7.14/7.15 shim at all (verified against its release
assets), so caffe2[memefficient] cannot be built against the 10.0 stack with
it. 0.13b adds rocm7.15 and keeps the libaotriton_v2 ABI.
r90 is kept: it is the ebuild for a 7.2.x stack, whose libraries have the
narrower 7.* arch list.
Build-verified on gfx1150 with USE="rocm memefficient" against the installed
10.0 closure. Note this needed llvm-core/rocm-llvm to stop shipping an
assertions build first -- see that commit; the failure was a clang abort in
rocPRIM's merge-sort kernels, not anything in caffe2.
commit ce0b61a5c61be8ab69d77a2ddeec2b5e18495e6a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 15:01:40 2026 +0200
sci-libs/aotriton-bin: wire the ROCm 10.0 shim for 0.13b
The 6.4/7.0/7.1/7.2 -> 7.14/7.15 gap in upstream's shim set is AMD renumbering
its releases, not missing releases: the 7.14/7.15 line became ROCm 10.0, and
dev-util/hip-10.0.0 reports HIP version 7.15.0 via `hipconfig --version`. So
rocm7.15 is the shim for hip-10.0, and the earlier note here -- that these were
"deliberately not wired ... until those HIP versions are confirmed to exist" --
is now settled.
src_unpack selects the archive by the installed hip's $(ver_cut 1-2), which is
"10.0", so fetch the rocm7.15 asset under a -rocm10.0- name. Cap moves to
<dev-util/hip-11 accordingly.
Also make the selection loud. The RDEPEND range is necessarily wider than the
set of shims SRC_URI lists, since upstream skips HIP versions; an unmatched
version previously unpacked only the gfx images and produced a package with no
libaotriton_v2.so in it, with nothing failing until something tried to link
against it. Now it dies in src_unpack naming the version it wanted.
0.12b is deliberately left alone: it ships no 7.14/7.15 shim at all (verified
against its release assets), so its <dev-util/hip-7.3 cap is accurate.
Verified against hip-10.0.0: libaotriton_v2.so.0.13.0 installs at 19.7 MB with
the amd-gfx115x images, i.e. the shim really is selected.
commit 4549e288fc4e2c10ad4e099e99a6a91f31647e49
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 09:41:47 2026 +0200
sci-ml/caffe2: normalize copyright header
Use the overlay-wide required copyright range consistently across retained versions.
commit 92ff8d876330c20911a390ee46f7c08248bd7687
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 08:43:54 2026 +0200
sci-ml/caffe2: block older pytorch before 2.13 merge
Ownership of torch._C moves into caffe2 in 2.13. Force older pytorch frontends out
before the replacement file is merged to prevent a deterministic collision during
upgrades.
commit e0498d45b0529f310260b5800f5226c5bc440984
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 7 17:41:51 2026 +0200
sci-ml/caffe2: cap FBGEMM below 1.5 on frozen forks; fix 2.13 configure
FBGEMM 1.7 dropped a template parameter from fbgemm::Quantize (Quantize<T,
LEGACY> -> Quantize<T>). The frozen pytorch 2.11/2.12/2.13 sources still call
the 2-arg form in aten/src/ATen/native/QuantizedLinear.cpp, so rebuilding any
of them against 1.7 fails to compile. FBGEMM's only consumer here is caffe2, so
cap fbgemm? to <sci-ml/FBGEMM-1.5 on 2.11.0-r90, 2.12.0-r91 and 2.13.0-r90 to
keep them rebuildable without patching source.
2.13 also moved setup.py's file mirroring into cmake/FileMirroring.cmake, which
FATAL_ERRORs on CUDA builds when the cutlass submodule's Blackwell CuTeDSL
grouped_gemm.py is absent -- it is, since this fork builds against system
dev-libs/cutlass. That template is an optional torch/_inductor vendored kernel;
downgrade the sole FATAL_ERROR to STATUS so the unbundled build configures.
commit 3569c06e03e9fcd976ca1b744ffa39310eda3b15
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Jul 22 22:51:06 2026 +0200
sci-ml/caffe2: relocate torch _C, drop 2.13.0 python-package leak
2.13.0 folds the torch Python package install into caffe2's cmake:
torch/CMakeLists.txt installs the torch._C extension and version.py with
DESTINATION ".", and cmake/PackageData.cmake installs the type stubs the
same way. Those destinations are relative to CMAKE_INSTALL_PREFIX, which
upstream expects to be the wheel's <root>/torch but which is /usr here, so
the whole payload lands in /usr/ root and also leaves stray symlinks in
/usr/lib mirroring the host's /usr/lib64 top-level symlinks --
/usr/lib/terminfo collides with sys-libs/ncurses and aborts the merge.
The compiled torch._C extension is the one piece that must be kept: 2.13.0
moved its build from setuptools into cmake, and sci-ml/pytorch skips the
cmake build, so nothing else produces it. Without it `import torch` loads
the torch/_C stub directory instead of the extension and fails. Relocate
_C into site-packages/torch (its NEEDED libtorch* resolve from /usr/lib64
on ldconfig's default path) and drop the rest of the leaked payload -- the
pure-python package is installed by sci-ml/pytorch.
commit 98d9f5848da3641f9cd5da269d76f28e980dcddb
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Jul 18 16:03:48 2026 +0200
sci-ml/caffe2: add 2.13.0-r90
PyTorch 2.13.0. Carries the 2.12.0-r91 patch stack forward unchanged plus
four new-in-2.13.0 build fixes, and depends on aotriton-bin-0.12*:
- prebuildsteps-tarball-guard: 2.13.0 moved setup.py's submodule
completeness check into cmake/PreBuildSteps.cmake, which FATAL_ERRORs on
the (intentionally empty) third_party submodules of a source-tarball
build; guard the whole block on the presence of .git.
- glog-exception-init: 2.13.0 duplicated the "is glog initialized" hack
into c10/util/Exception.cpp, still calling the internal
glog_internal_namespace_::IsGoogleLoggingInitialized(); system glog 0.6.0
only exports the public ::google::IsGoogleLoggingInitialized() (mirror
the existing Logging.cpp glog patch).
- wrap-headers-destdir: cmake/PostBuildSteps.cmake runs wrap_headers.py via
install(CODE) against $/include with no $,
writing to the live /usr instead of the image; prepend $ENV.
- HIP_CLANG_PATH: 2.13.0's LoadHIP.cmake hardcodes the HIP compiler at
$/lib/llvm/bin, but Gentoo slots llvm; export it from
hipconfig -l in the rocm configure path.
Plus an fmt sed (drop the new FMT_NO_UNIQUE_ADDRESS compile-defs on the
unbundled fmt target) and a src_install cleanup (drop the aotriton libs
2.13.0 bundles into /usr/lib via hardcoded DESTINATION "lib" -- trips
multilib-strict and duplicates sci-libs/aotriton-bin).
commit bfb3f2583d4f86577dcc789426a0861b34b75439
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Jun 17 21:48:26 2026 +0200
sci-ml/caffe2: Keyword 2.12.0-r91 for ~arm64
2.11.0-r90 already carried ~arm64 but 2.12.0-r91 had dropped it. The
C++ libtorch core builds CPU-only on aarch64 (default USE: no
cuda/rocm/fbgemm — fbgemm is x86-AVX-only), so restore ~arm64 and keep
the newest in sync. C++ (CMake); aarch64-portable.
commit 5653f839b39ea876b1a0e1ab42c0e7adc43b59e8
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Tue Jun 2 12:58:53 2026 +0200
sci-ml/caffe2: fix gcc-15 -Wtemplate-body in List_inl.h
gcc-15 enforces -Wtemplate-body in CUDA .cu compilation, rejecting
typename decltype(impl_->list)::difference_type as a dependent scope.
Replace with std::ptrdiff_t which is equivalent for std::vector iterators.
commit 7c384c8cac70105c06770e0ccb7af0fc295f2471
Author: Raukaan Cogbrother <cogbrother@raukaan.local>
Date: Sun May 31 03:29:35 2026 +0200
metadata: normalize tab → 2-space indent in 42 imported files
42 metadata.xml files inherited tab indentation from their origin
overlays (mostly the ROCm cluster, plus a handful of dev-python and
AMD-tooling imports). The overlay convention is 2-space; this brings
them in line with the other 498 files. Maintainer attributions
preserved verbatim.
commit 6b7192253e216334bb27f7c41b79803fc8ac8a36
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 09:37:30 2026 +0200
sci-ml/caffe2: relax pybind11 cap to <3.0.5
The <3.0.2 cap was inherited from the 2.11.0-era fork; ::gentoo's 2.12.0-r2 allows
<3.0.5 for the same upstream, and the stale cap blocked pybind11 from upgrading past
3.0.1.
commit b32aff8fbb037bea69577857ec2da162856a8e4f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 01:26:23 2026 +0200
sci-ml/caffe2: fix xz-compressed patches breaking src_prepare
e8b75c54 compressed files/*.patch to clear a TotalSizeViolation, but eapply does not
decompress patches, so each one failed on the raw xz bytes. Decompress them into $
and repoint PATCHES in src_prepare, keeping files/ compressed so the cap stays
satisfied.
commit e8b75c54469f970239ba45e0cd17beb8159b8f9a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 27 15:02:11 2026 +0200
sci-ml/caffe2: xz-compress files/*.patch to clear TotalSizeViolation
eapply decompresses .patch.xz transparently. Compression takes the
files/ payload from 58.2 KiB to 21.5 KiB (well under pkgcheck's 50
KiB threshold). Chosen over the extra-stuff bundle migration to keep
the patches in-tree; per-patch edits become 'xz -d' → edit → 'xz -9'
but the workflow stays single-repo.
commit e348703397f16e5622b4834a13833dc79ccc35c8
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu May 21 14:27:48 2026 +0200
sci-ml/caffe2: add 2.12.0-r90
Fork from ::gentoo 2.12.0-r1 (which drops rocm-fix-std-cpp17.patch to
resolve the C++20 .contains() failure under gcc-16). Carry forward our
overlay-only mkl-public-scrub patch:
- mkl-public-scrub: scrub MKL MPI/cluster libs and force GNU OpenMP
threading in caffe2::mkl's public link interface (ported from 2.11.0;
cmake/public/mkl.cmake context identical at 2.12.0).
Also: drop python3_11 from PYTHON_COMPAT and arm64 from KEYWORDS (not
tested in this overlay); drop nonexistent !sci-libs/QNNPACK blocker.
commit 253cd7b826e77eb7e68d1d3548d5ac9d36ab4ccc
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 13 14:35:42 2026 +0200
sci-ml/caffe2: disable py3.11
commit f1e7ab9cddb007e7de2e4e4e87ffedd361b212ad
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat May 9 00:44:08 2026 +0200
sci-ml/caffe2: drop dead !sci-libs/QNNPACK blocker
QNNPACK has never existed as a separate Gentoo package (verified
2026-05-09 — no entry in ::gentoo's sci-libs/, no diff-filter=A hits in
::gentoo history). Upstream PyTorch vendors it and exposes the build
toggle as USE_PYTORCH_QNNPACK, which we already wire via the qnnpack
USE flag. The blocker was protecting against a separate package that
never landed. pkgcheck NonexistentBlocker.
commit f19ffa2a53036faae2f64194c8b2f7c7c15e3fb0
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 8 19:29:35 2026 +0200
sci-ml/caffe2: drop two orphan patches inherited from ::gentoo fork
Both patches target older PVs (2.6.0 / 2.10.0) that 2.11.0-r90's
PATCHES= no longer references; the array uses $- variants for the
relevant fixes (mimalloc, rocm-fix-std-cpp17). Came along when forking
from ::gentoo's -r3 to add the MKL public-link scrub.
commit 78e2fba68899c5e58faba344002934108b58e32c
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 8 01:25:42 2026 +0200
sci-ml/caffe2: fork ::gentoo's 2.11.0-r3 to -r90 with MKL public-link scrub
::gentoo's caffe2 ships caffe2/cmake/public/mkl.cmake unchanged from
upstream pytorch — calls find_package(MKL) and dumps the full result
into caffe2::mkl's INTERFACE_LINK_LIBRARIES with no filter. On hosts
with Intel oneAPI installed, the resolver is Intel's own MKLConfig.cmake
which by default returns the full HPC / Cluster Edition lib set,
including libmkl_scalapack_ilp64, libmkl_cdft_core,
libmkl_blacs_intelmpi_ilp64 (all MPI-distributed), and libmkl_intel_thread
(Intel-OpenMP threading layer). Those libs only exist when Intel's
separate Cluster + Compiler oneAPI packages are also installed; the
basic intel-oneapi-mkl package omits them.
Because they end up in caffe2::mkl's public link interface, every
downstream consumer that links against torch::torch (vllm's
cumem_allocator, custom torch C++ extensions, etc.) inherits them
and the link fails with "cannot find -lmkl_scalapack_ilp64" etc.
torch::torch's public APIs never reach BLACS / ScaLAPACK / distributed
FFT — distributed-tensor paths use NCCL / Gloo / MPI directly, not
MKL's BLACS. So filtering them is purely subtraction of incorrect
public-link-interface clutter. Forcing MKL_THREADING=gnu_thread before
the find pulls libmkl_gnu_thread (always available) instead of
libmkl_intel_thread, and pairs cleanly with system libgomp — also
avoids the multi-OpenMP-runtime mixing trap (libgomp + libiomp5 in
the same process can oversubscribe / deadlock).
Verified 2026-05-08 against caffe2-2.11.0 with USE="cuda distributed
fbgemm gloo memefficient mkl mpi nnpack numpy onednn openblas opencl
openmp": full rebuild + import torch succeed identically; downstream
consumers (vllm USE=cuda's cumem_allocator extension) link cleanly.
Upstream pytorch's cmake/public/mkl.cmake on main still has the same
unfiltered behaviour as of 2026-05-08; no upstream PR open for this
specific scrub. Drop this -r90 fork when an equivalent upstream fix
lands.
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Sep 12 14:54:30 2026 +0200
sci-ml/caffe2: pin CUTLASS for 2.13.0-r91
PyTorch 2.13.0 pins CUTLASS 4.4.2. CUTLASS 4.6.1 removed the
TileScheduler overload used by AsyncMM.cu, breaking CUDA compilation.
Limit the backport to the dependency pin so the current ebuild retains
its supported 7.5 fallback and existing cleanups.
commit a801173b7f6ed1c34e6b72014164c2de13bd38b6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:47 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.14.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 930ad200b90eb56d3095877cf81ac044294b4569
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:43 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.13.0-r91
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. CUDA configuration and an actual
translation-unit compile passed. A full staged install was not
completed.
Bug: https://github.com/istitov/stuff/issues/287
commit 94746e468d2055ebd5abe3391fe8fcc7e0667e25
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:41 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.13.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit cc42158e004e54a9e588807cbb3a55dc8abe81e2
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:38 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.12.0-r91
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 3cb83177927e7a517c0d611377a39d189d2127dd
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 11 23:55:35 2026 +0200
sci-ml/caffe2: use supported CUDA fallback in 2.11.0-r90
The ebuild requires CUDA 12.9 or newer, which no longer accepts
compute capability 3.5. Its unconfigured fallback therefore failed at
the first CUDA compile.
Use 7.5, the oldest target also accepted by the current CUDA 13 series,
while preserving user overrides. The identical CUDA path was exercised
on 2.13.0-r91; this revision was source-checked only and did not receive
a separate staged install.
Bug: https://github.com/istitov/stuff/issues/287
commit 6ffac649160cdbd969745670b5b39d4786cfa68a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit ea6c1c966f61fd3f2a5d4e2647422435e7db6aff
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 4663a4cb8b242c47608a3326263309a196731256
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 44cb4a12e53e57538b156f1464c6681b998f9847
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 60c55f7130475141b3472b068f815db0f1e6e3b6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:49:31 2026 +0200
sci-ml/caffe2: tighten warning workaround comments
commit 52c737f4832dce9705e56bc4d67b388650996940
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:19 2026 +0200
sci-ml/caffe2: resolve stale 2.14.0-r90 backend TODOs
commit 99d4f5fe95485dae5c132045c0e847aaf47dfd1e
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:19 2026 +0200
sci-ml/caffe2: resolve stale 2.13.0-r91 backend TODOs
commit 271d0e357f0f2f06f4576af90c8137d507f9353e
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.13.0-r90 backend TODOs
commit a1fce06756acae4ae57ade8ed053f1d8076f2ce0
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.12.0-r91 backend TODOs
commit 6db1d1b6302d01ca3dcb7f336d10e66902991de7
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 09:12:18 2026 +0200
sci-ml/caffe2: resolve stale 2.11.0-r90 backend TODOs
commit e5cde6d905877d96652b85e63a4a2443236083a6
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:09:25 2026 +0200
sci-ml/caffe2: trim 2.11.0-r90 comment prose
commit 68d1ec78c265cbd71bd63c5e63d275218561a68f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:08:36 2026 +0200
sci-ml/caffe2: trim 2.12.0-r91 comment prose
commit ae69e6e1d71c37066ab385e2005e7e6729ce6c22
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:08:07 2026 +0200
sci-ml/caffe2: trim 2.13.0-r90 comment prose
commit 7247f056ec3faa86fd59329d8c660116b52936e2
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:07:28 2026 +0200
sci-ml/caffe2: trim 2.13.0-r91 comment prose
commit 8410ac8e9fe97454c92c390261ffd3b576ee4ceb
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 10 00:05:55 2026 +0200
sci-ml/caffe2: trim 2.14.0-r90 comment prose
commit dbee56dc1f40ab04027af3fdcdf65c2cde67a070
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Sep 4 02:23:37 2026 +0200
sci-ml/caffe2: correct the hipFile rationale for 2.14.0-r90
The ebuild comment and the patch header both said hipFile is packaged
neither in ::gentoo nor here. sys-libs/hipFile has since landed in this
overlay, so the stated reason for dropping REQUIRED from
find_package(hipfile) was wrong even though the conclusion was right.
The actual reason: PyTorch's USE_CUFILE option is still CUDA-only and
the ROCm build consumes no hipfile target, so requiring it would pull
hipFile and rocprofiler-register into every ROCm dependency closure
without buying any GDS support. Optional discovery finds an installed
hipFile and otherwise leaves the already-disabled USE_CUFILE path
untouched.
Prose only -- the patch hunks are byte-identical, so no revbump.
commit 956c384b4b09287afe613d8d00c45c082ff9c0b9
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu Sep 3 01:54:15 2026 +0200
sci-ml/caffe2: add 2.14.0-r90
Forked from 2.13.0-r91 onto pytorch 2.14.0. Build-verified on haarmek
with the host USE set (rocm memefficient distributed fbgemm mkl gloo mpi
nnpack onednn xnnpack qnnpack openblas), gfx1150, ROCm 10.0 / HIP 7.15.
Not just a green build: USE_ROCM ON with ROCM_VERSION 10.0.0 and
TORCH_HIP_VERSION 715, USE_ROCM_CK_GEMM ON, "Using Preinstalled AOTriton
at /usr", libtorch_hip.so installed to /usr/lib64 and linked against the
real HIP stack, _C.cpython-313 relocated into site-packages/torch, and a
914.7 MiB installed tree -- so no silent CPU-only degradation.
Nineteen patches: eight $-named rebases, one new, ten shared files
carried unchanged. Of the eight, three did not apply to 2.14.0 at all
(unbundle_pocketfft, gentoo, removekineto) and five applied only with
fuzz, two of them at fuzz 2. eapply tolerates fuzz, so those would have
built, but a guessed hunk is not a verified one; all nineteen now apply
with zero failures, zero fuzz, zero rejects.
The rebase baseline is NOT the raw tarball -- src_prepare's own sed pass
runs before cmake_src_prepare applies PATCHES and rewrites the very
files the patches touch. A first attempt against the unmodified tarball
verified clean and then died in the real prepare phase; the PATCHES
block now says so.
The rebased gentoo patch also REPAIRS a hunk its predecessor lost rather
than carrying the loss forward. caffe2-2.10.0-gentoo.patch has a bare
blank line between its second and third CMakeLists.txt hunks, and GNU
patch stops parsing that file's hunks there -- so the hunk dropping
`append_cxx_flag_if_supported("-Werror=format")` has never applied on
any caffe2 version, silently, because patch still exits 0. Deleting the
blank line makes it apply at 1221. The 2.14.0 patch carries the hunk;
configure now shows no -Werror=format in CXX flags where the first
2.14.0 build had it. The older shared patch is untouched here -- fixing
it changes four shipping ebuilds and wants its own revbumped commit.
Three upstream changes drove the rest:
- composable_kernel moves to 5a74dec0, the commit 2.14.0 pins, and the
gfx101x ISA patch collapses from twelve hunks to one: upstream CK has
taken everything except kernel_gemm_dpp's guard.
- 2.14.0 pins __AOTRITON_VER 0.13b itself, so the aotriton atom no
longer diverges from upstream -- it now simply agrees.
- 2.14.0 adds a REQUIRED find_package(hipfile) for ROCm >= 7.14.
hipFile is unpackaged (its repo is retired into ROCm/rocm-systems) and
the REQUIRED makes upstream's own `if(USE_CUFILE AND NOT
hipfile_FOUND)` fallback unreachable. Dropping only the REQUIRED lets
that fallback run: configure reports "Optional package hipfile not
found" and USE_CUFILE OFF.
src_install hardening, both cases the 2.13.0 forge learned the hard way:
the torch._C relocation now dies instead of silently skipping when the
glob misses (an unmatched glob leaves the literal pattern behind, the
whitelist purge then deletes the extension, and the package merges green
while `import torch` dies at runtime), and the /usr/lib symlink purge is
gated on get_libdir so a libdir=lib profile cannot have its entire
payload deleted.
The FBGEMM <1.5 cap is kept but its justification changed: every
fbgemm::Quantize call site in 2.14.0 uses the 1-arg form, and 1.4
declares `template <typename T, bool LEGACY = true>`, so the API no
longer forces the cap. Only 1.4 is build-verified here, so lifting it
needs a check against >=1.7 rather than the API reading alone.
Both rebased patches that lost their predecessor's explanatory header --
rocm-distributed-link and removekineto-pr178960 -- carry it again. The
first regains its subject and the upstream bug link (pytorch#158725),
plus a note that it fixes a RUNTIME undefined-symbol failure (rsmi_init)
that no build-check can validate. The second was a three-commit
git-format-patch series that the rebase had to squash, so its header now
enumerates the three fixes against the files they touch.
commit 286c1c5ffb7b7707441d55ea0a4031f4f439c55a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 18:57:20 2026 +0200
sci-ml/caffe2: drop the dead AOTRITON_* variables, move the rationale
AOTRITON_PV, AOTRITON_PN, AOTRITON_P and AOTRITON_tar were defined and never
read: they appear in no SRC_URI, no dependency, and no phase function. The
ebuild's own comment conceded they were "documentation only". Left in place
they read like a functional version pin, so a future bump would spend effort
keeping a string in sync that nothing consumes, and anyone diffing r90 against
r91 would look for the aotriton pin in the wrong place.
The pin is the =sci-libs/aotriton-bin-0.13*:= atom in RDEPEND, so the
rationale moved there -- above the RDEPEND= line, since a '#' inside a
dependency string is parsed as a package atom rather than a comment.
No functional change: metadata regenerates and `emerge -p` resolves the
package identically, with the same USE and AMDGPU_TARGETS.
commit a0bc52523d25a05540fe74d81521ce05e8a49498
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 17:14:07 2026 +0200
sci-ml/caffe2: add 2.13.0-r91, built against ROCm 10.0
Relax the eighteen <...-7.3 ROCm caps to <...-11 and move ROCM_VERSION 6.1 ->
10.0. The caps dated from when AMD numbered releases 7.x; ROCm 10.0 is that
same line renumbered (hip-10.0.0 reports HIP 7.15.0), so <7.3 excluded it by
accident of numbering. ROCM_VERSION only selects rocm.eclass's target-list
branch for THIS package's amdgpu_targets_* IUSE and ROCM_REQUIRED_USE -- the
10.* branch is a superset of 6.1's, adding gfx1152/gfx1153 and promoting
gfx1102/1103/1150/1151 to official. Nothing here uses ROCM_USEDEP, so it does
not constrain the ROCm libraries themselves.
memefficient now wants =sci-libs/aotriton-bin-0.13*, and this DIVERGES FROM
UPSTREAM on purpose: pytorch v2.13.0's cmake/External/aotriton.cmake sets
__AOTRITON_VER "0.12b" with __AOTRITON_ROCM_LIST = rocm6.4/7.0/7.1/7.2. But
0.12b ships no rocm7.14/7.15 shim at all (verified against its release
assets), so caffe2[memefficient] cannot be built against the 10.0 stack with
it. 0.13b adds rocm7.15 and keeps the libaotriton_v2 ABI.
r90 is kept: it is the ebuild for a 7.2.x stack, whose libraries have the
narrower 7.* arch list.
Build-verified on gfx1150 with USE="rocm memefficient" against the installed
10.0 closure. Note this needed llvm-core/rocm-llvm to stop shipping an
assertions build first -- see that commit; the failure was a clang abort in
rocPRIM's merge-sort kernels, not anything in caffe2.
commit ce0b61a5c61be8ab69d77a2ddeec2b5e18495e6a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sun Aug 30 15:01:40 2026 +0200
sci-libs/aotriton-bin: wire the ROCm 10.0 shim for 0.13b
The 6.4/7.0/7.1/7.2 -> 7.14/7.15 gap in upstream's shim set is AMD renumbering
its releases, not missing releases: the 7.14/7.15 line became ROCm 10.0, and
dev-util/hip-10.0.0 reports HIP version 7.15.0 via `hipconfig --version`. So
rocm7.15 is the shim for hip-10.0, and the earlier note here -- that these were
"deliberately not wired ... until those HIP versions are confirmed to exist" --
is now settled.
src_unpack selects the archive by the installed hip's $(ver_cut 1-2), which is
"10.0", so fetch the rocm7.15 asset under a -rocm10.0- name. Cap moves to
<dev-util/hip-11 accordingly.
Also make the selection loud. The RDEPEND range is necessarily wider than the
set of shims SRC_URI lists, since upstream skips HIP versions; an unmatched
version previously unpacked only the gfx images and produced a package with no
libaotriton_v2.so in it, with nothing failing until something tried to link
against it. Now it dies in src_unpack naming the version it wanted.
0.12b is deliberately left alone: it ships no 7.14/7.15 shim at all (verified
against its release assets), so its <dev-util/hip-7.3 cap is accurate.
Verified against hip-10.0.0: libaotriton_v2.so.0.13.0 installs at 19.7 MB with
the amd-gfx115x images, i.e. the shim really is selected.
commit 4549e288fc4e2c10ad4e099e99a6a91f31647e49
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 09:41:47 2026 +0200
sci-ml/caffe2: normalize copyright header
Use the overlay-wide required copyright range consistently across retained versions.
commit 92ff8d876330c20911a390ee46f7c08248bd7687
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 14 08:43:54 2026 +0200
sci-ml/caffe2: block older pytorch before 2.13 merge
Ownership of torch._C moves into caffe2 in 2.13. Force older pytorch frontends out
before the replacement file is merged to prevent a deterministic collision during
upgrades.
commit e0498d45b0529f310260b5800f5226c5bc440984
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri Aug 7 17:41:51 2026 +0200
sci-ml/caffe2: cap FBGEMM below 1.5 on frozen forks; fix 2.13 configure
FBGEMM 1.7 dropped a template parameter from fbgemm::Quantize (Quantize<T,
LEGACY> -> Quantize<T>). The frozen pytorch 2.11/2.12/2.13 sources still call
the 2-arg form in aten/src/ATen/native/QuantizedLinear.cpp, so rebuilding any
of them against 1.7 fails to compile. FBGEMM's only consumer here is caffe2, so
cap fbgemm? to <sci-ml/FBGEMM-1.5 on 2.11.0-r90, 2.12.0-r91 and 2.13.0-r90 to
keep them rebuildable without patching source.
2.13 also moved setup.py's file mirroring into cmake/FileMirroring.cmake, which
FATAL_ERRORs on CUDA builds when the cutlass submodule's Blackwell CuTeDSL
grouped_gemm.py is absent -- it is, since this fork builds against system
dev-libs/cutlass. That template is an optional torch/_inductor vendored kernel;
downgrade the sole FATAL_ERROR to STATUS so the unbundled build configures.
commit 3569c06e03e9fcd976ca1b744ffa39310eda3b15
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Jul 22 22:51:06 2026 +0200
sci-ml/caffe2: relocate torch _C, drop 2.13.0 python-package leak
2.13.0 folds the torch Python package install into caffe2's cmake:
torch/CMakeLists.txt installs the torch._C extension and version.py with
DESTINATION ".", and cmake/PackageData.cmake installs the type stubs the
same way. Those destinations are relative to CMAKE_INSTALL_PREFIX, which
upstream expects to be the wheel's <root>/torch but which is /usr here, so
the whole payload lands in /usr/ root and also leaves stray symlinks in
/usr/lib mirroring the host's /usr/lib64 top-level symlinks --
/usr/lib/terminfo collides with sys-libs/ncurses and aborts the merge.
The compiled torch._C extension is the one piece that must be kept: 2.13.0
moved its build from setuptools into cmake, and sci-ml/pytorch skips the
cmake build, so nothing else produces it. Without it `import torch` loads
the torch/_C stub directory instead of the extension and fails. Relocate
_C into site-packages/torch (its NEEDED libtorch* resolve from /usr/lib64
on ldconfig's default path) and drop the rest of the leaked payload -- the
pure-python package is installed by sci-ml/pytorch.
commit 98d9f5848da3641f9cd5da269d76f28e980dcddb
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat Jul 18 16:03:48 2026 +0200
sci-ml/caffe2: add 2.13.0-r90
PyTorch 2.13.0. Carries the 2.12.0-r91 patch stack forward unchanged plus
four new-in-2.13.0 build fixes, and depends on aotriton-bin-0.12*:
- prebuildsteps-tarball-guard: 2.13.0 moved setup.py's submodule
completeness check into cmake/PreBuildSteps.cmake, which FATAL_ERRORs on
the (intentionally empty) third_party submodules of a source-tarball
build; guard the whole block on the presence of .git.
- glog-exception-init: 2.13.0 duplicated the "is glog initialized" hack
into c10/util/Exception.cpp, still calling the internal
glog_internal_namespace_::IsGoogleLoggingInitialized(); system glog 0.6.0
only exports the public ::google::IsGoogleLoggingInitialized() (mirror
the existing Logging.cpp glog patch).
- wrap-headers-destdir: cmake/PostBuildSteps.cmake runs wrap_headers.py via
install(CODE) against $/include with no $,
writing to the live /usr instead of the image; prepend $ENV.
- HIP_CLANG_PATH: 2.13.0's LoadHIP.cmake hardcodes the HIP compiler at
$/lib/llvm/bin, but Gentoo slots llvm; export it from
hipconfig -l in the rocm configure path.
Plus an fmt sed (drop the new FMT_NO_UNIQUE_ADDRESS compile-defs on the
unbundled fmt target) and a src_install cleanup (drop the aotriton libs
2.13.0 bundles into /usr/lib via hardcoded DESTINATION "lib" -- trips
multilib-strict and duplicates sci-libs/aotriton-bin).
commit bfb3f2583d4f86577dcc789426a0861b34b75439
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed Jun 17 21:48:26 2026 +0200
sci-ml/caffe2: Keyword 2.12.0-r91 for ~arm64
2.11.0-r90 already carried ~arm64 but 2.12.0-r91 had dropped it. The
C++ libtorch core builds CPU-only on aarch64 (default USE: no
cuda/rocm/fbgemm — fbgemm is x86-AVX-only), so restore ~arm64 and keep
the newest in sync. C++ (CMake); aarch64-portable.
commit 5653f839b39ea876b1a0e1ab42c0e7adc43b59e8
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Tue Jun 2 12:58:53 2026 +0200
sci-ml/caffe2: fix gcc-15 -Wtemplate-body in List_inl.h
gcc-15 enforces -Wtemplate-body in CUDA .cu compilation, rejecting
typename decltype(impl_->list)::difference_type as a dependent scope.
Replace with std::ptrdiff_t which is equivalent for std::vector iterators.
commit 7c384c8cac70105c06770e0ccb7af0fc295f2471
Author: Raukaan Cogbrother <cogbrother@raukaan.local>
Date: Sun May 31 03:29:35 2026 +0200
metadata: normalize tab → 2-space indent in 42 imported files
42 metadata.xml files inherited tab indentation from their origin
overlays (mostly the ROCm cluster, plus a handful of dev-python and
AMD-tooling imports). The overlay convention is 2-space; this brings
them in line with the other 498 files. Maintainer attributions
preserved verbatim.
commit 6b7192253e216334bb27f7c41b79803fc8ac8a36
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 09:37:30 2026 +0200
sci-ml/caffe2: relax pybind11 cap to <3.0.5
The <3.0.2 cap was inherited from the 2.11.0-era fork; ::gentoo's 2.12.0-r2 allows
<3.0.5 for the same upstream, and the stale cap blocked pybind11 from upgrading past
3.0.1.
commit b32aff8fbb037bea69577857ec2da162856a8e4f
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 29 01:26:23 2026 +0200
sci-ml/caffe2: fix xz-compressed patches breaking src_prepare
e8b75c54 compressed files/*.patch to clear a TotalSizeViolation, but eapply does not
decompress patches, so each one failed on the raw xz bytes. Decompress them into $
and repoint PATCHES in src_prepare, keeping files/ compressed so the cap stays
satisfied.
commit e8b75c54469f970239ba45e0cd17beb8159b8f9a
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 27 15:02:11 2026 +0200
sci-ml/caffe2: xz-compress files/*.patch to clear TotalSizeViolation
eapply decompresses .patch.xz transparently. Compression takes the
files/ payload from 58.2 KiB to 21.5 KiB (well under pkgcheck's 50
KiB threshold). Chosen over the extra-stuff bundle migration to keep
the patches in-tree; per-patch edits become 'xz -d' → edit → 'xz -9'
but the workflow stays single-repo.
commit e348703397f16e5622b4834a13833dc79ccc35c8
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Thu May 21 14:27:48 2026 +0200
sci-ml/caffe2: add 2.12.0-r90
Fork from ::gentoo 2.12.0-r1 (which drops rocm-fix-std-cpp17.patch to
resolve the C++20 .contains() failure under gcc-16). Carry forward our
overlay-only mkl-public-scrub patch:
- mkl-public-scrub: scrub MKL MPI/cluster libs and force GNU OpenMP
threading in caffe2::mkl's public link interface (ported from 2.11.0;
cmake/public/mkl.cmake context identical at 2.12.0).
Also: drop python3_11 from PYTHON_COMPAT and arm64 from KEYWORDS (not
tested in this overlay); drop nonexistent !sci-libs/QNNPACK blocker.
commit 253cd7b826e77eb7e68d1d3548d5ac9d36ab4ccc
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Wed May 13 14:35:42 2026 +0200
sci-ml/caffe2: disable py3.11
commit f1e7ab9cddb007e7de2e4e4e87ffedd361b212ad
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Sat May 9 00:44:08 2026 +0200
sci-ml/caffe2: drop dead !sci-libs/QNNPACK blocker
QNNPACK has never existed as a separate Gentoo package (verified
2026-05-09 — no entry in ::gentoo's sci-libs/, no diff-filter=A hits in
::gentoo history). Upstream PyTorch vendors it and exposes the build
toggle as USE_PYTORCH_QNNPACK, which we already wire via the qnnpack
USE flag. The blocker was protecting against a separate package that
never landed. pkgcheck NonexistentBlocker.
commit f19ffa2a53036faae2f64194c8b2f7c7c15e3fb0
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 8 19:29:35 2026 +0200
sci-ml/caffe2: drop two orphan patches inherited from ::gentoo fork
Both patches target older PVs (2.6.0 / 2.10.0) that 2.11.0-r90's
PATCHES= no longer references; the array uses $- variants for the
relevant fixes (mimalloc, rocm-fix-std-cpp17). Came along when forking
from ::gentoo's -r3 to add the MKL public-link scrub.
commit 78e2fba68899c5e58faba344002934108b58e32c
Author: Ivan S. Titov <iohann.s.titov@gmail.com>
Date: Fri May 8 01:25:42 2026 +0200
sci-ml/caffe2: fork ::gentoo's 2.11.0-r3 to -r90 with MKL public-link scrub
::gentoo's caffe2 ships caffe2/cmake/public/mkl.cmake unchanged from
upstream pytorch — calls find_package(MKL) and dumps the full result
into caffe2::mkl's INTERFACE_LINK_LIBRARIES with no filter. On hosts
with Intel oneAPI installed, the resolver is Intel's own MKLConfig.cmake
which by default returns the full HPC / Cluster Edition lib set,
including libmkl_scalapack_ilp64, libmkl_cdft_core,
libmkl_blacs_intelmpi_ilp64 (all MPI-distributed), and libmkl_intel_thread
(Intel-OpenMP threading layer). Those libs only exist when Intel's
separate Cluster + Compiler oneAPI packages are also installed; the
basic intel-oneapi-mkl package omits them.
Because they end up in caffe2::mkl's public link interface, every
downstream consumer that links against torch::torch (vllm's
cumem_allocator, custom torch C++ extensions, etc.) inherits them
and the link fails with "cannot find -lmkl_scalapack_ilp64" etc.
torch::torch's public APIs never reach BLACS / ScaLAPACK / distributed
FFT — distributed-tensor paths use NCCL / Gloo / MPI directly, not
MKL's BLACS. So filtering them is purely subtraction of incorrect
public-link-interface clutter. Forcing MKL_THREADING=gnu_thread before
the find pulls libmkl_gnu_thread (always available) instead of
libmkl_intel_thread, and pairs cleanly with system libgomp — also
avoids the multi-OpenMP-runtime mixing trap (libgomp + libiomp5 in
the same process can oversubscribe / deadlock).
Verified 2026-05-08 against caffe2-2.11.0 with USE="cuda distributed
fbgemm gloo memefficient mkl mpi nnpack numpy onednn openblas opencl
openmp": full rebuild + import torch succeed identically; downstream
consumers (vllm USE=cuda's cumem_allocator extension) link cleanly.
Upstream pytorch's cmake/public/mkl.cmake on main still has the same
unfiltered behaviour as of 2026-05-08; no upstream PR open for this
specific scrub. Drop this -r90 fork when an equivalent upstream fix
lands.

