Split gotoblas dispatch table by group / operation (2/2) - #6088
Open
topolarity wants to merge 7 commits into
Open
topolarity wants to merge 7 commits into
topolarity wants to merge 7 commits into
Conversation
No functional change. With DYNAMIC_ARCH the small-matrix kernel tables in interface/gemm*.c hold FUNC_OFFSET()s, which are offsets into gotoblas_t, and add them to gotoblas. Spell that base FUNC_BASE(func), next to FUNC_OFFSET(func), and reach it through GEMM_SMALL_KERNEL_BASE in the same way as the offsets themselves. This lets the kernels move out of gotoblas_t without interface/gemm*.c knowing where they went. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for numbering the DYNAMIC_ARCH cores; nothing includes the header yet. dyn_cores.h holds one X-macro, OPENBLAS_CORE_LIST(X, arg), which expands to X(CORE, arg) for each core of the build, in sorted order and without duplicates. Makefile.system and cmake/arch.cmake each write it where they finish computing DYNAMIC_CORE. dyn_cores.h is written by the top-level make only - sub-makes inherit OPENBLAS_DYN_CORES - so that parallel sub-makes never rewrite it under a compile. It is emitted through $(HASH) because GNU Make 3.81 reads a literal '#' inside $(shell ...) as a comment. cmake rewrites it only when the list changes, and puts it in the build directory, which the kernel targets now have on their include path. A stale copy left in the source tree by make would shadow cmake's, so cmake refuses to configure with one there, as it already does for config_kernel.h. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
topolarity
force-pushed
the
ct/dispatch-split
branch
from
October 2, 2026 00:55
1508a72 to
4d4170c
Compare
Preparation for splitting the kernels out of gotoblas_t; nothing uses the
new macros yet.
* common_param.h numbers the cores of the build, OPENBLAS_CORE_<CORE>,
from OPENBLAS_CORE_LIST() in dyn_cores.h, and every core's gotoblas_t
records its number in the new "core" field.
* OPENBLAS_DISPATCH(group) is the running core's table for a group of
kernels, from the array openblas_<group>_dispatch[], indexed by
gotoblas->core. driver/others/dispatch.c will define those arrays,
with DEFINE_DISPATCH_TABLE(group). OPENBLAS_DISPATCH_OFFSET/_BASE do
for the offset tables of interface/gemm*.c what FUNC_OFFSET/FUNC_BASE
do today.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for splitting the kernels out of gotoblas_t, later in this series. After that split, each core's kernel tables are defined together in setparam_<CORE>.o, and all the per-group arrays together in dispatch.o; a static link would keep all of them as soon as it kept any, and with them every kernel. Separate sections let --gc-sections keep only the groups a program calls. On its own this changes no static link size: every kernel is still reachable from gotoblas_t. Only for compilers known to accept the flags, and not for MSVC-style frontends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No functional change. The gotoblas_t initializer was positional across roughly 1300 entries and two dozen nested #if blocks, so which value lands in which field could only be worked out by lining it up against the struct definition by hand. Label every entry with its field, one entry per line. Generated by dispatch-split/add_explicit_inits.py, which pairs the entries with the fields of gotoblas_t for every combination of the macros that guard those fields (BUILD_*, SMALL_MATRIX_OPT, EXPRECISION, ARCH_*: 1024 configurations) and stops if any entry would get a different field in any of them. setparam_<CORE>.o is byte-identical before and after for all 14 x86_64 DYNAMIC_ARCH cores, both as built and with BFLOAT16/HFLOAT16 enabled, EXPRECISION and SMALL_MATRIX_OPT disabled, or only one of BUILD_SINGLE, BUILD_DOUBLE, BUILD_COMPLEX and BUILD_COMPLEX16 enabled. Generated-by: dispatch-split/add_explicit_inits.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With DYNAMIC_ARCH, gotoblas_t holds every kernel of a core, and the cpu detection code in driver/others/dynamic*.c takes the address of every core's gotoblas_t. One reference to a core's table thus keeps all of that core's kernels, and a static link keeps every kernel of every core, of which --gc-sections can remove almost nothing. Move the 904 kernel pointers into 233 small tables, one per kernel group (the kernels whose names share a leading token: dgemm_kernel, dgemm_beta, dgemm_incopy, ... are "dgemm"), each instantiated per core in setparam-ref.c and reached through OPENBLAS_DISPATCH(group), from an array of them that driver/others/dispatch.c defines. gotoblas_t keeps only the tuning parameters and init(), so cpu detection no longer references any kernel, and a program links only the groups it calls: text of a static executable before after dgemm only 19.5 MB 0.7 MB dgesv + daxpy + dgemv + dnrm2 19.5 MB 1.6 MB every cblas_d* routine 19.8 MB 3.5 MB (x86_64, 14 cores, linked with --gc-sections; dispatch-split/measure-size.sh.) This commit is the output of dispatch-split/split_dispatch_groups.py and nothing else; see there for the exact rules. "gotoblas -> dgemm_kernel" becomes "OPENBLAS_DISPATCH(dgemm) -> dgemm_kernel", that is openblas_dgemm_dispatch[gotoblas->core]->dgemm_kernel. That costs nothing on real work (single-threaded dgemm, n=2048: 57.1 GFLOP/s before, 57.2 after) but shows on the smallest calls (ddot, n=8: 5.05 ns per call before, 5.6-5.8 ns after). For all 14 x86_64 cores, every kernel pointer (resolved to its symbol) and every tuning parameter (after init()) is the same as before, 11830 entries in all (dispatch-split/compare-tables.py). Generated-by: dispatch-split/split_dispatch_groups.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing uses them since the kernels moved out of gotoblas_t: the small-matrix kernel tables use OPENBLAS_DISPATCH_OFFSET/_BASE instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
topolarity
force-pushed
the
ct/dispatch-split
branch
from
October 2, 2026 01:35
4d4170c to
37da176
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Dependent on #6087. This is PR 2 of 2 to split the
gotoblasdispatch table by "group" (mostly by operation / kernel). The primary motivation is to make OpenBLAS amenable to pruning by the linker via--Wl,--gc-sectionseven in the presence ofDYNAMIC_ARCH.For applications linking statically against OpenBLAS, this can make a big difference to binary size:
dgemmonlydgesv+daxpy+dgemv+dnrm2cblas_d*routineThis is a very big, script-generated diff.
The changes are:
setparam-ref.cto use designated initializers. This change is not required, but it makes the remaining change more mechanical and less dangerous. This is done by add_explicit_inits.py.gotoblas->group_suffixtoOPENBLAS_DISPATCH(group)->group_suffix, so that all dispatches use the group-specific dispatch table. It also updatessetparam-ref.cto initialize the group-specific dispatch tables and defines them conditionally indispatch.c. This is done by split_dispatch_groups.py.I have tried to make this reviewable, to the extent that I can.
This was generated with the assistance of Claude Opus 5.5 🤖, but the change was done entirely (except for the last 5-line commit) by applying the above scripts to the repo to make the changes in bulk. Hopefully this avoids being at the whim / vigilance of the AI, although care still needs to be taken to sanity check the result.
I have reviewed the core changes to the extent that I can, and I am still running some checks on other platforms.
This is only tested thoroughly on
x86_64-linuxat the moment. I will test soon onaarch64-darwin.