Skip to content

Split gotoblas dispatch table by group / operation (2/2) - #6088

Open
topolarity wants to merge 7 commits into
OpenMathLib:developfrom
topolarity:ct/dispatch-split
Open

topolarity wants to merge 7 commits into
OpenMathLib:developfrom
topolarity:ct/dispatch-split

Conversation

@topolarity

Copy link
Copy Markdown

Dependent on #6087. This is PR 2 of 2 to split the gotoblas dispatch table by "group" (mostly by operation / kernel). The primary motivation is to make OpenBLAS amenable to pruning by the linker via --Wl,--gc-sections even in the presence of DYNAMIC_ARCH.

For applications linking statically against OpenBLAS, this can make a big difference to binary size:

Test program Upstream PR
dgemm only 19.5 MB 0.7 MB
dgesv + daxpy + dgemv + dnrm2 19.5 MB 1.6 MB
every cblas_d* routine 19.8 MB 3.5 MB

This is a very big, script-generated diff.

The changes are:

  1. The first commit changes setparam-ref.c to use designated initializers. This change is not required, but it makes the remaining change more mechanical and less dangerous. This is done by add_explicit_inits.py.
  2. The second commit changes gotoblas->group_suffix to OPENBLAS_DISPATCH(group)->group_suffix, so that all dispatches use the group-specific dispatch table. It also updates setparam-ref.c to initialize the group-specific dispatch tables and defines them conditionally in dispatch.c. This is done by split_dispatch_groups.py.

I have tried to make this reviewable, to the extent that I can.

This was generated with the assistance of Claude Opus 5.5 🤖, but the change was done entirely (except for the last 5-line commit) by applying the above scripts to the repo to make the changes in bulk. Hopefully this avoids being at the whim / vigilance of the AI, although care still needs to be taken to sanity check the result.

I have reviewed the core changes to the extent that I can, and I am still running some checks on other platforms.
This is only tested thoroughly on x86_64-linux at the moment. I will test soon on aarch64-darwin.

topolarity and others added 2 commits October 2, 2026 00:52
No functional change.  With DYNAMIC_ARCH the small-matrix kernel tables in
interface/gemm*.c hold FUNC_OFFSET()s, which are offsets into gotoblas_t,
and add them to gotoblas.  Spell that base FUNC_BASE(func), next to
FUNC_OFFSET(func), and reach it through GEMM_SMALL_KERNEL_BASE in the same
way as the offsets themselves.

This lets the kernels move out of gotoblas_t without interface/gemm*.c
knowing where they went.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for numbering the DYNAMIC_ARCH cores; nothing includes the
header yet.  dyn_cores.h holds one X-macro, OPENBLAS_CORE_LIST(X, arg),
which expands to X(CORE, arg) for each core of the build, in sorted order
and without duplicates.  Makefile.system and cmake/arch.cmake each write it
where they finish computing DYNAMIC_CORE.

dyn_cores.h is written by the top-level make only - sub-makes inherit
OPENBLAS_DYN_CORES - so that parallel sub-makes never rewrite it under a
compile.  It is emitted through $(HASH) because GNU Make 3.81 reads a
literal '#' inside $(shell ...) as a comment.  cmake rewrites it only when
the list changes, and puts it in the build directory, which the kernel
targets now have on their include path.  A stale copy left in the source
tree by make would shadow cmake's, so cmake refuses to configure with one
there, as it already does for config_kernel.h.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
topolarity and others added 5 commits October 2, 2026 01:35
Preparation for splitting the kernels out of gotoblas_t; nothing uses the
new macros yet.

  * common_param.h numbers the cores of the build, OPENBLAS_CORE_<CORE>,
    from OPENBLAS_CORE_LIST() in dyn_cores.h, and every core's gotoblas_t
    records its number in the new "core" field.

  * OPENBLAS_DISPATCH(group) is the running core's table for a group of
    kernels, from the array openblas_<group>_dispatch[], indexed by
    gotoblas->core.  driver/others/dispatch.c will define those arrays,
    with DEFINE_DISPATCH_TABLE(group).  OPENBLAS_DISPATCH_OFFSET/_BASE do
    for the offset tables of interface/gemm*.c what FUNC_OFFSET/FUNC_BASE
    do today.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preparation for splitting the kernels out of gotoblas_t, later in this
series.  After that split, each core's kernel tables are defined together in
setparam_<CORE>.o, and all the per-group arrays together in dispatch.o; a
static link would keep all of them as soon as it kept any, and with them
every kernel.  Separate sections let --gc-sections keep only the groups a
program calls.

On its own this changes no static link size: every kernel is still
reachable from gotoblas_t.

Only for compilers known to accept the flags, and not for MSVC-style
frontends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No functional change.  The gotoblas_t initializer was positional across
roughly 1300 entries and two dozen nested #if blocks, so which value lands
in which field could only be worked out by lining it up against the struct
definition by hand.  Label every entry with its field, one entry per line.

Generated by dispatch-split/add_explicit_inits.py, which pairs the entries
with the fields of gotoblas_t for every combination of the macros that guard
those fields (BUILD_*, SMALL_MATRIX_OPT, EXPRECISION, ARCH_*: 1024
configurations) and stops if any entry would get a different field in any
of them.

setparam_<CORE>.o is byte-identical before and after for all 14 x86_64
DYNAMIC_ARCH cores, both as built and with BFLOAT16/HFLOAT16 enabled,
EXPRECISION and SMALL_MATRIX_OPT disabled, or only one of BUILD_SINGLE,
BUILD_DOUBLE, BUILD_COMPLEX and BUILD_COMPLEX16 enabled.

Generated-by: dispatch-split/add_explicit_inits.py
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With DYNAMIC_ARCH, gotoblas_t holds every kernel of a core, and the cpu
detection code in driver/others/dynamic*.c takes the address of every
core's gotoblas_t.  One reference to a core's table thus keeps all of that
core's kernels, and a static link keeps every kernel of every core, of which
--gc-sections can remove almost nothing.

Move the 904 kernel pointers into 233 small tables, one per kernel group
(the kernels whose names share a leading token: dgemm_kernel, dgemm_beta,
dgemm_incopy, ... are "dgemm"), each instantiated per core in setparam-ref.c
and reached through OPENBLAS_DISPATCH(group), from an array of them that
driver/others/dispatch.c defines.  gotoblas_t keeps only the tuning
parameters and init(), so cpu detection no longer references any kernel,
and a program links only the groups it calls:

  text of a static executable       before     after
  dgemm only                        19.5 MB    0.7 MB
  dgesv + daxpy + dgemv + dnrm2     19.5 MB    1.6 MB
  every cblas_d* routine            19.8 MB    3.5 MB

(x86_64, 14 cores, linked with --gc-sections; dispatch-split/measure-size.sh.)

This commit is the output of dispatch-split/split_dispatch_groups.py and
nothing else; see there for the exact rules.  "gotoblas -> dgemm_kernel"
becomes "OPENBLAS_DISPATCH(dgemm) -> dgemm_kernel", that is
openblas_dgemm_dispatch[gotoblas->core]->dgemm_kernel.  That costs nothing
on real work (single-threaded dgemm, n=2048: 57.1 GFLOP/s before, 57.2
after) but shows on the smallest calls (ddot, n=8: 5.05 ns per call before,
5.6-5.8 ns after).

For all 14 x86_64 cores, every kernel pointer (resolved to its symbol) and
every tuning parameter (after init()) is the same as before, 11830 entries
in all (dispatch-split/compare-tables.py).

Generated-by: dispatch-split/split_dispatch_groups.py
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing uses them since the kernels moved out of gotoblas_t: the
small-matrix kernel tables use OPENBLAS_DISPATCH_OFFSET/_BASE instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant