Commit de7d561
fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer (#2917)
* fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer
Each physical allocation, address reservation and mapping now lives in
a std::shared_ptr handle whose deleter knows the exact driver call to
undo it. A buffer owns a range of mappings through its device pointer
handle, so everything a buffer maps is released when the last buffer
that maps it closes, and a failed multi-step operation unwinds by
letting its local handles die. The module moves from Python to Cython.
The design is in cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md.
Behavior changes:
- modify_allocation returns a new VirtualMemoryBuffer and leaves the
input open; the two alias the same physical memory, which is freed
when the last of them closes. The pointer is preserved when the
driver grants the adjacent address range.
- Buffer.size after a grow is the aligned total.
- config= applies to the chunk the call adds and is not stored on the
resource.
- Buffers from allocate() free themselves on close and do not call
deallocate(), which now serves pointers wrapped with
Buffer.from_handle.
- A buffer records the stream passed to allocate(); the last close of
an aliased range synchronizes every recorded stream before it unmaps.
An explicit close on a capturing stream raises.
- location_type="host" requires handle_type=None. allocate(0) returns
an empty buffer without a driver call.
Fixes #2887
Fixes #2907
Fixes #2908
Fixes #2909
Fixes #2886
Fixes #2345
Addresses #2388 item 2 and the size-0, misaligned-probe and host
handle-type parts of #2910. Part of #2906.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(cuda.core): run the VMM shutdown test from an empty directory
The child interpreter inherited pytest's working directory, cuda_core/,
so `import cuda.core` resolved to the uncompiled source tree in CI and
failed on `cuda.core._version`. Use the shared run_python_snippet
helper, which starts the child in an empty temporary directory.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(cuda.core): harden VirtualMemoryResource after review
- Run the range deleter's stream sync in relaxed capture mode, so a
capture on an unrelated stream is not invalidated.
- Record a real stream on allocate(0) and inherit it on the grow.
- Require handle_type=None for location "host" only.
- Apply the constructor's option checks to a per-call
modify_allocation config, including the RDMA support check.
- Narrow the close() capture contract to non-default streams in the
docstring, design doc and release note.
- Tests: failed grow leaves the input intact, close during an unrelated
capture, GC release during capture, deterministic stream sync with a
sleep kernel, cuMemGetAccess on both chunks, graph retention across a
grow, forced-move leak on 2 MiB that fails rather than skips.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(cuda.core): pass the handle-type enum to the CUDA 13.0 bindings
The struct setter in cuda-bindings 13.0 accepts only the enum, and the
Cython helper returns a plain int.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(cuda.core): cover host VMM grow, default-stream capture skip, and size overflow
Add tests for a host-located grow that moves, for a release ordered on the
legacy default stream while a blocking stream in its context is capturing,
and for a size whose rounding to the granularity does not fit in size_t. The
forced-move test now asserts that neither the grow nor the closes warn
(#2877).
_align_up raises OverflowError instead of wrapping. The docstrings and the
release note say that config has no effect when the buffer already covers
the request and never changes the access of mapped memory, and that a
host-located resource records no default stream. VMM_DESIGN.md describes the
close() override.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(cuda.core): state the VirtualMemoryResource invariants in VMM_DESIGN.md
List the fifteen properties the handle-based VirtualMemoryResource
maintains: once-only and ordered release of reservations and allocations,
what a failed or successful grow leaves behind, stream ordering and graph
capture, ownership by graph nodes and aliases, per-chunk access, range
layout and rounding, the base-address registry, the deallocate() contract,
context independence, and interpreter shutdown. The wording names no
mechanism, so the list stays valid if the release is made stream-ordered.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* refactor(cuda.core): make VMM resource attributes read-only and share the default-token check
`VirtualMemoryResource.device` and `.config` are now `cdef readonly`; no
other resource in the layer exposes writable attributes. The raw-handle
default-stream check moves to the stream module as
`Stream_handle_is_default_token`, and `Stream_is_default_token` delegates
to it, so the two modules agree on one definition. `VirtualMemoryBuffer.
close()` treats an empty handle as "no stream recorded" explicitly instead
of folding it into the token check.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(cuda.core): check VMM release per address instead of device-wide free memory
The counter from cuMemGetInfo covers the whole device, so any other
process moves it and the tests fail on shared machines. Each test now
asks the driver about the exact addresses it used after close: the
mapping lookup must fail and freeing the reservation must fail because it
no longer exists. Closes run under assert_no_cuda_warning, so a failed
unmap, address free or release fails the test. The leak test lists one
reservation per allocate and one more per grow, for both the in-place and
the moved case. One deterministic pass replaces the eight-iteration loop.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(cuda.core): make VMM ranges immutable and give each buffer its own
A range is now an immutable list of mapping handles that lives in the
buffer's device pointer box. A grow copies the input's list, appends or
replaces mappings, and builds a new range for its result; the input's
range never changes. Mapping handles are shared between ranges, so each
mapping unmaps when the last range that holds it goes.
Two buffers therefore never share mutable state, which is what made
concurrent grows of aliased buffers unsafe: the shared range's mapping
vector was appended and iterated without a lock, and a grow that lost a
race could dereference a handle another thread had emptied. With one
range and one recorded stream per buffer, teardown follows the ordinary
Buffer model, so the stream union, the range mutex, the base-address
registry and the range header are gone.
modify_allocation never returns its input any more. A request the buffer
already covers returns a full alias without a driver call, so closing
the result never closes the buffer passed in. It also reads the input's
handle once, so a close from another thread defers the release instead
of emptying what the call reads.
The box behind a VMM handle is a VmmDevicePtrBox, a DevicePtrBox with
the range as a member and no virtual functions. Every handle on a
VirtualMemoryBuffer comes from deviceptr_create_vmm, including the
size-zero buffer, which now sits on a VMM box with an empty range, so
the class check in modify_allocation is what makes the downcast valid.
Adds a test that grows and closes aliases of one buffer from four
threads, reduced from the report on this pull request.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(cuda.core): return an empty stream for an empty device pointer handle
deallocation_stream() was the one accessor in the layer that
dereferenced an empty handle instead of returning empty, as its sibling
set_deallocation_stream and every as_cu() overload do. A buffer closed
by one thread while another still reads its handle now gets an empty
stream rather than a crash.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(cuda.core): bound the joins in the concurrent VMM grow test
Join each worker with the suite's sanitizer-aware timeout and assert
that none is still alive before the shared buffers are closed, as the
other threading tests do. Drop "undefined" from the docstring.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(cuda.core): keep the pending exception across a MemoryResource deleter
The deleter behind Buffer.from_handle(..., mr) acquired the GIL and
called deallocate() while an exception could be propagating through
the caller that released the last reference. The handler that reports
a failed deallocate() then cleared that exception, and the caller
returned an error with no exception set, which Python reports as
SystemError. report_message already saved and restored the pending
exception inline.
PendingExceptionGuard (py.hpp) saves the exception in flight and
restores it on scope exit, dropping anything the scope itself raised.
The deleter and report_message use it. The regression test releases a
temporary buffer whose deallocate() fails while a TypeError propagates
and expects the TypeError and a CUDAWarning.
Found in review of #2917.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* docs(cuda.core): restore __weakref__ on VirtualMemoryResource and state the release wait
The Python class had __weakref__ as a subclass of an extension type;
the cdef class declares it, as Buffer and the pool-backed resources do.
VMM_DESIGN.md compared the blocking release to pool-backed deleters,
which do not block: cuMemFreeAsync is stream-ordered. The precedent is
_SynchronousMemoryResource and LegacyPinnedMemoryResource, which wait
in deallocate() on the same deleter path. The design doc, the class
docstring, close() and the release note now say that closing a buffer
waits for the work on its deallocation stream and how to control when
that happens. The release note also drops a sentence about the range
deleter synchronizing every recorded stream, which immutable ranges
made false. The follow-up for a stream-ordered release is #2989.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(cuda.core): query capture state with cuStreamIsCapturing in the VMM deleter
cuStreamGetCaptureInfo has two slots in the cuda-bindings loader (v2 and
v3) and a different arity per CUDA major, which needed a build-major
fence and broke the driver-table test that builds a fake table from the
first slot per name. cuStreamIsCapturing is the query that
cuStreamGetCaptureInfo makes first: same status, including the implicit
capture error for the legacy stream, one slot, one signature.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>1 parent 54d9b3e commit de7d561
27 files changed
Lines changed: 2822 additions & 945 deletions
File tree
- cuda_core
- cuda/core
- _cpp/rt
- _memory
- _utils
- docs/source
- release
- tests
- graph
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
63 | 63 | | |
64 | 64 | | |
65 | 65 | | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
66 | 71 | | |
67 | 72 | | |
68 | 73 | | |
| |||
Large diffs are not rendered by default.
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
5 | 5 | | |
6 | 6 | | |
7 | 7 | | |
| 8 | + | |
8 | 9 | | |
9 | 10 | | |
10 | 11 | | |
| |||
232 | 233 | | |
233 | 234 | | |
234 | 235 | | |
| 236 | + | |
235 | 237 | | |
236 | 238 | | |
237 | 239 | | |
| |||
240 | 242 | | |
241 | 243 | | |
242 | 244 | | |
| 245 | + | |
| 246 | + | |
| 247 | + | |
| 248 | + | |
| 249 | + | |
| 250 | + | |
| 251 | + | |
| 252 | + | |
| 253 | + | |
| 254 | + | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
| 259 | + | |
| 260 | + | |
| 261 | + | |
| 262 | + | |
| 263 | + | |
| 264 | + | |
| 265 | + | |
| 266 | + | |
| 267 | + | |
| 268 | + | |
| 269 | + | |
| 270 | + | |
| 271 | + | |
| 272 | + | |
| 273 | + | |
| 274 | + | |
| 275 | + | |
| 276 | + | |
| 277 | + | |
| 278 | + | |
| 279 | + | |
| 280 | + | |
| 281 | + | |
| 282 | + | |
| 283 | + | |
| 284 | + | |
| 285 | + | |
| 286 | + | |
| 287 | + | |
| 288 | + | |
| 289 | + | |
| 290 | + | |
| 291 | + | |
| 292 | + | |
| 293 | + | |
| 294 | + | |
| 295 | + | |
| 296 | + | |
| 297 | + | |
| 298 | + | |
| 299 | + | |
| 300 | + | |
| 301 | + | |
| 302 | + | |
| 303 | + | |
| 304 | + | |
| 305 | + | |
| 306 | + | |
| 307 | + | |
| 308 | + | |
243 | 309 | | |
244 | 310 | | |
245 | 311 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
126 | 126 | | |
127 | 127 | | |
128 | 128 | | |
129 | | - | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
| 132 | + | |
| 133 | + | |
| 134 | + | |
| 135 | + | |
| 136 | + | |
| 137 | + | |
| 138 | + | |
| 139 | + | |
| 140 | + | |
| 141 | + | |
130 | 142 | | |
131 | 143 | | |
132 | 144 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
62 | 62 | | |
63 | 63 | | |
64 | 64 | | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
65 | 71 | | |
66 | 72 | | |
67 | 73 | | |
| |||
126 | 132 | | |
127 | 133 | | |
128 | 134 | | |
| 135 | + | |
| 136 | + | |
| 137 | + | |
129 | 138 | | |
130 | 139 | | |
131 | 140 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
123 | 123 | | |
124 | 124 | | |
125 | 125 | | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
| 132 | + | |
| 133 | + | |
| 134 | + | |
126 | 135 | | |
127 | 136 | | |
128 | 137 | | |
| |||
137 | 146 | | |
138 | 147 | | |
139 | 148 | | |
140 | | - | |
| 149 | + | |
| 150 | + | |
141 | 151 | | |
142 | | - | |
| 152 | + | |
143 | 153 | | |
144 | 154 | | |
145 | 155 | | |
| |||
300 | 310 | | |
301 | 311 | | |
302 | 312 | | |
| 313 | + | |
| 314 | + | |
| 315 | + | |
| 316 | + | |
| 317 | + | |
| 318 | + | |
| 319 | + | |
| 320 | + | |
| 321 | + | |
| 322 | + | |
| 323 | + | |
| 324 | + | |
| 325 | + | |
| 326 | + | |
| 327 | + | |
| 328 | + | |
| 329 | + | |
| 330 | + | |
| 331 | + | |
| 332 | + | |
| 333 | + | |
| 334 | + | |
| 335 | + | |
| 336 | + | |
| 337 | + | |
| 338 | + | |
| 339 | + | |
| 340 | + | |
| 341 | + | |
| 342 | + | |
303 | 343 | | |
304 | 344 | | |
305 | 345 | | |
| |||
325 | 365 | | |
326 | 366 | | |
327 | 367 | | |
| 368 | + | |
| 369 | + | |
| 370 | + | |
| 371 | + | |
328 | 372 | | |
329 | 373 | | |
330 | 374 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
94 | 94 | | |
95 | 95 | | |
96 | 96 | | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
| 111 | + | |
| 112 | + | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
| 119 | + | |
| 120 | + | |
| 121 | + | |
| 122 | + | |
| 123 | + | |
| 124 | + | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
| 132 | + | |
| 133 | + | |
| 134 | + | |
| 135 | + | |
97 | 136 | | |
98 | 137 | | |
99 | 138 | | |
| |||
205 | 244 | | |
206 | 245 | | |
207 | 246 | | |
| 247 | + | |
| 248 | + | |
| 249 | + | |
| 250 | + | |
| 251 | + | |
| 252 | + | |
| 253 | + | |
| 254 | + | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
208 | 259 | | |
209 | 260 | | |
210 | 261 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
48 | 48 | | |
49 | 49 | | |
50 | 50 | | |
51 | | - | |
52 | | - | |
53 | | - | |
54 | | - | |
55 | | - | |
56 | | - | |
57 | | - | |
58 | | - | |
59 | | - | |
60 | | - | |
61 | | - | |
62 | | - | |
63 | | - | |
64 | | - | |
65 | | - | |
66 | | - | |
67 | | - | |
68 | | - | |
69 | | - | |
70 | | - | |
71 | | - | |
72 | | - | |
73 | | - | |
74 | | - | |
75 | | - | |
76 | | - | |
77 | | - | |
78 | | - | |
79 | | - | |
80 | | - | |
81 | | - | |
82 | | - | |
83 | | - | |
84 | 51 | | |
85 | 52 | | |
86 | 53 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
34 | 34 | | |
35 | 35 | | |
36 | 36 | | |
37 | | - | |
38 | | - | |
39 | | - | |
40 | | - | |
41 | | - | |
42 | | - | |
| 37 | + | |
43 | 38 | | |
44 | 39 | | |
45 | 40 | | |
| |||
52 | 47 | | |
53 | 48 | | |
54 | 49 | | |
55 | | - | |
56 | | - | |
57 | | - | |
58 | | - | |
59 | | - | |
60 | 50 | | |
61 | 51 | | |
62 | 52 | | |
| |||
0 commit comments