Skip to content

Commit de7d561

Browse files
Andy-Jostclaude
andauthored
fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer (#2917)
* fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer Each physical allocation, address reservation and mapping now lives in a std::shared_ptr handle whose deleter knows the exact driver call to undo it. A buffer owns a range of mappings through its device pointer handle, so everything a buffer maps is released when the last buffer that maps it closes, and a failed multi-step operation unwinds by letting its local handles die. The module moves from Python to Cython. The design is in cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md. Behavior changes: - modify_allocation returns a new VirtualMemoryBuffer and leaves the input open; the two alias the same physical memory, which is freed when the last of them closes. The pointer is preserved when the driver grants the adjacent address range. - Buffer.size after a grow is the aligned total. - config= applies to the chunk the call adds and is not stored on the resource. - Buffers from allocate() free themselves on close and do not call deallocate(), which now serves pointers wrapped with Buffer.from_handle. - A buffer records the stream passed to allocate(); the last close of an aliased range synchronizes every recorded stream before it unmaps. An explicit close on a capturing stream raises. - location_type="host" requires handle_type=None. allocate(0) returns an empty buffer without a driver call. Fixes #2887 Fixes #2907 Fixes #2908 Fixes #2909 Fixes #2886 Fixes #2345 Addresses #2388 item 2 and the size-0, misaligned-probe and host handle-type parts of #2910. Part of #2906. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): run the VMM shutdown test from an empty directory The child interpreter inherited pytest's working directory, cuda_core/, so `import cuda.core` resolved to the uncompiled source tree in CI and failed on `cuda.core._version`. Use the shared run_python_snippet helper, which starts the child in an empty temporary directory. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cuda.core): harden VirtualMemoryResource after review - Run the range deleter's stream sync in relaxed capture mode, so a capture on an unrelated stream is not invalidated. - Record a real stream on allocate(0) and inherit it on the grow. - Require handle_type=None for location "host" only. - Apply the constructor's option checks to a per-call modify_allocation config, including the RDMA support check. - Narrow the close() capture contract to non-default streams in the docstring, design doc and release note. - Tests: failed grow leaves the input intact, close during an unrelated capture, GC release during capture, deterministic stream sync with a sleep kernel, cuMemGetAccess on both chunks, graph retention across a grow, forced-move leak on 2 MiB that fails rather than skips. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): pass the handle-type enum to the CUDA 13.0 bindings The struct setter in cuda-bindings 13.0 accepts only the enum, and the Cython helper returns a plain int. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): cover host VMM grow, default-stream capture skip, and size overflow Add tests for a host-located grow that moves, for a release ordered on the legacy default stream while a blocking stream in its context is capturing, and for a size whose rounding to the granularity does not fit in size_t. The forced-move test now asserts that neither the grow nor the closes warn (#2877). _align_up raises OverflowError instead of wrapping. The docstrings and the release note say that config has no effect when the buffer already covers the request and never changes the access of mapped memory, and that a host-located resource records no default stream. VMM_DESIGN.md describes the close() override. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(cuda.core): state the VirtualMemoryResource invariants in VMM_DESIGN.md List the fifteen properties the handle-based VirtualMemoryResource maintains: once-only and ordered release of reservations and allocations, what a failed or successful grow leaves behind, stream ordering and graph capture, ownership by graph nodes and aliases, per-chunk access, range layout and rounding, the base-address registry, the deallocate() contract, context independence, and interpreter shutdown. The wording names no mechanism, so the list stays valid if the release is made stream-ordered. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * refactor(cuda.core): make VMM resource attributes read-only and share the default-token check `VirtualMemoryResource.device` and `.config` are now `cdef readonly`; no other resource in the layer exposes writable attributes. The raw-handle default-stream check moves to the stream module as `Stream_handle_is_default_token`, and `Stream_is_default_token` delegates to it, so the two modules agree on one definition. `VirtualMemoryBuffer. close()` treats an empty handle as "no stream recorded" explicitly instead of folding it into the token check. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): check VMM release per address instead of device-wide free memory The counter from cuMemGetInfo covers the whole device, so any other process moves it and the tests fail on shared machines. Each test now asks the driver about the exact addresses it used after close: the mapping lookup must fail and freeing the reservation must fail because it no longer exists. Closes run under assert_no_cuda_warning, so a failed unmap, address free or release fails the test. The leak test lists one reservation per allocate and one more per grow, for both the in-place and the moved case. One deterministic pass replaces the eight-iteration loop. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cuda.core): make VMM ranges immutable and give each buffer its own A range is now an immutable list of mapping handles that lives in the buffer's device pointer box. A grow copies the input's list, appends or replaces mappings, and builds a new range for its result; the input's range never changes. Mapping handles are shared between ranges, so each mapping unmaps when the last range that holds it goes. Two buffers therefore never share mutable state, which is what made concurrent grows of aliased buffers unsafe: the shared range's mapping vector was appended and iterated without a lock, and a grow that lost a race could dereference a handle another thread had emptied. With one range and one recorded stream per buffer, teardown follows the ordinary Buffer model, so the stream union, the range mutex, the base-address registry and the range header are gone. modify_allocation never returns its input any more. A request the buffer already covers returns a full alias without a driver call, so closing the result never closes the buffer passed in. It also reads the input's handle once, so a close from another thread defers the release instead of emptying what the call reads. The box behind a VMM handle is a VmmDevicePtrBox, a DevicePtrBox with the range as a member and no virtual functions. Every handle on a VirtualMemoryBuffer comes from deviceptr_create_vmm, including the size-zero buffer, which now sits on a VMM box with an empty range, so the class check in modify_allocation is what makes the downcast valid. Adds a test that grows and closes aliases of one buffer from four threads, reduced from the report on this pull request. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cuda.core): return an empty stream for an empty device pointer handle deallocation_stream() was the one accessor in the layer that dereferenced an empty handle instead of returning empty, as its sibling set_deallocation_stream and every as_cu() overload do. A buffer closed by one thread while another still reads its handle now gets an empty stream rather than a crash. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(cuda.core): bound the joins in the concurrent VMM grow test Join each worker with the suite's sanitizer-aware timeout and assert that none is still alive before the shared buffers are closed, as the other threading tests do. Drop "undefined" from the docstring. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cuda.core): keep the pending exception across a MemoryResource deleter The deleter behind Buffer.from_handle(..., mr) acquired the GIL and called deallocate() while an exception could be propagating through the caller that released the last reference. The handler that reports a failed deallocate() then cleared that exception, and the caller returned an error with no exception set, which Python reports as SystemError. report_message already saved and restored the pending exception inline. PendingExceptionGuard (py.hpp) saves the exception in flight and restores it on scope exit, dropping anything the scope itself raised. The deleter and report_message use it. The regression test releases a temporary buffer whose deallocate() fails while a TypeError propagates and expects the TypeError and a CUDAWarning. Found in review of #2917. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(cuda.core): restore __weakref__ on VirtualMemoryResource and state the release wait The Python class had __weakref__ as a subclass of an extension type; the cdef class declares it, as Buffer and the pool-backed resources do. VMM_DESIGN.md compared the blocking release to pool-backed deleters, which do not block: cuMemFreeAsync is stream-ordered. The precedent is _SynchronousMemoryResource and LegacyPinnedMemoryResource, which wait in deallocate() on the same deleter path. The design doc, the class docstring, close() and the release note now say that closing a buffer waits for the work on its deallocation stream and how to control when that happens. The release note also drops a sentence about the range deleter synchronizing every recorded stream, which immutable ranges made false. The follow-up for a stream-ordered release is #2989. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(cuda.core): query capture state with cuStreamIsCapturing in the VMM deleter cuStreamGetCaptureInfo has two slots in the cuda-bindings loader (v2 and v3) and a different arity per CUDA major, which needed a build-major fence and broke the driver-table test that builds a fake table from the first slot per name. cuStreamIsCapturing is the query that cuStreamGetCaptureInfo makes first: same status, including the implicit capture error for the legacy stream, one slot, one signature. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
1 parent 54d9b3e commit de7d561

27 files changed

Lines changed: 2822 additions & 945 deletions

‎cuda_core/cuda/core/_cpp/rt/DESIGN.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -63,6 +63,11 @@ Internally, handles use **shared pointer aliasing**: the actual managed object i
6363
"box" containing the resource, its dependencies, and any state needed for destruction.
6464
The public handle points only to the raw resource field, keeping the API minimal.
6565

66+
The virtual memory resource adds `MemAllocationHandle`, `VaReservationHandle` and
67+
`VaMappingHandle`. Their values are `TaggedHandle<T, N>` wrappers, because
68+
`CUmemGenericAllocationHandle` and `CUdeviceptr` are both `unsigned long long` and the
69+
accessor overloads must stay distinct. See [VMM_DESIGN.md](VMM_DESIGN.md).
70+
6671
### Why shared_ptr?
6772

6873
- **Automatic reference counting**: Resources are released when the last reference

‎cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md‎

Lines changed: 269 additions & 0 deletions
Large diffs are not rendered by default.

‎cuda_core/cuda/core/_cpp/rt/api.hpp‎

Lines changed: 66 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,6 +5,7 @@
55
#pragma once
66

77
#include "types.hpp"
8+
#include <vector>
89
#include <cuda.h>
910
#include <nvrtc.h>
1011
#include <cstddef>
@@ -232,6 +233,7 @@ DevicePtrHandle deviceptr_import_ipc(
232233

233234
// Access the deallocation stream for a device pointer handle (read-only).
234235
// For non-owning handles, the stream is not used but can still be accessed.
236+
// Returns an empty handle for an empty device pointer handle.
235237
StreamHandle deallocation_stream(const DevicePtrHandle& h) noexcept;
236238

237239
// Set the deallocation stream for a device pointer handle.
@@ -240,6 +242,70 @@ StreamHandle deallocation_stream(const DevicePtrHandle& h) noexcept;
240242
CUresult set_deallocation_stream(
241243
const DevicePtrHandle& h, const StreamHandle& h_stream) noexcept;
242244

245+
// ============================================================================
246+
// Virtual memory management (VMM_DESIGN.md)
247+
//
248+
// A VirtualMemoryResource buffer is a range of mappings. Each mapping holds
249+
// one physical allocation and one address reservation; the mapping deleter
250+
// unmaps, then the allocation is released and the reservation freed as their
251+
// last references go. A buffer's DevicePtrHandle owns the range. Ranges are
252+
// immutable: a grow builds a new range for its result.
253+
// ============================================================================
254+
255+
// Create a physical allocation via cuMemCreate. The access descriptors are
256+
// applied to every mapping of this allocation. When the last reference is
257+
// released, cuMemRelease is called; the memory is freed once no mapping
258+
// remains. Returns empty handle on error (caller must check).
259+
MemAllocationHandle create_mem_allocation_handle(size_t size, const CUmemAllocationProp& prop,
260+
const CUmemAccessDesc* descs, size_t count);
261+
262+
// Size of the allocation; the only size cuMemMap accepts for it.
263+
size_t mem_allocation_size(const MemAllocationHandle& h) noexcept;
264+
265+
// Reserve an address range via cuMemAddressReserve. Pass alignment 0 for the
266+
// driver default. When the last reference is released, cuMemAddressFree is
267+
// called with the exact reserved pair. Returns empty handle on error.
268+
VaReservationHandle create_va_reservation_handle(size_t size, size_t alignment, CUdeviceptr hint);
269+
270+
// Size of the reservation.
271+
size_t va_reservation_size(const VaReservationHandle& h) noexcept;
272+
273+
// Map the whole allocation at ptr inside the reservation via cuMemMap and
274+
// apply the allocation's access descriptors. The mapping structurally depends
275+
// on both handles. When the last reference is released, cuMemUnmap is called
276+
// first. Returns empty handle on error, including a range outside the
277+
// reservation; a failed cuMemSetAccess unmaps before returning.
278+
VaMappingHandle create_va_mapping_handle(CUdeviceptr ptr, const MemAllocationHandle& h_alloc,
279+
const VaReservationHandle& h_res);
280+
281+
// Mapping accessors.
282+
size_t va_mapping_size(const VaMappingHandle& h) noexcept;
283+
MemAllocationHandle va_mapping_allocation(const VaMappingHandle& h) noexcept;
284+
285+
// Build an immutable range from mappings in ascending, contiguous order. A
286+
// grow builds a new range for its result and never changes the input's.
287+
// May throw std::bad_alloc.
288+
VmmRangeHandle create_vmm_range(const std::vector<VaMappingHandle>& mappings);
289+
290+
// The range of a device pointer handle created by deviceptr_create_vmm. Only
291+
// for such handles: the caller (VirtualMemoryBuffer) guarantees the origin.
292+
// Empty for an empty handle.
293+
VmmRangeHandle vmm_range(const DevicePtrHandle& h) noexcept;
294+
295+
// Range accessors: a copy of the mapping list, which a grow extends and turns
296+
// into a new range, and the range total. Reads of an immutable range need no
297+
// synchronization.
298+
std::vector<VaMappingHandle> vmm_range_mappings(const VmmRangeHandle& range); // may throw
299+
size_t vmm_range_total(const VmmRangeHandle& range) noexcept;
300+
301+
// Create a device pointer handle whose box owns a range (an empty range for
302+
// a size-zero buffer). The box records no deallocation stream; set one with
303+
// set_deallocation_stream. When the last reference is released, the recorded
304+
// stream is synchronized (skipped with a report when that would disturb a
305+
// capture) and the box freed, which unmaps every mapping this buffer was the
306+
// last to hold. Returns empty handle for a null range handle.
307+
DevicePtrHandle deviceptr_create_vmm(CUdeviceptr base, const VmmRangeHandle& range);
308+
243309
// ============================================================================
244310
// Library handle functions
245311
// ============================================================================

‎cuda_core/cuda/core/_cpp/rt/driver_api.hpp‎

Lines changed: 13 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -126,7 +126,19 @@ namespace cuda_core::rt {
126126
X(cuTexObjectCreate, 5000) \
127127
X(cuTexObjectDestroy, 5000) \
128128
X(cuSurfObjectCreate, 5000) \
129-
X(cuSurfObjectDestroy, 5000)
129+
X(cuSurfObjectDestroy, 5000) \
130+
/* Virtual memory management (VMM_DESIGN.md) */ \
131+
X(cuMemCreate, 10020) \
132+
X(cuMemRelease, 10020) \
133+
X(cuMemAddressReserve, 10020) \
134+
X(cuMemAddressFree, 10020) \
135+
X(cuMemMap, 10020) \
136+
X(cuMemUnmap, 10020) \
137+
X(cuMemSetAccess, 10020) \
138+
/* cuda-bindings requests 7000 (PTDS) or 2000 (legacy) */ \
139+
X(cuStreamSynchronize, 7000) \
140+
X(cuStreamIsCapturing, 10000) \
141+
X(cuThreadExchangeStreamCaptureMode, 10010)
130142

131143
#define CUDA_CORE_DECLARE_DRIVER_FN(name, introduced) extern decltype(&name) p_##name;
132144
CUDA_CORE_DRIVER_FUNCTIONS(CUDA_CORE_DECLARE_DRIVER_FN)

‎cuda_core/cuda/core/_cpp/rt/internal.hpp‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -62,6 +62,12 @@ ContextHandle deallocation_context(const DeallocationStream& stream) noexcept;
6262
// Implemented in stream.cpp
6363
bool make_deallocation_stream(const StreamHandle& h, DeallocationStream& out) noexcept;
6464

65+
// Implemented in virtual_memory.cpp. Synchronize the stream a VMM buffer
66+
// recorded before its mappings are released. Skips the sync, with a report
67+
// when a capture is the reason, if no stream was recorded, the interpreter is
68+
// finalizing, or the sync would disturb a graph capture (VMM_DESIGN.md).
69+
void vmm_sync_before_release(const DeallocationStream& stream) noexcept;
70+
6571
// Decorate a status-returning cleanup call to report whenever it fails. CUDA
6672
// calls (CUresult) are reported with the error name and description; NVRTC,
6773
// NVVM and nvJitLink calls (integer status codes) with the raw code.
@@ -126,6 +132,9 @@ const WarnOnFailure<p_cuSurfObjectDestroy> pw_cuSurfObjectDestroy{"cuSurfObjectD
126132
const WarnOnFailure<p_cuGreenCtxDestroy> pw_cuGreenCtxDestroy{"cuGreenCtxDestroy"};
127133
const WarnOnFailure<p_cuMemPoolDestroy> pw_cuMemPoolDestroy{"cuMemPoolDestroy"};
128134
const WarnOnFailure<p_cuMemFreeHost> pw_cuMemFreeHost{"cuMemFreeHost"};
135+
const WarnOnFailure<p_cuMemRelease> pw_cuMemRelease{"cuMemRelease"};
136+
const WarnOnFailure<p_cuMemUnmap> pw_cuMemUnmap{"cuMemUnmap"};
137+
const WarnOnFailure<p_cuMemAddressFree> pw_cuMemAddressFree{"cuMemAddressFree"};
129138
const WarnOnFailure<p_cuGraphDestroy> pw_cuGraphDestroy{"cuGraphDestroy"};
130139
const WarnOnFailure<p_cuGraphExecDestroy> pw_cuGraphExecDestroy{"cuGraphExecDestroy"};
131140
const WarnOnFailure<p_cuGraphicsUnregisterResource> pw_cuGraphicsUnregisterResource{"cuGraphicsUnregisterResource"};

‎cuda_core/cuda/core/_cpp/rt/memory.cpp‎

Lines changed: 46 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -123,6 +123,15 @@ struct DevicePtrBox {
123123
// stream tokens carry a bound context.
124124
mutable DeallocationStream deallocation;
125125
};
126+
127+
// The box behind a VirtualMemoryResource buffer: the range of mappings it
128+
// owns (VMM_DESIGN.md). Every handle on a VirtualMemoryBuffer comes from
129+
// deviceptr_create_vmm, so vmm_range() downcasts without a tag; there are no
130+
// virtual functions, so other boxes pay nothing. A size-zero buffer has an
131+
// empty range.
132+
struct VmmDevicePtrBox : DevicePtrBox {
133+
VmmRangeHandle range;
134+
};
126135
} // namespace
127136

128137
// Recovers the owning DevicePtrBox from the aliased CUdeviceptr pointer.
@@ -137,9 +146,10 @@ static DevicePtrBox* get_box(const DevicePtrHandle& h) {
137146
);
138147
}
139148

140-
// Return the stream that orders a device pointer's deallocation.
149+
// Return the stream that orders a device pointer's deallocation; empty for an
150+
// empty handle, as set_deallocation_stream rejects one.
141151
StreamHandle deallocation_stream(const DevicePtrHandle& h) noexcept {
142-
return get_box(h)->deallocation.h_stream;
152+
return h ? get_box(h)->deallocation.h_stream : StreamHandle{};
143153
}
144154

145155
// Replace the stream that orders a device pointer's deallocation.
@@ -300,6 +310,36 @@ DevicePtrHandle deviceptr_create_mapped_graphics(
300310
return DevicePtrHandle(box, &box->resource);
301311
}
302312

313+
// ============================================================================
314+
// Virtual memory ranges (VMM_DESIGN.md)
315+
// ============================================================================
316+
317+
DevicePtrHandle deviceptr_create_vmm(CUdeviceptr base, const VmmRangeHandle& range) {
318+
if (!range) { // a null handle; an empty range is valid
319+
err = CUDA_ERROR_INVALID_VALUE;
320+
return {};
321+
}
322+
auto box = std::shared_ptr<VmmDevicePtrBox>(
323+
new VmmDevicePtrBox{{base, DeallocationStream{}}, range},
324+
[](VmmDevicePtrBox* b) {
325+
GILReleaseGuard gil;
326+
// Order the release on this buffer's recorded stream, then free
327+
// the box, which drops the range: every mapping this buffer was
328+
// the last to hold unmaps, and its reservation and allocation
329+
// follow.
330+
vmm_sync_before_release(b->deallocation);
331+
delete b;
332+
}
333+
);
334+
return DevicePtrHandle(box, &box->resource);
335+
}
336+
337+
VmmRangeHandle vmm_range(const DevicePtrHandle& h) noexcept {
338+
// Only for handles from deviceptr_create_vmm; the VirtualMemoryBuffer
339+
// class guarantees that for its callers.
340+
return h ? static_cast<VmmDevicePtrBox*>(get_box(h))->range : VmmRangeHandle{};
341+
}
342+
303343
// ============================================================================
304344
// MemoryResource-owned Device Pointer Handles
305345
// ============================================================================
@@ -325,6 +365,10 @@ DevicePtrHandle deviceptr_create_with_mr(CUdeviceptr ptr, size_t size, PyObject*
325365
[mr, size](DevicePtrBox* b) {
326366
GILAcquireGuard gil;
327367
if (gil.acquired()) {
368+
// The last reference may go while an exception propagates
369+
// through the releasing caller; deallocate() must run with a
370+
// clean error state and leave that exception in place.
371+
PendingExceptionGuard pending;
328372
if (mr_dealloc_cb) {
329373
const DeallocationStream& stream = b->deallocation;
330374
cleanup_in_context(

‎cuda_core/cuda/core/_cpp/rt/py.hpp‎

Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -94,6 +94,45 @@ class GILAcquireGuard {
9494
bool acquired_;
9595
};
9696

97+
// Save the Python exception in flight and restore it on scope exit, dropping
98+
// anything the scope itself raised. A deleter may run while an exception is
99+
// propagating through the caller that released the last reference; Python
100+
// code it calls (a MemoryResource's deallocate, the warnings machinery) must
101+
// start with a clean error state and must not leave that exception cleared,
102+
// or the caller returns an error with no exception set. Allocates nothing.
103+
// Construct with the GIL held and destroy before releasing it.
104+
class PendingExceptionGuard {
105+
public:
106+
PendingExceptionGuard() noexcept {
107+
#if PY_VERSION_HEX >= 0x030C0000
108+
pending_ = PyErr_GetRaisedException();
109+
#else
110+
PyErr_Fetch(&type_, &value_, &tb_);
111+
#endif
112+
}
113+
114+
~PendingExceptionGuard() {
115+
PyErr_Clear();
116+
#if PY_VERSION_HEX >= 0x030C0000
117+
PyErr_SetRaisedException(pending_);
118+
#else
119+
PyErr_Restore(type_, value_, tb_);
120+
#endif
121+
}
122+
123+
PendingExceptionGuard(const PendingExceptionGuard&) = delete;
124+
PendingExceptionGuard& operator=(const PendingExceptionGuard&) = delete;
125+
126+
private:
127+
#if PY_VERSION_HEX >= 0x030C0000
128+
PyObject* pending_ = nullptr;
129+
#else
130+
PyObject* type_ = nullptr;
131+
PyObject* value_ = nullptr;
132+
PyObject* tb_ = nullptr;
133+
#endif
134+
};
135+
97136
// as_py() - convert handle to Python wrapper object (returns new reference)
98137
namespace detail {
99138
// n.b. class lookup is not cached to avoid deadlock hazard, see DESIGN.md
@@ -205,6 +244,18 @@ inline PyObject* as_py(const SurfObjectHandle& h) noexcept {
205244
return detail::make_py("cuda.bindings.driver", "CUsurfObject", as_intptr(h));
206245
}
207246

247+
inline PyObject* as_py(const MemAllocationHandle& h) noexcept {
248+
return detail::make_py("cuda.bindings.driver", "CUmemGenericAllocationHandle", as_intptr(h));
249+
}
250+
251+
inline PyObject* as_py(const VaReservationHandle& h) noexcept {
252+
return detail::make_py("cuda.bindings.driver", "CUdeviceptr", as_intptr(h));
253+
}
254+
255+
inline PyObject* as_py(const VaMappingHandle& h) noexcept {
256+
return detail::make_py("cuda.bindings.driver", "CUdeviceptr", as_intptr(h));
257+
}
258+
208259
// ============================================================================
209260
// Python-coupled API: the prototypes that take or return PyObject*
210261
// ============================================================================

‎cuda_core/cuda/core/_cpp/rt/py_driver_fns.cpp‎

Lines changed: 0 additions & 33 deletions
Original file line numberDiff line numberDiff line change
@@ -48,39 +48,6 @@ std::atomic<bool> unavailable_reported[kTables];
4848
std::mutex fill_mutex;
4949
char fill_error[kTables][512] = {}; // guarded by fill_mutex
5050

51-
// Saves the pending Python exception on construction and restores it on
52-
// destruction. The Python calls in between start from a clean error state, and
53-
// the caller's exception survives. Requires the GIL.
54-
class PendingExceptionGuard {
55-
public:
56-
PendingExceptionGuard() noexcept {
57-
#if PY_VERSION_HEX >= 0x030C0000
58-
exc_ = PyErr_GetRaisedException();
59-
#else
60-
PyErr_Fetch(&type_, &value_, &traceback_);
61-
#endif
62-
}
63-
~PendingExceptionGuard() {
64-
PyErr_Clear(); // drop anything the guarded calls left set
65-
#if PY_VERSION_HEX >= 0x030C0000
66-
PyErr_SetRaisedException(exc_);
67-
#else
68-
PyErr_Restore(type_, value_, traceback_);
69-
#endif
70-
}
71-
PendingExceptionGuard(const PendingExceptionGuard&) = delete;
72-
PendingExceptionGuard& operator=(const PendingExceptionGuard&) = delete;
73-
74-
private:
75-
#if PY_VERSION_HEX >= 0x030C0000
76-
PyObject* exc_ = nullptr;
77-
#else
78-
PyObject* type_ = nullptr;
79-
PyObject* value_ = nullptr;
80-
PyObject* traceback_ = nullptr;
81-
#endif
82-
};
83-
8451
std::size_t index_of(FnTable table) noexcept { return static_cast<std::size_t>(table); }
8552

8653
const char* module_name(FnTable table) noexcept {

‎cuda_core/cuda/core/_cpp/rt/py_report.cpp‎

Lines changed: 1 addition & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -34,12 +34,7 @@ void report_message(const char* message) noexcept {
3434
GILAcquireGuard gil;
3535
if (gil.acquired()) {
3636
// Deleters can run while a Python exception is propagating; keep it.
37-
#if PY_VERSION_HEX >= 0x030C0000
38-
PyObject* pending = PyErr_GetRaisedException();
39-
#else
40-
PyObject *pending_type, *pending_value, *pending_tb;
41-
PyErr_Fetch(&pending_type, &pending_value, &pending_tb);
42-
#endif
37+
PendingExceptionGuard pending;
4338
bool interrupted = false;
4439
if (PyErr_WarnEx(category, message, 1) != 0) {
4540
interrupted = PyErr_ExceptionMatches(PyExc_KeyboardInterrupt);
@@ -52,11 +47,6 @@ void report_message(const char* message) noexcept {
5247
Py_XDECREF(subject);
5348
}
5449
}
55-
#if PY_VERSION_HEX >= 0x030C0000
56-
PyErr_SetRaisedException(pending);
57-
#else
58-
PyErr_Restore(pending_type, pending_value, pending_tb);
59-
#endif
6050
if (interrupted) {
6151
PyErr_SetInterrupt();
6252
}

0 commit comments

Comments
 (0)