Repository navigation
Replies: 1 comment
|
Your code trips before the atomic part: use cuda_device::atomic::{AtomicOrdering, BlockAtomicU32};
let bin = val as usize;
let p = unsafe { SharedArray::as_raw_mut_ptr(&raw mut TILE).add(bin) }; // 1
let counter = unsafe { BlockAtomicU32::from_ptr(p) }; // 2
counter.fetch_add(1, AtomicOrdering::Relaxed); // 3
Why Another rule to keep in mind: plain and atomic access to the same slot must not overlap in time, so structure the kernel in phases: Note, your original cast also works once the pointer is taken correctly ( PS: The GitHub discussions tab isn't actively monitored. We have a community discord server for questions like these. Feel free to check it out |
Uh oh!
There was an error while loading. Please reload this page.
I am trying to implement a histogram kernel as shown below. However, I am unable to figure out how to atomically increment an element of the SharedArray.
All reactions