GPU acceleration across the FEM pipeline
Enter CPU stage timings, choose stages to offload, and account for transfers and changes in solver work.
From element work to a global vector
6 bilinear quadrilaterals · scalar diffusion3D unavailable. The operator and shared-node calculations remain available.
Enter non-overlapping CPU stage totals from the same run. Presets are synthetic milliseconds; selecting a workload restores its preset.
Faster kernels eventually stop helping
What does an operator application store?
Kernel limit: bandwidth or computation?
Independent model · illustrative hardwareThe point is a theoretical bound, not measured throughput. Use bytes moved at one specified memory level. Index traffic counts as bytes; cache reuse changes intensity. This plot does not set the pipeline speedup above. Roofline reference ↗
Data residency across repeated iterations
mesh · boundary data
state → residual → operator / solve → updated state ↻
selected results
Repeated CPU access inside the loop may introduce copies and synchronization. Keep required state resident when possible; count unavoidable output, contact, and distributed communication in the measured overhead.