FEM PERFORMANCE · END-TO-END COSTIllustrative timing model — not a benchmark

GPU acceleration across the FEM pipeline

Enter CPU stage timings, choose stages to offload, and account for transfers and changes in solver work.

From element work to a global vector

6 bilinear quadrilaterals · scalar diffusion
Click a node or element · drag to orbit · scroll to zoom

Enter non-overlapping CPU stage totals from the same run. Presets are synthetic milliseconds; selecting a workload restores its preset.

CPU reference100 units
Proposed execution
Overall speedup

Faster kernels eventually stop helping

What does an operator application store?

Kernel limit: bandwidth or computation?

Independent model · illustrative hardware

The point is a theoretical bound, not measured throughput. Use bytes moved at one specified memory level. Index traffic counts as bytes; cache reuse changes intensity. This plot does not set the pipeline speedup above. Roofline reference ↗

Data residency across repeated iterations

CPU initialization
mesh · boundary data
upload →
GPU iteration loop
state → residual → operator / solve → updated state ↻
→ download
Output
selected results

Repeated CPU access inside the loop may introduce copies and synchronization. Keep required state resident when possible; count unavoidable output, contact, and distributed communication in the measured overhead.