GTaP Documentation
GTaP is a pragma-based task-parallel programming system for GPUs, implemented in CUDA C++. In GTaP, tasks execute within a persistent CUDA kernel; each task is not a separate kernel launch.
GTaP combines a header-only CUDA runtime with a Clang extension that lowers GTaP pragmas into CUDA device code. The source code is available on GitHub.
GTaP is a research prototype under active development. Its interfaces and internal mechanisms may change.
Documentation
Tutorial
Start with Installation, then run the first program in Quickstart. Continue with Execution Modes, Rules and Pitfalls, and Profiling as needed.
API Reference
Look up exact syntax and interfaces in Pragmas, Runtime Functions, Configuration, and the Profiling API.
Programming at a Glance
A CUDA kernel uses entry to create the root task. A task function uses task to spawn direct children and taskwait to join them:
__device__ int d_result;
#pragma gtap function
__device__ int fib(int n) {
if (n < 2) return n;
int left, right;
#pragma gtap task
left = fib(n - 1);
#pragma gtap task
right = fib(n - 2);
#pragma gtap taskwait
return left + right;
}
__global__ void exec_kernel(int n) {
#pragma gtap entry
d_result = fib(n);
}Features
- Fork-join task parallelism on GPUs. Express task creation and join points with
#pragma gtap taskand#pragma gtap taskwait. - Thread and block modes. Execute one task on one CUDA thread, or cooperatively with an entire CUDA thread block.
- GPU-resident scheduling. A GPU-resident scheduler distributes runnable tasks through randomized work stealing.
- Divergence-aware queueing. Group tasks with similar expected execution paths to reduce inter-task warp divergence in thread mode.
Publication
Yuki Maeda and Kenjiro Taura.
GTaP: A GPU-Resident Fork-Join Task-Parallel System with a Pragma-Based Interface.
arXiv:2604.05982, 2026.