Skip to content

GTaP Documentation ​

GTaP is a pragma-based task-parallel programming system for GPUs, implemented in CUDA C++. In GTaP, tasks execute within a persistent CUDA kernel; each task is not a separate kernel launch.

GTaP combines a header-only CUDA runtime with a Clang extension that lowers GTaP pragmas into CUDA device code. The source code is available on GitHub.

GTaP is a research prototype under active development. Its interfaces and internal mechanisms may change.

Documentation ​

Tutorial ​

Start with Installation, then run the first program in Quickstart. Continue with Execution Modes, Rules and Pitfalls, and Profiling as needed.

API Reference ​

Look up exact syntax and interfaces in Pragmas, Runtime Functions, Configuration, and the Profiling API.

Programming at a Glance ​

A CUDA kernel uses entry to create the root task. A task function uses task to spawn direct children and taskwait to join them:

cpp
__device__ int d_result;

#pragma gtap function
__device__ int fib(int n) {
    if (n < 2) return n;
    int left, right;
    #pragma gtap task
    left = fib(n - 1);
    #pragma gtap task
    right = fib(n - 2);
    #pragma gtap taskwait
    return left + right;
}

__global__ void exec_kernel(int n) {
    #pragma gtap entry
    d_result = fib(n);
}

Features ​

  • Fork-join task parallelism on GPUs. Express task creation and join points with #pragma gtap task and #pragma gtap taskwait.
  • Thread and block modes. Execute one task on one CUDA thread, or cooperatively with an entire CUDA thread block.
  • GPU-resident scheduling. A GPU-resident scheduler distributes runnable tasks through randomized work stealing.
  • Divergence-aware queueing. Group tasks with similar expected execution paths to reduce inter-task warp divergence in thread mode.

Publication ​

Yuki Maeda and Kenjiro Taura.
GTaP: A GPU-Resident Fork-Join Task-Parallel System with a Pragma-Based Interface.
arXiv:2604.05982, 2026.