Quickstart
This tutorial walks through a minimal GTaP program. It introduces the two execution modes, builds and runs thread-mode Fibonacci, and then explains task functions, child tasks, joins, and host-side launch.
All commands are run from the root of the cloned GTaP repository and assume that the CUDA Toolkit is available.
1. GTaP's two execution modes
GTaP provides two execution modes:
| Mode | One task runs on | Use when |
|---|---|---|
| Thread mode | One CUDA thread | The task body is fine-grained and mostly sequential |
| Block mode | One CUDA thread block | The task body uses shared memory, block synchronization, or cooperative parallel work |
This tutorial uses thread mode because each Fibonacci task is fine-grained and sequential. See Execution Modes for further details and the block-mode programming model.
2. Build and run Fibonacci
Set CUDA_PATH to the CUDA Toolkit root and choose the GPU architecture. Then build and run the example:
export CUDA_PATH=/path/to/cuda
export CUDA_ARCH=sm_90
"${CUDA_PATH}/bin/nvcc" --version
cd examples/fib
make
./bin/fib_threadYou can override paths and architecture settings on the command line:
make GTAP_ROOT=/path/to/GTaP \
CUDA_PATH=/path/to/cuda \
CUDA_ARCH=sm_903. Understand the programming model
The example includes the thread-mode runtime:
#include "gtap_thread.cuh"The work performed by a GTaP task must be defined in a separate __device__ function, called a task function. Place #pragma gtap function immediately before its definition:
#pragma gtap function
__device__ int fib(int n) {
if (n < 2) return n;
int a, b;
#pragma gtap task
a = fib(n - 1);
#pragma gtap task
b = fib(n - 2);
#pragma gtap taskwait
return a + b;
}This function uses three GTaP concepts:
#pragma gtap functiondeclares a task function.#pragma gtap taskspawns the immediately following task-function call as a child task.#pragma gtap taskwaitwaits for the direct child tasks and then resumes the parent.
A task spawn does not block the parent. A taskwait waits only for direct children spawned since the previous taskwait; it is not a device-wide barrier.
Place #pragma gtap entry before the root task-function call in a __global__ kernel. This kernel is the entry point for the GTaP computation:
__global__ void exec_kernel(int n) {
#pragma gtap entry
d_result = fib(n);
}These four pragmas are sufficient for the first program. See Rules and Pitfalls for execution constraints and the Pragma Reference for the exact syntax and restrictions of each pragma.
4. Launch the computation
The host initializes GTaP, then uses gtap_launch to launch the CUDA kernel containing the entry directive. This starts the root task. gtap_synchronize waits for the computation to finish, and gtap_finalize releases the runtime:
gtap_thread_config config{
.grid_size = 4000,
.block_size = 32,
.max_tasks_per_warp = 100000,
.num_queues = 1,
};
cudaError_t status = gtap_initialize(config);
if (status != cudaSuccess) return 1;
status = gtap_launch(exec_kernel, 40);
if (status == cudaSuccess)
status = gtap_synchronize();
gtap_finalize();For every function argument, return value, error, and repeated-run behavior, see the Runtime API Reference. See the Configuration Reference for all configuration fields and validation rules.