Skip to content
vast-cow's blog
Go back

Building llama.cpp with CUDA in an NVIDIA HPC SDK Environment

Edit page

If you are working in an NVIDIA HPC SDK environment and want to build llama.cpp with CUDA support, one reliable approach is to use GCC/G++ for the C/C++ parts and NVCC for the CUDA parts.

This setup is practical because some compiler warning flags used by projects like ggml/llama.cpp are commonly supported by GCC/Clang, but may not be accepted by other C++ compilers. By explicitly selecting gcc and g++, you reduce the risk of compiler-flag incompatibilities, while still enabling CUDA with nvcc.

Run the following commands from the project root directory:

cmake -S . -B build \
  -DGGML_CUDA=ON \
  -DCMAKE_C_COMPILER=gcc \
  -DCMAKE_CXX_COMPILER=g++ \
  -DCMAKE_CUDA_COMPILER=nvcc

cmake --build build -j

What These Options Do

Summary

In an NVIDIA HPC SDK environment, explicitly selecting gcc/g++ for host compilation and nvcc for CUDA compilation is a simple and effective way to build llama.cpp with CUDA enabled. This approach is also easy to reproduce across systems and tends to avoid compiler option conflicts.


Edit page
Share this post:

Comments


Previous Post
Building llama.cpp in an Environment Without curl Headers
Next Post
A Simple Windows Tool to Hide Chrome Picture-in-Picture Behind Other Windows