This public downstream tracks current upstream llama.cpp while
developing the ROCmFPX weight-format family for AMD GPUs. Normal Vulkan support
remains enabled; HIP and CPU are also supported build targets. Existing NVFP4
GGUF tensors can be loaded natively and remain bit-exact when an NVFP4 model is
completed with the NVFP4 quantization preset.
Carlo Pasquale (Charlie12345) is the creator and founder of the ROCmFPX format family and of ROCmFP3, ROCmFP4, ROCmFP6, and ROCmFP8. See ROCmFPX documentation, NOTICE, and the same upstream MIT license terms with the ROCmFPX copyright line.
ROCmFPX preserves the upstream llama.cpp lineage and the public legacy
ROCmFPX lineage. Original commit authors and commit IDs remain reachable. The
legacy lineage is connected by an ancestry-only merge whose source tree is
identical to the current ROCmFPX tree, so legacy code does not replace the
current implementation.
GitHub contributor displays can lag behind repository history. The complete
commit-level author record is also available with git shortlog -sne main.
The currently qualified formats are experimental. The on-disk layouts of ROCmFP2/3/4/6/8 are frozen for compatibility; new kernel and quantizer work must preserve their encoded sizes and semantics.
Canonical dual-scale S40 ROCmFP2 uses GGUF tensor type 111. Its payload is
still exactly 10 bytes per 32 weights (2.5 bpw); only the tensor type tag moved.
Legacy type 107 is reserved as ambiguous because both dual-scale S40 and an
affine code * scale - offset layout were emitted with that same ID and block
size. ROCmFPX refuses type 107 instead of guessing and silently corrupting a
model.
New quantizations write type 111 automatically. To audit an older file, run:
scripts/rocmfpx/retag-legacy-rocmfp2.py MODEL.ggufOnly when the file's provenance confirms that it uses dual-scale S40, make a backup and retag its tensor headers in place:
scripts/rocmfpx/retag-legacy-rocmfp2.py MODEL.gguf \
--layout s40-dual-scale-v1 --applyThe tool does not inspect or infer the layout because the two interpretations cannot be distinguished from the bytes. Do not use it on affine ROCmFP2 files.
The optional ROCmFPXVulkan extension
preserves the proven Charlie-era ROCmFP4 Vulkan path as a separate backend. It
does not replace or patch llama.cpp's normal Vulkan0 backend. The extension
is disabled by default and appears as ROCmFPXVulkan0 only when it is built
and explicitly loaded.
Build both the normal Vulkan backend and the optional extension on Linux:
cmake -S . -B build-rocmfpx-vulkan -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DBUILD_SHARED_LIBS=ON \
-DGGML_VULKAN=ON \
-DROCMFPX_VULKAN_PLUGIN=ON
cmake --build build-rocmfpx-vulkan \
--target llama-cli rocmfpx-vulkan-plugin -jLoad the plugin and select its device:
export ROCMFPX_PLUGIN_PATH="$PWD/build-rocmfpx-vulkan/bin/rocmfpx-vulkan-plugin.so"
build-rocmfpx-vulkan/bin/llama-cli --list-devices
build-rocmfpx-vulkan/bin/llama-cli \
-m /absolute/path/to/model.gguf \
-dev ROCmFPXVulkan0 -ngl 999Keep rocmfpx-vulkan-plugin.so and its sibling
libggml-rocmfpx-vulkan.so together. Both files must come from the same build
and ROCmFPX commit. --list-devices should show both Vulkan0 and
ROCmFPXVulkan0; remove ROCMFPX_PLUGIN_PATH to return to the normal backend.
See the extension guide for matched
ROCmFP4 benchmarks and qualification limits.
ROCmFPX plugin ABI v1 lets trusted native shared libraries register a standard
ggml backend or receive model-open and model-close notifications for PLE, KV
checkpoint, repack-cache, and expert-streaming sidecars. Set
ROCMFPX_PLUGIN_PATH to a plugin file or directory before starting a ROCmFPX
tool. Use : between entries on Linux and macOS and ; on Windows. Directories
are scanned once, non-recursively, in sorted order; the working directory is
never searched automatically.
A minimal C plugin looks like this:
#define ROCMFPX_PLUGIN_BUILD
#include "rocmfpx-plugin.h"
static int on_load(void) {
return 0;
}
static const struct rocmfpx_plugin_v1 plugin = {
ROCMFPX_PLUGIN_ABI_VERSION,
sizeof(struct rocmfpx_plugin_v1),
"example-sidecar",
"0.1.0",
ROCMFPX_PLUGIN_CAP_SIDECAR,
on_load,
0,
0,
0,
};
ROCMFPX_PLUGIN_EXPORT const struct rocmfpx_plugin_v1 * rocmfpx_plugin_query(
uint32_t host_abi_version,
const struct rocmfpx_plugin_host_v1 * host) {
if (host_abi_version != ROCMFPX_PLUGIN_ABI_VERSION || !host ||
host->abi_version != ROCMFPX_PLUGIN_ABI_VERSION) {
return 0;
}
return &plugin;
}Build and load it on Linux:
cc -shared -fPIC -I/path/to/ROCmFPX/include \
example-sidecar.c -o example-sidecar.so
ROCMFPX_PLUGIN_PATH="$PWD/example-sidecar.so" \
/path/to/ROCmFPX/build/bin/llama-cli --list-devicesEvery plugin must export rocmfpx_plugin_query, validate the ABI, return a
static size-versioned descriptor, and keep that descriptor and its strings
alive until unload. A backend plugin copies backend_load and plugin_path
during the query, then registers its matching ggml backend from on_load. A
sidecar uses on_model_open and on_model_close and keys its state by
model_id. ABI v1 provides discovery, backend registration, and lifecycle
notifications; capability flags alone do not intercept tensors or token
generation. Plugins run as native code with the same permissions as ROCmFPX,
so load only libraries you trust. See the complete plugin and sidecar ABI
guide and the tested
rocmfpx-test-plugin example.
Windows users with two AMD GPUs can evaluate Charlie12345's external Windows AMD Multi-GPU Bridge. For ROCmFPX and llama.cpp, use its dedicated Windows installation guide, not the separate vLLM adapter instructions.
The bridge keeps the ROCmFPX source tree unchanged. It supplies an external
roc::rccl CMake package to the existing HIP collective interface, so build
ROCmFPX with GGML_HIP_RCCL=ON, point rccl_DIR at the installed bridge, and
use the bridge's run-with-llama-plugin.ps1 launcher. The tested Windows path
also requires GGML_CUDA_NO_PEER_COPY=ON; follow the external guide for the
matching ROCm version, GPU target, DLL layout, health probes, and model-specific
launch flags.
Despite the launcher's name, this transport is not loaded through
ROCMFPX_PLUGIN_PATH and is not a ROCmFPX ABI v1 plugin. ABI v1 can register a
ggml backend or receive model lifecycle notifications, but it cannot inject a
link-time RCCL provider or intercept collectives. Pointing
ROCMFPX_PLUGIN_PATH at rccl.dll will therefore do nothing. The current
external-package design is the upstream-safe integration: update ROCmFPX
normally, then configure a fresh HIP build against the bridge package.
The bridge is experimental, source-only, and currently qualified only on the hardware and revisions listed by its maintainers. Validate its small parity and transport probes before loading a large model; a selectable GPU target is not the same as a runtime-qualified configuration.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain


