I am trying to use sigma to make inference with llama2 7b model weights with GPU-MPC/experiments/sigma/sigma.cu or sigma_offline_online.cu.
However, I found that the created GPULlama were only initiated both in key generation phase and online inference phase. Could you show me where the server's model weights are loaded and used?
I am trying to use sigma to make inference with llama2 7b model weights with GPU-MPC/experiments/sigma/sigma.cu or sigma_offline_online.cu.
However, I found that the created GPULlama were only initiated both in key generation phase and online inference phase. Could you show me where the server's model weights are loaded and used?