Skip to content

Deploy LLM with AidGen

Introduction

Edge deployment of a large language model (LLM) refers to compressing, quantizing, and deploying a large model that originally runs in the cloud onto a local device, enabling offline, low-latency natural language understanding and generation. This chapter is based on the AidGen inference engine and demonstrates how to complete the deployment, loading, and conversation process of a large language model on an edge device.

In this case, large language model inference runs on the device, and C++ code calls relevant interfaces to receive user input and return conversation results in real time.

  • Device: IQ8275
  • System: Ubuntu 24.04
  • Model: Qwen2.5-0.5B-Instruct

Supported Platforms

PlatformExecution Method
IQ8275Ubuntu 24.04

Prerequisites

  1. IQ8275 hardware

  2. Ubuntu 24.04 system

  3. Prepare the model file

Visit Model Farm: Qwen2.5-0.5B-Instruct to download the model resource files

💡Note

This model does not yet support IQ8. You can use the QCS8550 chip model as a substitute, or choose another LLM model that is supported on IQ8.

System Dependency Configuration

Configure the AidLux Package Source

bash
# Download the correct public key
sudo wget -O- https://archive.aidlux.com/ubuntu24/public.key | gpg --dearmor | sudo tee /etc/apt/trusted.gpg.d/private-aidlux.gpg > /dev/null

# Edit the source list file
sudo vim /etc/apt/sources.list.d/private-aidlux.list

# Add the repository provided by AidLux to the source file
deb [arch=arm64 signed-by=/etc/apt/trusted.gpg.d/private-aidlux.gpg] https://archive.aidlux.com/ubuntu24 noble main

# Update the package cache
sudo apt update

After the update is complete, you can use the following command to list the SDK dependencies officially provided by AidLux:

bash
sudo apt list | grep aid | grep unknown
bash
# Install software
# Must be installed first because they are not included in the system by default
sudo apt install python3 python3-pip libopencv-dev python3-opencv  net-tools
# Must be installed before aidlite
sudo apt install aidlux-aistack-base aidrtcm

# Install aidlite and its dependencies
sudo apt install aid-lms aidlms-sdk aidlite-sdk cmake
sudo apt-get install libfmt-dev nlohmann-json3-dev
sudo apt install aidlite-*

# Enable DSP support
sudo apt-get install qcom-fastrpc1
sudo apt-get install qcom-fastrpc-dev

# Install aidgen-sdk
sudo apt install aidgen-sdk
sudo apt install aidgen-qnn*

# Install the mms service
sudo apt install aid-mms

# Enable GPU support
sudo apt-add-repository -s ppa:ubuntu-qcom-iot/qcom-ppa
sudo apt install qcom-adreno-cl1
sudo ln -s /usr/lib/aarch64-linux-gnu/libOpenCL.so.1 /usr/lib/aarch64-linux-gnu/libOpenCL.so

After the installation is complete, check that the aidlite and aidgen directories have been added under /usr/local/share.

Device Authorization

Get the Device SN

bash
cat  /sys/devices/soc0/serial_number

Get the License File

Provide the SN to APLUX technical support so that they can generate the device-specific license file. Place the generated file under /etc/opt/aidlux/license/AidLuxLics.

Activate the License

bash
sudo /opt/aidlux/cpf/aid-lms/manager.sh restart

Case Deployment

Step 1: Copy the AidGen SDK Code Example

bash
# Copy the test code
cd /home/ubuntu/aidllm

cp -r /usr/local/share/aidgen/examples/ ./

Step 2: Upload & Unzip the Model Resources

  • Upload the downloaded model resources to the edge device.

  • Unzip the model resources to the /home/ubuntu/aidllm directory:

bash
cd /home/ubuntu/aidllm
unzip qnn229_qcs8550_cl4096.zip

Step 3: Confirm the Resource Files

The files are distributed as follows:

bash
/home/ubuntu/aidllm/qnn229_qcs8550_cl4096
├── tokenizer_config.json
├── tokenizer.json
├── qwen2.5-0.5b-instruct_qnn229_qcs8550_4096_2_of_2.serialized.bin
├── qwen2.5-0.5b-instruct_qnn229_qcs8550_4096_1_of_2.serialized.bin
├── qwen2.5-0.5b-instruct-tokenizer.json
├── qwen2.5-0.5b-instruct-htp.json
├── prompt.conf
├── metadata.json
├── htp_backend_ext_config.json
├── genie_config.json
├── config_linux.json
├── chat.txt
├── aidgen_config.json
├── aidgen_chat_template.txt

Step 4: Set the Conversation Template

💡Note

Refer to the aidgen_chat_template.txt file in the model resource package for the conversation template.

Modify the test_aidgen_t2t.cpp file according to the large model template:

cpp
    // ========================================================================
    // 5. Build the prompt template (Qwen2 format)
    // ========================================================================
    std::string system_prompt =
        "<|im_start|>system\n"
        "You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n";

    auto make_user_turn = [](const std::string& text) -> std::string {
        return "<|im_start|>user\n" + text + "<|im_end|>\n<|im_start|>assistant\n";
    };

Step 5: Compile and Run

bash
cd /home/ubuntu/aidllm/examples

# Compile
mkdir build && cd build
cmake .. && make

cp test_t2t ../../qnn229_qcs8550_cl4096

cd /home/ubuntu/aidllm/qnn229_qcs8550_cl4096
./test_t2t aidgen_config.json 'Introduce large language models' qnn240
  • After the model runs successfully, the following log output is displayed:
bash
============================================================
  Turn 1: Single-turn inference
============================================================
User: Introduce large language models
Assistant: [API] Generator::run(prompt, callback)
[BOS]Large language models are an artificial intelligence technology that can simulate human thinking and behavior. In large language models, we can see many advanced algorithms and technologies, such as deep learning and neural networks. These technologies enable machines to make complex decisions and reason, in order to achieve better results. In addition, large language models can also understand and generate natural language through natural language processing, which enables them to have more natural conversations with humans.

In general, large language models are a powerful tool that can help us better understand and predict natural language, thereby improving our intelligence.<|im_end|>[EOS]

--- Turn 1 Performance (Generator::get_profiler) ---
  Init Time (us):           2120063
  Prompt Token Count:       33
  Time-to-First-Token (us): 32534
  Prompt TPS (tok/s):       1967.17
  Generated Token Count:    105
  Generate Time (us):       1077451
  Generate TPS (tok/s):     97.4522