|
| 1 | +# MiniCPM-V-4.6 |
| 2 | + |
| 3 | +本工程实现BM1684X/BM1688部署多模态大模型[MiniCPM-V-4.6](https://huggingface.co/openbmb/MiniCPM-V-4.6)。通过[TPU-MLIR](https://github.com/sophgo/tpu-mlir)编译器将模型转换成bmodel,并将其部署到PCIE环境,或者SoC环境。 |
| 4 | + |
| 5 | +该模型支持图片和视频的识别,有python版本的demo。 |
| 6 | + |
| 7 | +本文包括如何编译bmodel,和如何在BM1684X/BM1688环境运行bmodel。如何编译bmodel环节可以省去,直接用以下链接下载: |
| 8 | + |
| 9 | +``` shell |
| 10 | +# =============== 1684x ===================== |
| 11 | +python3 -m dfss --url=open@sophgo.com:/ext_model_information/LLM/LLM-TPU/minicpm-v-4.6_bf16_seq2048_bm1684x_1dev_dynamic_20260630_105643.bmodel |
| 12 | +python3 -m dfss --url=open@sophgo.com:/ext_model_information/LLM/LLM-TPU/minicpm-v-4.6_bf16_seq2048_bm1688_2core_dynamic_20260630_155028.bmodel |
| 13 | +``` |
| 14 | + |
| 15 | +## 模型架构 |
| 16 | + |
| 17 | +MiniCPM-V-4.6 参数量 1.3B,由视觉编码器和文本模型两部分组成: |
| 18 | + |
| 19 | +- **文本模型**:与 Qwen3.5 相同架构 |
| 20 | +- **视觉编码器**:SigLIP ViT + Merger 降采样,支持两种模式: |
| 21 | + - `16x`(默认):两级 2×2 合并,共 16 倍降采样,token 数少,推理快 |
| 22 | + - `4x`:一级 2×2 合并,4 倍降采样,保留更多视觉细节 |
| 23 | + |
| 24 | +## 编译bmodel |
| 25 | + |
| 26 | +此处介绍如何将模型编译成bmodel。 |
| 27 | + |
| 28 | +#### 1. 下载模型 |
| 29 | + |
| 30 | +``` shell |
| 31 | +git clone https://huggingface.co/openbmb/MiniCPM-V-4.6 |
| 32 | +``` |
| 33 | + |
| 34 | +#### 2. 下载docker,启动容器 |
| 35 | + |
| 36 | +``` shell |
| 37 | +docker pull sophgo/tpuc_dev:latest |
| 38 | + |
| 39 | +# myname1234 is just an example, you can set your own name |
| 40 | +docker run --privileged --name myname1234 -v $PWD:/workspace -it sophgo/tpuc_dev:latest |
| 41 | +``` |
| 42 | +后文假定环境都在docker的`/workspace`目录。 |
| 43 | + |
| 44 | +#### 3. 下载`TPU-MLIR`代码并编译 |
| 45 | + |
| 46 | +``` shell |
| 47 | +cd /workspace |
| 48 | +git clone git@github.com:sophgo/tpu-mlir.git |
| 49 | +cd tpu-mlir |
| 50 | +source ./envsetup.sh #激活环境变量 |
| 51 | +./build.sh #编译mlir |
| 52 | +``` |
| 53 | + |
| 54 | +#### 4. 编译模型生成bmodel |
| 55 | + |
| 56 | +``` shell |
| 57 | +# 这里max_input_length指定最大输入长度,如果不指定则为-s指定的长度 |
| 58 | +llm_convert.py -m /workspace/MiniCPM-V-4.6 -s 2048 --max_input_length 1024 -q bf16 -c bm1684x -o minicpm_v4_6 --max_pixels 448,448 |
| 59 | +``` |
| 60 | +编译完成后,在指定目录生成`minicpm-v-4.6-xxx.bmodel`和`config`。 |
| 61 | + |
| 62 | +## 编译与运行程序(python) |
| 63 | + |
| 64 | +### 1. 环境准备 |
| 65 | + |
| 66 | +需要 python3.10 环境。如果不满足,参考[此文档](https://github.com/sophgo/sophon-demo/blob/release/docs/FAQ.md#13-se7%E5%AE%89%E8%A3%85python310)安装。 |
| 67 | + |
| 68 | +``` shell |
| 69 | +sudo apt-get update |
| 70 | +sudo apt-get install pybind11-dev |
| 71 | + |
| 72 | +pip3 install torch==2.6.0 torchvision==0.21.0 transformers==5.7.0 |
| 73 | +``` |
| 74 | + |
| 75 | +### 2. 编译库文件 |
| 76 | + |
| 77 | +编译C++库文件,生成`chat.cpython*.so`: |
| 78 | + |
| 79 | +``` shell |
| 80 | +cd python_demo |
| 81 | +mkdir build |
| 82 | +cd build && cmake .. && make && cp *cpython* .. && cd .. |
| 83 | +``` |
| 84 | + |
| 85 | +### 3. 运行 |
| 86 | + |
| 87 | +``` shell |
| 88 | +# 交互模式 |
| 89 | +python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config |
| 90 | + |
| 91 | +# 单次推理模式(图片) |
| 92 | +python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config \ |
| 93 | + --prompt "描述这张图片" --media_path test.jpg |
| 94 | + |
| 95 | +# 单次推理模式(视频) |
| 96 | +python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config \ |
| 97 | + --prompt "描述视频中发生了什么" --media_path test.mp4 |
| 98 | +``` |
| 99 | + |
| 100 | +### CLI 参数 |
| 101 | + |
| 102 | +| 参数 | 默认值 | 说明 | |
| 103 | +|------|--------|------| |
| 104 | +| `-m, --model_path` | 必填 | bmodel 文件路径 | |
| 105 | +| `-c, --config_path` | `../config` | processor 配置文件目录 | |
| 106 | +| `-d, --devid` | `0` | 设备 ID | |
| 107 | +| `--downsample_mode` | `16x` | ViT 降采样模式,可选 `4x` / `16x` | |
| 108 | +| `--max_slice_nums` | None | 图片最大切片数。不指定时图片默认36,视频默认1;指定后统一使用 | |
| 109 | +| `--max_num_frames` | `16` | 视频最大采样帧数 | |
| 110 | +| `-p, --prompt` | None | 指定后进入单次推理模式 | |
| 111 | +| `-t, --prompt_file` | None | 从文件加载 prompt | |
| 112 | +| `--media_path` | 空 | 图片或视频路径,配合 `--prompt` 使用 | |
| 113 | + |
| 114 | +## 进阶应用 |
| 115 | + |
| 116 | +参考[Qwen3.5 README](../Qwen3_5/README.md)。 |
| 117 | + |
| 118 | +## 常见问题 |
| 119 | + |
| 120 | +#### 一张图片占多少 Token ? |
| 121 | + |
| 122 | +计算公式:$token数 = \frac{h_{patches} \times w_{patches}}{merge\_size^2}$ |
| 123 | + |
| 124 | +其中 $h_{patches} = \frac{height}{14}$,$w_{patches} = \frac{width}{14}$。 |
| 125 | + |
| 126 | +以 448×448 图片为例: |
| 127 | +- patches = 32 × 32 = 1024 |
| 128 | +- 16x 模式:1024 / 16 = **64 tokens** |
| 129 | +- 4x 模式:1024 / 4 = **256 tokens** |
| 130 | + |
| 131 | +高分辨率图片会被切片处理,`max_slice_nums` 控制最大切片数。 |
| 132 | + |
| 133 | +#### 视频占多少 Token ? |
| 134 | + |
| 135 | +视频每帧独立处理,`max_slice_nums=1`(不切片)。 |
| 136 | + |
| 137 | +以 16x 模式、448×448 每帧为例,每帧 64 tokens。 |
| 138 | + |
| 139 | +16 帧视频:$64 × 16 = 1024$ tokens |
| 140 | + |
| 141 | +#### 视频最多支持多少帧 ? |
| 142 | + |
| 143 | +| 降采样模式 | 每帧 Token | 最大帧数 (SEQLEN=2048) | |
| 144 | +|-----------|-----------|----------------------| |
| 145 | +| 16x | ~64 | ~25 | |
| 146 | +| 4x | ~256 | ~6 | |
| 147 | + |
| 148 | +实际可用帧数还需减去 prompt 和输出占用的 tokens。 |
0 commit comments