Skip to content

Commit e281580

Browse files
authored
Merge pull request #159 from sophgo/xin.zhang
feat: add MiniCPM-V-4.6 multimodal model deployment
2 parents fbb5526 + bea5f3e commit e281580

16 files changed

Lines changed: 1241171 additions & 1 deletion

‎README.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -31,7 +31,8 @@
3131

3232
| 日期 | 更新内容 |
3333
| :--- | :--- |
34-
| 🔥 **2026.05.21** | **Gemma4** 已支持 BM1684X / BM1688,Python Demo,支持图片 / 视频 / 音频 → [查看](./models/Gemma4/) |
34+
| 🔥 **2026.06.30** | **MiniCPM-V-4.6** 已支持 BM1684X / BM1688,Python Demo,支持图片与视频 → [查看](./models/MiniCPMV4_6/) |
35+
| **2026.05.21** | **Gemma4** 已支持 BM1684X / BM1688,Python Demo,支持图片 / 视频 / 音频 → [查看](./models/Gemma4/) |
3536
| **2026.04.15** | **Qwen3.5** 已支持 BM1684X / BM1688,提供 Python 与 C++ Demo,支持图片与视频 → [查看](./models/Qwen3_5/) |
3637
| **2025.10.15** | **Qwen3-VL** 已支持 BM1684X / BM1688,Python / C++ Demo,支持图片与视频 → [查看](./models/Qwen3_VL/) |
3738
| **2025.05.22** | **InternVL3** 已支持 BM1684X / BM1688,支持图片与视频 → [查看](./models/InternVL3/) |
@@ -159,6 +160,7 @@ cd LLM-TPU
159160
[Llama3_2-Vision](./models/Llama3_2-Vision) ·
160161
[MiniCPM-V-2_6](./models/MiniCPM-V-2_6) ·
161162
[MiniCPMV4](./models/MiniCPMV4) ·
163+
[MiniCPMV4_6](./models/MiniCPMV4_6) ·
162164
[Molmo](./models/Molmo) ·
163165
[NVILA](./models/NVILA) ·
164166
[Qwen2_5_Omni](./models/Qwen2_5_Omni) ·

‎models/MiniCPMV4_6/README.md‎

Lines changed: 148 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,148 @@
1+
# MiniCPM-V-4.6
2+
3+
本工程实现BM1684X/BM1688部署多模态大模型[MiniCPM-V-4.6](https://huggingface.co/openbmb/MiniCPM-V-4.6)。通过[TPU-MLIR](https://github.com/sophgo/tpu-mlir)编译器将模型转换成bmodel,并将其部署到PCIE环境,或者SoC环境。
4+
5+
该模型支持图片和视频的识别,有python版本的demo。
6+
7+
本文包括如何编译bmodel,和如何在BM1684X/BM1688环境运行bmodel。如何编译bmodel环节可以省去,直接用以下链接下载:
8+
9+
``` shell
10+
# =============== 1684x =====================
11+
python3 -m dfss --url=open@sophgo.com:/ext_model_information/LLM/LLM-TPU/minicpm-v-4.6_bf16_seq2048_bm1684x_1dev_dynamic_20260630_105643.bmodel
12+
python3 -m dfss --url=open@sophgo.com:/ext_model_information/LLM/LLM-TPU/minicpm-v-4.6_bf16_seq2048_bm1688_2core_dynamic_20260630_155028.bmodel
13+
```
14+
15+
## 模型架构
16+
17+
MiniCPM-V-4.6 参数量 1.3B,由视觉编码器和文本模型两部分组成:
18+
19+
- **文本模型**:与 Qwen3.5 相同架构
20+
- **视觉编码器**:SigLIP ViT + Merger 降采样,支持两种模式:
21+
- `16x`(默认):两级 2×2 合并,共 16 倍降采样,token 数少,推理快
22+
- `4x`:一级 2×2 合并,4 倍降采样,保留更多视觉细节
23+
24+
## 编译bmodel
25+
26+
此处介绍如何将模型编译成bmodel。
27+
28+
#### 1. 下载模型
29+
30+
``` shell
31+
git clone https://huggingface.co/openbmb/MiniCPM-V-4.6
32+
```
33+
34+
#### 2. 下载docker,启动容器
35+
36+
``` shell
37+
docker pull sophgo/tpuc_dev:latest
38+
39+
# myname1234 is just an example, you can set your own name
40+
docker run --privileged --name myname1234 -v $PWD:/workspace -it sophgo/tpuc_dev:latest
41+
```
42+
后文假定环境都在docker的`/workspace`目录。
43+
44+
#### 3. 下载`TPU-MLIR`代码并编译
45+
46+
``` shell
47+
cd /workspace
48+
git clone git@github.com:sophgo/tpu-mlir.git
49+
cd tpu-mlir
50+
source ./envsetup.sh #激活环境变量
51+
./build.sh #编译mlir
52+
```
53+
54+
#### 4. 编译模型生成bmodel
55+
56+
``` shell
57+
# 这里max_input_length指定最大输入长度,如果不指定则为-s指定的长度
58+
llm_convert.py -m /workspace/MiniCPM-V-4.6 -s 2048 --max_input_length 1024 -q bf16 -c bm1684x -o minicpm_v4_6 --max_pixels 448,448
59+
```
60+
编译完成后,在指定目录生成`minicpm-v-4.6-xxx.bmodel`和`config`。
61+
62+
## 编译与运行程序(python)
63+
64+
### 1. 环境准备
65+
66+
需要 python3.10 环境。如果不满足,参考[此文档](https://github.com/sophgo/sophon-demo/blob/release/docs/FAQ.md#13-se7%E5%AE%89%E8%A3%85python310)安装。
67+
68+
``` shell
69+
sudo apt-get update
70+
sudo apt-get install pybind11-dev
71+
72+
pip3 install torch==2.6.0 torchvision==0.21.0 transformers==5.7.0
73+
```
74+
75+
### 2. 编译库文件
76+
77+
编译C++库文件,生成`chat.cpython*.so`:
78+
79+
``` shell
80+
cd python_demo
81+
mkdir build
82+
cd build && cmake .. && make && cp *cpython* .. && cd ..
83+
```
84+
85+
### 3. 运行
86+
87+
``` shell
88+
# 交互模式
89+
python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config
90+
91+
# 单次推理模式(图片)
92+
python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config \
93+
--prompt "描述这张图片" --media_path test.jpg
94+
95+
# 单次推理模式(视频)
96+
python3 pipeline.py -m minicpm-v-4.6.bmodel -c ../config \
97+
--prompt "描述视频中发生了什么" --media_path test.mp4
98+
```
99+
100+
### CLI 参数
101+
102+
| 参数 | 默认值 | 说明 |
103+
|------|--------|------|
104+
| `-m, --model_path` | 必填 | bmodel 文件路径 |
105+
| `-c, --config_path` | `../config` | processor 配置文件目录 |
106+
| `-d, --devid` | `0` | 设备 ID |
107+
| `--downsample_mode` | `16x` | ViT 降采样模式,可选 `4x` / `16x` |
108+
| `--max_slice_nums` | None | 图片最大切片数。不指定时图片默认36,视频默认1;指定后统一使用 |
109+
| `--max_num_frames` | `16` | 视频最大采样帧数 |
110+
| `-p, --prompt` | None | 指定后进入单次推理模式 |
111+
| `-t, --prompt_file` | None | 从文件加载 prompt |
112+
| `--media_path` | 空 | 图片或视频路径,配合 `--prompt` 使用 |
113+
114+
## 进阶应用
115+
116+
参考[Qwen3.5 README](../Qwen3_5/README.md)。
117+
118+
## 常见问题
119+
120+
#### 一张图片占多少 Token ?
121+
122+
计算公式:$token数 = \frac{h_{patches} \times w_{patches}}{merge\_size^2}$
123+
124+
其中 $h_{patches} = \frac{height}{14}$,$w_{patches} = \frac{width}{14}$。
125+
126+
以 448×448 图片为例:
127+
- patches = 32 × 32 = 1024
128+
- 16x 模式:1024 / 16 = **64 tokens**
129+
- 4x 模式:1024 / 4 = **256 tokens**
130+
131+
高分辨率图片会被切片处理,`max_slice_nums` 控制最大切片数。
132+
133+
#### 视频占多少 Token ?
134+
135+
视频每帧独立处理,`max_slice_nums=1`(不切片)。
136+
137+
以 16x 模式、448×448 每帧为例,每帧 64 tokens。
138+
139+
16 帧视频:$64 × 16 = 1024$ tokens
140+
141+
#### 视频最多支持多少帧 ?
142+
143+
| 降采样模式 | 每帧 Token | 最大帧数 (SEQLEN=2048) |
144+
|-----------|-----------|----------------------|
145+
| 16x | ~64 | ~25 |
146+
| 4x | ~256 | ~6 |
147+
148+
实际可用帧数还需减去 prompt 和输出占用的 tokens。

0 commit comments

Comments
 (0)