adapt s100

This commit is contained in:
cyy_mac
2026-07-29 21:00:26 +08:00
parent 8e2f5d6de5
commit 8a44857314
6 changed files with 1881 additions and 0 deletions

View File

@@ -0,0 +1,207 @@
# S100 BPU 部署测试
这个目录是 S100 平台的隔离部署路径,不覆盖现有 X5 BPU 脚本。
当前默认模型:
```text
deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/policy_robotlab_26000_s100_int16_gemm.hbm
```
S100 量化参数:
- 原始模型:`deploy_45dim_rl_gym/policy_robotlab_26000.onnx`
- 历史长度RobotLab 10 帧
- 输入:`obs_4d [1, 1, 1, 450]`float32 featuremap
- 输出:`actions [1, 12, 1, 1]`
- `march``nash-e`
- Docker 镜像:`registry.d-robotics.cc/deliver/ai_toolchain_ubuntu_22_s100_s600_cpu:v3.7.0`
## 本机量化
在 Mac 的仓库根目录执行:
```bash
cd /Users/chenyouyuan/cyy_ws/deploy_go1_pro
bash deploy_45dim_rl_gym/bpu_quantization/quantize_policy_s100.sh
```
等价显式命令:
```bash
cd /Users/chenyouyuan/cyy_ws/deploy_go1_pro
bash deploy_45dim_rl_gym/bpu_quantization/quantize_policy_s100.sh \
--policy ../policy_robotlab_26000.onnx \
--round 26000 \
--name policy_robotlab_26000 \
--history-len 10 \
--samples 64 \
--min-samples 32 \
--log-prefix robotlab_go1_deploy \
--cal-tag robotlab \
--march nash-e
```
输出文件:
```text
deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/policy_robotlab_26000_s100_int16_gemm.hbm
```
## 同步到 S100
S100 板端地址:
```text
root@192.168.11.144
```
如果仓库已经通过 git 同步,直接在板端拉取即可。如果只同步产物,可以从 Mac 执行Docker 只在 Mac 上用于量化S100 板端不运行 Docker
```bash
scp \
deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/policy_robotlab_26000_s100_int16_gemm.hbm \
root@192.168.11.144:/root/go1_pro_deploy/deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/
```
同时确保校准输入存在,离线测速会用到:
```bash
scp \
deploy_45dim_rl_gym/bpu_quantization/calibration_data_26000_robotlab_fast64/00000.bin \
root@192.168.11.144:/root/go1_pro_deploy/deploy_45dim_rl_gym/bpu_quantization/calibration_data_26000_robotlab_fast64/
```
## 板端安装 hbm_runtime
```bash
ssh root@192.168.11.144
cd /usr/hobot/lib/hbm_runtime
./build.sh install
```
S100 使用官方 `hbm_runtime` Python 绑定加载 `.hbm`,不复用 X5 的
`/usr/include/dnn/hb_dnn.h` C++ wrapper。
## 离线推理测速
先用官方 `hrt_model_exec` 看模型信息:
```bash
cd /root/go1_pro_deploy
/usr/hobot/bin/hrt_model_exec model_info \
--model_file deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/policy_robotlab_26000_s100_int16_gemm.hbm
```
官方 `hrt_model_exec` 稳态测速:
```bash
cd /root/go1_pro_deploy
/usr/hobot/bin/hrt_model_exec perf \
--model_file deploy_45dim_rl_gym/bpu_quantization/mapper_output_26000_s100_gemm/policy_robotlab_26000_s100_int16_gemm.hbm \
--model_name policy_robotlab_26000_s100_int16_gemm \
--input_file deploy_45dim_rl_gym/bpu_quantization/calibration_data_26000_robotlab_fast64/00000.bin \
--frame_count 1000 \
--thread_num 1
```
当前板端 `root@192.168.11.144` 已验证:
```text
Average latency: 0.394 ms
FPS: 2442.456
```
部署脚本使用的 Python `hbm_runtime` wrapper
```bash
cd /root/go1_pro_deploy
PYTHONPATH=/root/go1_pro_sdk:/root/go1_pro_deploy \
python3 deploy_45dim_rl_gym/bpu_deploy_s100/test_bpu_policy.py \
--repeat 1000
```
当前板端结果:
```text
Backend: hbm_runtime_s100
Input: obs_4d (1, 1, 1, 450)
Output: actions (1, 12)
repeat=1000 avg_ms=0.733020
```
## 离线推理检查
这一步会连接 MCU 读取状态,但不会发送电机指令:
```bash
cd /root/go1_pro_deploy
PYTHONPATH=/root/go1_pro_sdk:/root/go1_pro_deploy \
python3 deploy_45dim_rl_gym/bpu_deploy_s100/deploy_go1_robotlab_bpu_s100_fastcpp.py \
--infer-check \
--log-dir logs \
--print-every 50 \
--max-steps 500
```
## 悬空状态机测试
先不要加 `--enable-rl`,确认 R2 只能推进到 `INFER_TEST`
```bash
cd /root/go1_pro_deploy
PYTHONPATH=/root/go1_pro_sdk:/root/go1_pro_deploy \
python3 deploy_45dim_rl_gym/bpu_deploy_s100/deploy_go1_robotlab_bpu_s100_fastcpp.py \
--kill-sport \
--log-dir logs \
--kp 28 --kd 0.7 \
--kp-cal 20 --kd-cal 1.0 \
--power-factor 7 \
--position-protect-limit 0.0 \
--action-clip 5.0 \
--action-trip-limit 8.0 \
--action-hard-trip-limit 16.0 \
--max-target-step 0.025 \
--max-roll-deg 35 \
--max-pitch-deg 35 \
--swap-vy-yaw \
--rc-vx-scale 0.3 \
--rc-vy-scale 0.3 \
--rc-wz-scale 0.6 \
--log-timing
```
## 实际 RL 启动
只有悬空测试正常后,再启用 RL
```bash
cd /root/go1_pro_deploy
PYTHONPATH=/root/go1_pro_sdk:/root/go1_pro_deploy \
python3 deploy_45dim_rl_gym/bpu_deploy_s100/deploy_go1_robotlab_bpu_s100_fastcpp.py \
--kill-sport \
--enable-rl \
--log-dir logs \
--kp 28 --kd 0.7 \
--kp-cal 20 --kd-cal 1.0 \
--power-factor 7 \
--position-protect-limit 0.0 \
--action-clip 5.0 \
--action-trip-limit 8.0 \
--action-hard-trip-limit 16.0 \
--max-target-step 0.025 \
--max-roll-deg 35 \
--max-pitch-deg 35 \
--swap-vy-yaw \
--rc-vx-scale 0.3 \
--rc-vy-scale 0.3 \
--rc-wz-scale 0.6 \
--log-timing
```
如果要临时指定其它 S100 `.hbm`
```bash
python3 deploy_45dim_rl_gym/bpu_deploy_s100/deploy_go1_robotlab_bpu_s100_fastcpp.py \
--bpu-model /absolute/path/to/model.hbm
```

View File

@@ -0,0 +1,77 @@
#!/usr/bin/env python3
"""S100 HBM policy runtime wrapper."""
from pathlib import Path
import numpy as np
class BpuInferLibPolicy:
"""S100 backend: official hbm_runtime Python binding."""
backend_name = "hbm_runtime_s100"
def __init__(self, model_path, priority=0, bpu_cores=(0,), cpp_lib=None):
del cpp_lib
self.model_path = Path(model_path).expanduser().resolve()
if not self.model_path.exists():
raise FileNotFoundError(f"S100 HBM model not found: {self.model_path}")
try:
from hbm_runtime import HB_HBMRuntime
except ImportError as exc:
raise RuntimeError(
"hbm_runtime is required on S100. Install it on the board with: "
"cd /usr/hobot/lib/hbm_runtime && ./build.sh install"
) from exc
self.priority = int(priority)
self.bpu_cores = tuple(int(core) for core in bpu_cores)
self.runtime = HB_HBMRuntime(str(self.model_path))
self.version = getattr(self.runtime, "version", "")
model_names = list(self.runtime.model_names)
if len(model_names) != 1:
raise RuntimeError(f"expected one model in {self.model_path}, got {model_names}")
self.model_name = model_names[0]
input_names = list(self.runtime.input_names[self.model_name])
output_names = list(self.runtime.output_names[self.model_name])
if len(input_names) != 1 or len(output_names) != 1:
raise RuntimeError(
f"expected 1 input and 1 output, got {input_names} / {output_names}"
)
self.input_name = input_names[0]
self.output_name = output_names[0]
self.input_shape = tuple(int(x) for x in self.runtime.input_shapes[self.model_name][self.input_name])
self.output_shape = tuple(int(x) for x in self.runtime.output_shapes[self.model_name][self.output_name])
self.input_size = int(np.prod(self.input_shape))
self.output_size = int(np.prod(self.output_shape))
print(f"[INFO] BPU model: {self.model_path}")
print(f"[INFO] Backend: {self.backend_name}")
print(f"[INFO] Runtime: {self.version}")
print(f"[INFO] Input : {self.input_name} {self.input_shape}")
print(f"[INFO] Output : {self.output_name} {self.output_shape}")
def close(self):
self.runtime = None
def __call__(self, flat_input):
arr = np.asarray(flat_input, dtype=np.float32)
if arr.size != self.input_size:
raise ValueError(f"S100 policy input has {arr.size} values, expected {self.input_size}")
input_tensor = np.ascontiguousarray(arr.reshape(self.input_shape), dtype=np.float32)
outputs = self.runtime.run(input_tensor)
action = np.asarray(outputs[self.model_name][self.output_name], dtype=np.float32).reshape(-1)
if action.size != self.output_size:
raise RuntimeError(
f"S100 policy output has {action.size} values, expected {self.output_size}"
)
if not np.all(np.isfinite(action)):
raise RuntimeError(f"S100 policy output is not finite: {action}")
return action.copy()
BpuInferLibPythonPolicy = BpuInferLibPolicy

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,55 @@
#!/usr/bin/env python3
"""Offline BPU policy smoke test for S100.
This does not connect to the robot. It loads a S100 BPU .hbm and one raw
float32 input file, then runs the hbm_runtime backend repeatedly.
"""
import argparse
import time
from pathlib import Path
import numpy as np
from bpu_policy import BpuInferLibPolicy
HERE = Path(__file__).parent.resolve()
DEFAULT_MODEL = (
HERE.parent / "bpu_quantization" / "mapper_output_26000_s100_gemm" /
"policy_robotlab_26000_s100_int16_gemm.hbm"
)
DEFAULT_INPUT = (
HERE.parent / "bpu_quantization" /
"calibration_data_26000_robotlab_fast64" / "00000.bin"
)
def main():
parser = argparse.ArgumentParser(description="Offline BPU policy smoke test")
parser.add_argument("--bpu-model", default=str(DEFAULT_MODEL))
parser.add_argument("--input-bin", default=str(DEFAULT_INPUT))
parser.add_argument("--repeat", type=int, default=1000)
args = parser.parse_args()
input_path = Path(args.input_bin).expanduser().resolve()
data = np.fromfile(input_path, dtype=np.float32)
policy = BpuInferLibPolicy(args.bpu_model)
if data.size != policy.input_size:
raise ValueError(
f"{input_path} has {data.size} float32 values, "
f"but model expects {policy.input_size}"
)
action = policy(data)
print("action", np.array2string(action, precision=6))
print("action_max_abs", float(np.max(np.abs(action))))
repeats = max(1, int(args.repeat))
t0 = time.perf_counter()
for _ in range(repeats):
policy(data)
elapsed_ms = (time.perf_counter() - t0) * 1000.0
print(f"repeat={repeats} avg_ms={elapsed_ms / repeats:.6f}")
if __name__ == "__main__":
main()