🔴 AMD Radeon Cloud 模型接入配置
AMD Radeon 云端免费模型 API 接入 DSH(DeepSeek Harness)与 OpenCode 的配置。
🗓️ 记录:2026-10-02 整理。配置格式参照 SenseNova模型接入配置 推导(同为只认 system 不认 developer 的 OpenAI 兼容端点),尚未实机验证。
🔌 接口
两套协议,baseURL 结尾差一个 /v1,这个坑很容易踩:
| 协议 | Endpoint | baseURL 填什么 | 认证 |
|---|---|---|---|
| OpenAI | POST /v1/chat/completions | https://developer.amd.com.cn/radeon/api/v1 | Authorization: Bearer $RADEON_API_KEY |
| Anthropic | POST /v1/messages | https://developer.amd.com.cn/radeon/api | x-api-key + anthropic-version: 2023-06-01 |
Anthropic SDK 自己拼 /v1/messages,所以走 SDK 时 baseURL 只到 /api。/v1 与 /api/v1 双路径都通,同一端点。
一个 Key 通吃所有共享模型,换模型只改 model 字段。
export RADEON_API_KEY=rc-... # rc- + 48 位十六进制,共 51 字符Key 在 Token Factory 首次打开页面时自动签发。账号需先过审,pending 状态会报 403 account_not_verified。一人一 Key,用邮箱别名开第二个账号会被禁用。
📋 模型
Token Factory 共 9 个免费模型:5 个 VLM(支持图片输入)、4 个纯文本。
| Model ID | 厂商 | 上下文 | 图片 | 推理档位 | 默认思考 |
|---|---|---|---|---|---|
DeepSeek-V4-Flash | DeepSeek | 1048576 | ❌ | none minimal low medium high xhigh max | 不思考 |
DeepSeek-V4-Flash-Vision-Exp | DeepSeek | 1048576 | ✅ | 同上 | 不思考 |
DeepSeek-V4.1-Flash | DeepSeek | 1048576 | ✅ | 同上 | 不思考 |
MiMo-V2.6-Flash | 小米 | 1048576 | ✅ 🎙️ | none + 常规档位 | 默认思考 |
Qwen3.8-Flash-Next | 阿里 | 262144 | ✅ | none low medium xhigh | 默认 xhigh |
Qwen3.8-27B | 阿里 | 131072 | ✅ | low medium xhigh | 默认 xhigh |
GLM-5.3-Flash | 智谱(Z.ai) | 262144 | ❌ | low medium high | 默认思考 |
MiniCPM5-2B | 面壁(OpenBMB) | 131072 | ❌ | 不适用 | 不思考 |
MinerU2.5-Pro | MinerU | — | — | — | 特殊,见下 |
MiMo-V2.6-Flash 是唯一支持音频输入的,也是唯一接受严格 json_schema 的。
MinerU2.5-Pro 不是对话模型——它是文档 OCR,走独立端点 POST /v1/ocr,返回 Markdown,按页计费,不响应 /v1/chat/completions。别配进 agent 的模型列表。
⚠️ 卡片标题 ≠ API model id——看卡片下方的「Model」字段
Token Factory 卡片标题显示的是 DEEPSEEK-V4-FLASH-0731,但点进去看模型详情面板,「Model」那一行显示的是 DeepSeek-V4-Flash,旁边还有完整的 Base URL 和复制按钮。
model 参数值要填的是详情面板「Model」字段的值(DeepSeek-V4-Flash),不是卡片标题(...-0731)。用 DeepSeek-V4-Flash-0731 会被拒:
404 {"detail":{"error":{"message":"Model DeepSeek-V4-Flash-0731 is not available","code":"model_not_found"}}}
官方文档写的 DeepSeek-V4-Flash 是对的,我就是看错了标题。
⚠️ /v1/models 不是标准 OpenAI 格式:没有 object 字段,条目也没有 owned_by、created。按 OpenAI schema 读模型列表的客户端会找不到这些字段;DSH 的「获取可用模型」探测大概率失败(报「既没有 data 数组也没有 models 对象」),直接手动填 ID 即可。
⚠️ 容量状态(2026-10-02 实测)
Token Factory 每个卡片上有实时负载,饱和的模型会直接拒请求:
| 模型 | 状态 |
|---|---|
DeepSeek-V4.1-Flash | 🔴 At capacity 100% |
DeepSeek-V4-Flash-Vision-Exp | 🔴 At capacity 100% |
Qwen3.8-Flash-Next | 🔴 At capacity 100% |
MiMo-V2.6-Flash | 🟡 Busy 84.4% |
GLM-5.3-Flash | 🟡 Busy 68.8% |
DeepSeek-V4-Flash | 🟡 Busy 62.2% |
9 个里当时有 3 个满载。 配置多个模型做故障切换比只配一个实用得多。容量随时变,用前看 Token Factory 或撞了 429/503 就换。
⚠️ system 角色的支持差异
这条决定 compat 怎么配:
| 模型 | system 位置 | 接受 developer? |
|---|---|---|
| DeepSeek 三款 | 任意位置,可多条 | ✅ |
Qwen3.8-Flash-Next | 必须恰好一条且在最前 | ❌ |
Qwen3.8-27B | 最多一条且在最前 | ❌ |
GLM-5.3-Flash / MiniCPM5-2B / MiMo-V2.6-Flash | 任意位置,可多条 | ❌ |
只有 DeepSeek 三款认 developer,其余全部直接拒:
Failed to deserialize the JSON body into the target type: messages[0]: unknown role: developer
新版 OpenAI SDK 用 developer 替代 system,所以默认行为会在非 DeepSeek 模型上全挂。
所有模型都接受的形状:index 0 处一条 role: "system"。
📡 接入状态
| agent | 协议 | 配置位置 | 状态 |
|---|---|---|---|
| OpenCode | OpenAI | ~/.config/opencode/opencode.json | 待验证 |
| harness | OpenAI | $DSH_HOME/settings.yaml | 待验证 |
⚙️ OpenCode
配置文件 ~/.config/opencode/opencode.json。Windows 是 C:\Users\你\.config\opencode\opencode.json——注意是 .config 不是 %AppData%,目录不存在要手动建。
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"radeon": {
"npm": "@ai-sdk/openai-compatible",
"name": "Radeon Cloud",
"options": {
"baseURL": "https://developer.amd.com.cn/radeon/api/v1",
"apiKey": "rc-xxx"
},
"models": {
"DeepSeek-V4-Flash": {
"name": "DeepSeek V4 Flash",
"modalities": { "input": ["text"], "output": ["text"] },
"limit": { "context": 1048576, "output": 65536 }
},
"DeepSeek-V4-Flash-Vision-Exp": {
"name": "DeepSeek V4 Flash Vision Exp",
"modalities": { "input": ["text", "image"], "output": ["text"] },
"limit": { "context": 1048576, "output": 65536 }
},
"DeepSeek-V4.1-Flash": {
"name": "DeepSeek V4.1 Flash",
"modalities": { "input": ["text", "image"], "output": ["text"] },
"limit": { "context": 1048576, "output": 65536 }
},
"MiMo-V2.6-Flash": {
"name": "MiMo V2.6 Flash",
"modalities": { "input": ["text", "image"], "output": ["text"] },
"limit": { "context": 1048576, "output": 65536 }
},
"Qwen3.8-Flash-Next": {
"name": "Qwen3.8 Flash Next",
"modalities": { "input": ["text", "image"], "output": ["text"] },
"limit": { "context": 262144, "output": 65536 }
},
"Qwen3.8-27B": {
"name": "Qwen3.8 27B",
"modalities": { "input": ["text", "image"], "output": ["text"] },
"limit": { "context": 131072, "output": 32768 }
},
"GLM-5.3-Flash": {
"name": "GLM 5.3 Flash",
"modalities": { "input": ["text"], "output": ["text"] },
"limit": { "context": 262144, "output": 65536 }
},
"MiniCPM5-2B": {
"name": "MiniCPM5 2B",
"modalities": { "input": ["text"], "output": ["text"] },
"limit": { "context": 131072, "output": 32768 }
}
}
}
}
}不要配 MinerU2.5-Pro——它是 OCR 模型,走 /v1/ocr 返回 Markdown,不响应 chat completions。
模型 ID 区分大小写,必须和 /v1/models 返回的 id 完全一致。
OpenCode 的推理档位选项只映射 OpenAI 标准档,AMD 的 xhigh / max 选不到,low / medium 正常。
⚠️ 待验证:OpenCode 是否会发
reasoning_effort。不发的话 Qwen3.8 和 GLM-5.3-Flash 会走端点默认档(xhigh),又慢又费额度——这两个模型建议先在 DSH 里显式压到low验证。
🤖 harness(DSH)
配置在 $DSH_HOME/settings.yaml(设置页顶部「打开配置文件」可直接打开)。密钥单独存 $DSH_HOME/.credentials.yaml;手写 YAML 用 apiKeyEnv 指向环境变量。模型变更下一次请求就生效,不用重启。
1️⃣ 方式一:模型页表单
设置 → 模型 → 添加自定义提供方:
| 字段 | 值 |
|---|---|
| Provider ID | radeon(小写,永久不可改) |
| 显示名称 | Radeon Cloud |
| API 地址 | https://developer.amd.com.cn/radeon/api/v1 |
| API 协议 | openai-completions |
| API 密钥 | rc-xxx |
| 模型 | 至少一个,填 ID / 显示名 / 上下文窗口 / 最大输出 |
2️⃣ 方式二:settings.yaml
表单不开放推理等级和兼容开关,必须写文件:
llm-pi-ai:
providers:
radeon:
apiKeyEnv: RADEON_API_KEY
api: openai-completions
baseURL: https://developer.amd.com.cn/radeon/api/v1
compat:
supportsDeveloperRole: false
maxTokensField: max_tokens
models:
- id: DeepSeek-V4-Flash
reasoningEfforts:
off: none
low: low
medium: medium
high: high
max: max
- id: DeepSeek-V4-Flash-Vision-Exp
input: [text, image]
reasoningEfforts:
off: none
low: low
medium: medium
high: high
max: max
- id: DeepSeek-V4.1-Flash
input: [text, image]
reasoningEfforts:
off: none
low: low
medium: medium
high: high
max: max
- id: MiMo-V2.6-Flash
input: [text, image]
reasoningEfforts:
off: none
low: low
medium: medium
high: high
- id: Qwen3.8-Flash-Next
input: [text, image]
reasoningEfforts:
off: none
low: low
medium: medium
high: xhigh
- id: Qwen3.8-27B
input: [text, image]
reasoningEfforts:
low: low
medium: medium
high: xhigh
- id: GLM-5.3-Flash
reasoningEfforts:
low: low
medium: medium
high: high
- id: MiniCPM5-2B四个关键字段为什么必须写:
compat.supportsDeveloperRole: false⚠️ 最重要的一个。只有 DeepSeek 三款认developer,Qwen / GLM / MiMo / MiniCPM 全部直接拒(messages[0]: unknown role: developer)。pi-ai 默认按 OpenAI 规矩发developer,不声明就会在多数模型上全挂。设false后统一发system,所有模型都接受。input: [text, image]:手动录入的模型一律按纯文本对待,不声明的话附图片在发送前就被拒。5 个模型支持图片:DeepSeek 的 Vision-Exp 与 V4.1、MiMo-V2.6、Qwen3.8 两款。compat.maxTokensField: max_tokens:pi-ai 默认写max_completion_tokens,AMD 端点认max_tokens。reasoningEfforts:手动录入的模型不声明等级就出不来推理菜单。off必须写none而不是留空——AMD 端点是省略参数就不思考,但 Qwen3.8 / GLM / MiMo 默认就在思考,要关掉得显式传none。
⚠️ 每模型的档位不能照抄。Qwen3.8-27B 传 high 返回 400(只收 low/medium/xhigh),GLM-5.3-Flash 传 xhigh 返回 422。上面示例里 Qwen3.8-Flash-Next 的 high 映射到 xhigh 就是为了避开这个。要一套配置跑所有模型,只有 low 和 medium 是通用的。
⚠️ max_tokens 要留够。Qwen3.8 两款在省略 reasoning_effort 时,思考和正式输出共享同一个 max_tokens 预算——按非思考模型设的额度可能换来一个空 content。要么调大 max_tokens,要么显式传 reasoning_effort: "low"。
所有 compat 开关都必须给值,冒号后留空(supportsDeveloperRole:)会被拒而不是忽略。
⚠️ 端点侧的坑
1. 参数被静默丢弃
请求会按白名单校验后逐字段重建,不在白名单里的不报错、直接消失:
stop seed logit_bias logprobs top_logprobs top_k min_p repetition_penalty
agent 依赖 stop 做输出截断会失效,且没有任何提示。
2. thinking 报 400
Anthropic 风格的 thinking: {"type":"enabled","budget_tokens":N} 在两个端点上都返回 400。要用 reasoning_effort(OpenAI 端点)或 output_config.effort(Anthropic 端点)。Claude Code 默认发 budget_tokens 形式,接这个端点会直接失败。
3. 省略 reasoning_effort ≠ 不思考
| 模型 | 省略时 |
|---|---|
| DeepSeek 系列 | 不思考 |
| Qwen3.8 系列 | 仍在 xhigh 思考(最慢最贵) |
| GLM-5.3-Flash | 仍在思考 |
| MiniCPM5-2B | 不思考 |
要确定性行为就显式传值。
4. 思考 token 的位置
思考文本在 choices[0].message.reasoning(不是 reasoning_content)。计费用 usage.completion_tokens_details.reasoning_tokens——Qwen3.8-27B 例外,它省略这个字段,思考 token 混在 usage.completion_tokens 里。别依赖顶层 usage.reasoning_tokens,只有部分模型给。
5. 超时
非流式请求最长等 10 分钟。长生成务必开 stream: true,既能看到进度也能保持连接。
⏱️ 限流对 agent 的实际影响
| 层 | 限制 | 典型值 |
|---|---|---|
| 平台准入 | 每 key RPM / 每 key 并发 | 30 / 8 |
| 网关计量 | 每账号 RPM | 20(60 秒滑动窗口) |
| 每账号消费上限 | 滚动周期,示例 $10/天 |
⚠️ 瓶颈是「每账号 20 RPM」,比「每 key 30 RPM」更严。agent 跑多步任务(工具调用循环)每步都是一次请求,很容易撞上。并发上限 8 也偏低,多路并行要先限到 8 以内。
撞限流返回 429 + Retry-After:每分钟类给 Retry-After: 60,并发类给 1。按 Retry-After 退避加指数抖动,别硬编码间隔。
查当前用量:
curl "https://radeon-global.anruicloud.com/api/profile/model-usage?include_recent=true" \
-H "Authorization: Bearer $RADEON_API_KEY"daily_cost_remaining_usd 是剩余额度,每日窗口在 Asia/Shanghai 00:00 重置。
🏠 平台侧速览
AMD 提供两条路,本文用的是第一条:
| Public Free Model APIs | GPU 实例 | |
|---|---|---|
| 要什么 | 只要 API Key | 建模板、启实例 |
| 计费 | 免费(计日额度) | 消耗 credits,实例开着就计费 |
| 拿到什么 | 直接调模型 API | 带 ROCm 的 Radeon GPU 机器,自己部署 vLLM |
要跑自有模型才需要实例。API Key 轮换用 POST /api/profile/api-token,签发新 key 并立即作废旧的——所有还在用旧 key 的服务立刻 401,所以轮换要和密钥下发在同一次变更里做完。
⚠️ 三条使用前提
1. 官方明说「不要用于生产」
无 SLA、无可用性承诺。模型可能毫无预告地增减,key 可能被重启或吊销,容量共享、繁忙时直接拒。
2. 会记录 prompt 和模型输出
对话内容滚动保留两天后删除,元数据留更久。只用于容量规划、调试、滥用检测和执行条款。官方明确不用于训练模型、不卖不分享给第三方自用,但原话也是「不希望 prompt 被存下来就别发」。
别往里塞密钥、客户数据、未公开业务逻辑。
3. 模型是上游原版,禁转售
跑的是公开发布的原始权重,不微调不蒸馏不换壳,许可证适用于你的用途。禁止向第三方暴露 API(哪怕套一层代理)、把服务说成自己的或别家的产品。
🕘 状态记录
- 2026-10-02:依官方文档(Radeon Cloud Docs,最后更新 2026-09-23)整理。平台侧信息来自官方文档,未实机调用验证。
- 2026-10-02:订正——上一版把真实 ID 误写成
DeepSeek-V4-Flash-0731(那是卡片标题)。卡片详情面板的「Model」字段是DeepSeek-V4-Flash,实测用-0731调/v1/chat/completions返回 404 +model_not_found。全表、两套配置、容量快照已同步改回。 - 2026-10-02:补全模型(5 VLM + 4 纯文本)。据官方 model reference 补齐各模型上下文(131072~1048576,八倍差距)与图片支持。
supportsDeveloperRole: false理由细化为「只有 DeepSeek 三款认developer」,并附各模型报错原文。 - 2026-10-02:待验证——所有 agent 侧配置均未实机跑过。优先级:①OpenCode 是否发
reasoning_effort(不发则 Qwen/GLM/MiMo 走默认xhigh)②/v1/models探测是否失败 ③compat三项是否被 DSH 正确接受。
🔗 相关
- SenseNova模型接入配置 — 同类端点的完整实测记录,
compat字段的来源 - DeepSeek-Harness桌面版 — DSH 本身
- OpenCode Zen免费模型 — 另一条免费模型路线,隐私策略对比
- Radeon Cloud 官方文档
- API 使用条款