Merge branch 'master' of ssh://192.168.35.82:2222/GenarrativeAI/Genarrative into feat/scene_v1
This commit is contained in:
@@ -14,6 +14,14 @@
|
||||
- 关联:相关文件、文档、提交或 Issue
|
||||
```
|
||||
|
||||
## Chat 生成预算字段不能按模型名猜测或失败后自动重放
|
||||
|
||||
- 现象:同一个 OpenAI-compatible Chat endpoint 调用 reasoning 模型时返回 `Unsupported parameter: max_tokens`;直接把全局请求字段改成 `max_completion_tokens` 后,旧兼容网关又可能拒绝新字段。
|
||||
- 原因:内部生成预算语义与上游 wire dialect 被混在一起。Chat 当前字段是 `max_completion_tokens`,旧兼容层仍只接受 `max_tokens`;Responses 和 Anthropic 又分别使用自己的字段。模型名、base URL 和 `OpenAiCompatible` 标签都不能证明 endpoint 能力,收到 `400` 后重发还可能重复计费。
|
||||
- 处理:在 `LlmConfig` 上显式声明 Chat token budget field capability;通用兼容配置默认 legacy,已验证的 VectorEngine 专用 client opt-in `max_completion_tokens`,每次只发送一个字段。内部 `max_output_tokens` 与 AGC 持久指纹键 `maxOutputTokens` 保持不变。
|
||||
- 验证:序列化测试分别断言 modern / legacy Chat 只出现选定字段,请求级 model override 不改变字段;Responses 继续只发 `max_output_tokens`,Anthropic 继续只发 `max_tokens`;AppState 测试断言 VectorEngine client 已显式启用 modern capability。
|
||||
- 关联:`server-rs/crates/platform-llm/src/lib.rs`、`server-rs/crates/api-server/src/state.rs`、`scripts/test-ve-llm.mjs`、Issue #143。
|
||||
|
||||
## Runtime 状态写失败不能发生在公开失败消息之前
|
||||
|
||||
- 现象:用户提交长任务后只看到运行失败或任务直接消失,聊天里一条有用消息都没有;另一些失败又同时出现 Runtime event 和 conversation 两条近似提示。
|
||||
|
||||
Reference in New Issue
Block a user