收敛 Agent 工具计划格式修复
过滤 Provider 推理内容并受限归一化非最终工具旁白 为原生工具协议增加递归重复键校验和稳定错误分类 补充 repair 统计、失败证据与零泄漏真实验收门禁 同步 Agent Runtime 技术方案和项目决策记录
This commit is contained in:
@@ -1179,6 +1179,21 @@ V1.31 证明真实 Provider 可以自主形成 static + isolated 混合协作,
|
||||
|
||||
成功报告共记录 178 个 task snapshot、326 个 event、556 个 Agent DB record、30 个 action execution 和 38 个 receipt;重复 delivery/group/instance/result/join/claim/message/action/receipt/Provider lifecycle 与 pending/batch/finalization/confirmation sidecar 均为 0,私密正文、Provider payload、API Key、项目路径、正式配置路径和最终报告泄漏均为 0。`turn.report=settled` 且 reconciliation Agent 为 0。49 次 native tool plan 中发生 30 次格式修复,未破坏动作幂等与最终结果,但说明真实链路仍有明显延迟和 Provider 调用成本,后续应单独收敛工具合同表达和 repair 频率。
|
||||
|
||||
## V1.33 原生工具计划 repair 收敛与分类
|
||||
|
||||
V1.32 的真实基线是 `49` 次成功 native tool plan 对应 `30` 次格式修复。V1.33 不放宽动作、完成、权限或恢复门禁,只收敛 OpenAI-compatible Provider 的工具响应兼容边界,并让每次 repair 可以按稳定类别计数。
|
||||
|
||||
- `platform-llm` 的 Chat Completions 与 Responses content parts 只排除明确标记为 `reasoning / reasoning_content / analysis / thinking` 的内部推理 part;普通 `text / output_text` 继续进入可见正文,独立 `reasoning_content` 不提升为正文。不能因为响应同时包含 tool calls 就笼统丢弃 content。
|
||||
- Agent 原生工具解析可以移除完整、大小写不敏感且嵌套闭合的 `<think>...</think>` 块。对于不含 `respond_to_user` 和旧 `submit_agent_tool_plan` 的 native planning 响应,剩余普通文本只作为不具执行权的 `planner-commentary` 丢弃,以 function calls 作为权威动作;最终用户回复、legacy wrapper、未闭合 / 错配 thinking 标签仍按 `response-shape` 失败关闭。成功归一化的公共审计只保存固定 `complete-think-block / planner-commentary` kind、数量、原文本字符数和 SHA-256,不保存正文。
|
||||
- 工具计划错误使用固定类别 `response-shape / call-identity / unknown-function / arguments-json / arguments-schema / batch-constraint / plan-semantics / catalog-binding`。JSON 先递归遍历所有 object key,再做 schema 解析,既稳定区分语法与 schema,也拒绝顶层及任意嵌套 input 的重复字段。repair prompt 仍可在当前瞬时私有请求中使用过滤后的错误细节,Agent DB 与 E2E 报告只保存类别、计数和既有哈希;`catalog-binding` 属于本地目录冲突,必须直接失败,不能重问 Provider,也不能在成功 E2E 报告中出现非零计数。
|
||||
- OpenAI-compatible planning prompt 必须把原生 function schema 作为参数事实源。legacy text JSON schema 明确只供不支持 function tools 的 Provider 使用;所有动作函数参数统一为 `{"reason":"...","input":{...}}`,工具输入示例只描述 `input` 字段,禁止把 input 属性扁平到 arguments 顶层。
|
||||
- E2E 证据固定输出发生过 repair 的 loop 数、第二次 repair 数和按上述固定类别补零后的直方图;分类总和必须等于 repair 总数。完整报告和失败 partial report 使用同一聚合器,禁止输出单条错误、错误正文、preview、arguments、Agent/run/loop 身份或动态类别。
|
||||
- 确定性验收覆盖 standalone reasoning、reasoning content part、可见 content、完整 / 嵌套 / 未闭合 thinking block、非最终 commentary、最终回复正文冲突、重复 JSON 字段、八类错误、repair 审计零正文和报告聚合闭合。真实验收继续复用 V1.32 `supervisor-swarm-collaboration-policy-mixed-recovery`,在相同正式路由和隔离 AppData 口径下对比 `49 / 30` 基线,并同时检查总耗时、唯一 lifecycle、零重复、零泄漏和现场清理。
|
||||
|
||||
2026-07-18 第一轮诊断在旧保守正文规则下形成 `35` 次成功计划与 `23` 次 `response-shape` repair,比例未比 V1.32 下降;同时新 E2E 白名单遗漏 Agent DB 固有的 `schemaVersion / updatedAt`,把 58 条安全审计误判为 payload leak,因此该轮 **FAIL** 且不作为完成证据。第二次尝试在 2 次计划、0 repair 时因 Provider 给出的 isolated child write scope 不满足业务 fixture 提前停止,同样不作为完成证据。
|
||||
|
||||
最终代码对应的正式 `gpt-5.5 / openai_chat / high` 独立轮 **PASS**,总耗时 `631.6s`(约 10 分 32 秒)。46/46 次成功计划全部使用 `native_runtime_tools`,格式 repair 为 `0`,八类 repair 直方图全为 `0`;Provider lifecycle 从 V1.32 基线的 86 降为 54,started / terminal 均为 54,其中 53 completed、1 次瞬态失败通过新的 request identity 显式重试恢复,wrapper/text fallback 和协议审计 payload leak 均为 0。相同父 Session/run 完成 2 个 static delegate、1 个三 child isolated all-join、1 次业务 delivery repair、两类 Provider 真并行、pidfd Runner 强杀恢复、宿主验证和唯一 Supervisor assistant;重复 delivery/group/instance/result/join/claim/message/action/receipt/lifecycle、残留 sidecar、私密正文、API Key、项目 / 正式配置路径及报告泄漏均为 0,隔离 AppData 与 disposable 项目完整清理。
|
||||
|
||||
## 验收命令
|
||||
|
||||
- `cargo test --manifest-path apps/ai-game-creator-shell/src-tauri/Cargo.toml structured_plan_ -- --nocapture`
|
||||
|
||||
Reference in New Issue
Block a user