修复试玩回包遗漏双端玩法失败证据(#548) (#598)
Project CI / AI game creator shell Rust crates (push) Successful in 2m58s
Project CI / Backend tests (push) Failing after 4m25s
Project CI / AI game creator shell Rust lane 2/2 (push) Successful in 5m34s
Project CI / AI game creator shell Rust lane 1/2 (push) Successful in 5m44s
Project CI / Frontend tests (push) Successful in 2m50s
Project CI / AI game creator shell web tests (push) Successful in 2m45s
Project CI / Repository checks (push) Failing after 4m7s
Project CI / Native shell tests (push) Successful in 6m49s
Project CI / AI game creator shell Rust crates (push) Successful in 2m58s
Project CI / Backend tests (push) Failing after 4m25s
Project CI / AI game creator shell Rust lane 2/2 (push) Successful in 5m34s
Project CI / AI game creator shell Rust lane 1/2 (push) Successful in 5m44s
Project CI / Frontend tests (push) Successful in 2m50s
Project CI / AI game creator shell web tests (push) Successful in 2m45s
Project CI / Repository checks (push) Failing after 4m7s
Project CI / Native shell tests (push) Successful in 6m49s
试玩报告已记录玩法阻断,但工具回包只展示视觉结果,Agent 无法从“画面正常、整体失败”定位原因;返回的私有报告路径也不能通过通用文件工具读取。 本次增加逐视口 gameplayResults,返回有界诊断、断言和初末状态,并从同一份投影生成摘要。明确区分视觉检查与玩法检查、未通过与未执行;保留现有通过判定、验证预算和私有目录保护。同步更新证据合同、提示词与 skill pack。 验证: - Rust 定向测试、提示词构建与源码边界检查通过。 - 最终 MCP 成功/失败回包均检查逐视口状态、原因、断言及截图。 - 显式运行 4 个真实浏览器用例:双端成功、禁用控件、真实点击后 phase 未进入 playing、仅移动端失败;对照持久报告核对回包与摘要。 - rustfmt、编码、文档索引、skill pack、git diff --check 通过。 - 未运行全量测试或真实模型端到端验证。 Closes #548 Reviewed-on: #598
This commit was merged in pull request #598.
This commit is contained in:
@@ -195,6 +195,8 @@ UI 编辑器的“分析参考图”步骤、Rust 命令 `suggest_ui_design_sema
|
||||
### 分层验证与预算
|
||||
|
||||
- 复用客户端浏览器与现有固定玩法场景。视觉检查采集双端画面/布局/资源/诊断;玩法检查分别在 desktop/mobile 执行明确的固定场景和真实输入,按视口保存结果。旧报告缺少移动端玩法结果时保持未知,不补通过。报告必须声明检查层级;视觉通过不能宣称玩法通过,固定场景通过也不能宣称覆盖未执行的完整关卡。
|
||||
- 浏览器工具回执在 `mode=gameplay` 时新增 `gameplayResults` 双端有界投影,每个视口包含 `viewport`、`passed`、`diagnostics`、`assertions` 及 `initialPhase` / `initialSequence` / `initialLevel`、`finalPhase` / `finalSequence` / `finalLevel`。`assertions` 仅列当前固定场景的断言名称和通过状态;每端最多公开 4 条诊断,每条最多 512 字符。缺少某端结果时仍返回该视口 `passed=false`,诊断说明证据缺失,断言为空且六个状态字段为 `null`;断言的 `passed=false` 可包含未执行,不能据此断言该断言已独立执行且失败。`mode=visual` 的 `gameplayResults` 为空,并在摘要中明确玩法未执行。摘要从同一投影给出每端首条诊断和首个未通过断言,不能与结构化结果相矛盾。工具回执不得展开完整宿主报告;`reportPath` 只作宿主证据引用。宿主报告结构、验收门禁、持久化预算和私有路径保护保持原合同,旧回执不回填新字段。
|
||||
- 回执验收覆盖 start 禁用、点击后 phase 不符合和玩法成功;逐视口核对持久报告与摘要中的状态、首条诊断和首个未通过断言,并检查成功、失败两条最终 MCP 回包路径。
|
||||
- 输入或碰撞改变先做定点玩法检查;纯图像/颜色变化做视觉检查;首次交付和影响闭环的修改做所需玩法验证。新增失败或相关代码变化才重跑对应层,不因改说明文字重复完整验证。
|
||||
- 内置试玩和客户端托管的外部 Node/npm 验证共用当前 clientTurnId 的持久化预算。客户端分配递增执行序号,模型提供的旧 attempt 仅作兼容输入,不能减少计数或重置预算;同一轮错误反馈、工具切换和进程重启均不能刷新已消费次数。
|
||||
- 新增本地 validation.maxRuns(默认 3,正整数)独立于 llm.maxRetries;显式配置原样使用,不按角色或运行模式改写。超限直接返回已用/上限和最近证据,停止新的验证。预检与正常构建不计作重复试玩。
|
||||
|
||||
Reference in New Issue
Block a user