完善单Agent持久计划与恢复验收

新增结构化计划更新、单调修订、完成门禁和失败进度保留
升级上下文与最终回复恢复协议并绑定规划仓库指纹
收紧模型正文公共审计和CLI本地存储路径输出
补齐开发界面计划展示、Supervisor紧凑摘要与乱序竞态合并
强化真实Provider验收器并记录TLS外部门禁未通过
补充Rust前端回归及长期技术文档
This commit is contained in:
AIGameCreator App
2026-07-15 04:03:16 +08:00
parent 5ab664b22e
commit f5d2b3786a
12 changed files with 6208 additions and 316 deletions
File diff suppressed because it is too large Load Diff
File diff suppressed because one or more lines are too long
+60 -10
View File
@@ -446,6 +446,32 @@ pub(crate) fn parse_cli_command(args: &[String]) -> Result<Option<CliCommand>, S
}))
}
fn strip_agent_runtime_cli_private_paths(value: &mut serde_json::Value) {
match value {
serde_json::Value::Array(values) => {
for value in values {
strip_agent_runtime_cli_private_paths(value);
}
}
serde_json::Value::Object(object) => {
object.remove("sessionPath");
object.remove("eventPath");
object.remove("taskPath");
for value in object.values_mut() {
strip_agent_runtime_cli_private_paths(value);
}
}
_ => {}
}
}
fn serialize_agent_runtime_cli_payload<T: serde::Serialize>(payload: &T) -> Result<String, String> {
let mut value = serde_json::to_value(payload)
.map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))?;
strip_agent_runtime_cli_private_paths(&mut value);
serde_json::to_string(&value).map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))
}
pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
match command {
CliCommand::LlmStatus => {
@@ -595,8 +621,7 @@ pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
println!("requestedRunId={run_id}");
println!(
"runtimeJson={}",
serde_json::to_string(&runtime)
.map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))?
serialize_agent_runtime_cli_payload(&runtime)?
);
Ok(())
}
@@ -607,8 +632,7 @@ pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
let runtime = read_game_creator_agent_runtime_at(&project_path, &agent_id)?;
println!(
"runtimeJson={}",
serde_json::to_string(&runtime)
.map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))?
serialize_agent_runtime_cli_payload(&runtime)?
);
Ok(())
}
@@ -630,8 +654,7 @@ pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
println!("agent.confirm.accepted");
println!(
"runtimeJson={}",
serde_json::to_string(&runtime)
.map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))?
serialize_agent_runtime_cli_payload(&runtime)?
);
Ok(())
}
@@ -665,8 +688,7 @@ pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
println!("steerId={steer_id}");
println!(
"steerJson={}",
serde_json::to_string(&result)
.map_err(|error| format!("序列化 Agent steer 结果失败:{error}"))?
serialize_agent_runtime_cli_payload(&result)?
);
Ok(())
}
@@ -677,8 +699,7 @@ pub(crate) fn run_cli_command(command: CliCommand) -> Result<(), String> {
println!("agent.resume.accepted");
println!(
"runtimesJson={}",
serde_json::to_string(&runtimes)
.map_err(|error| format!("序列化 Agent Runtime 状态失败:{error}"))?
serialize_agent_runtime_cli_payload(&runtimes)?
);
Ok(())
}
@@ -763,6 +784,35 @@ mod tests {
use super::*;
use std::io::Cursor;
#[test]
fn runtime_cli_payload_omits_private_storage_paths_recursively() {
let payload = serde_json::json!({
"sessionPath": "/tmp/private/session.json",
"runtime": {
"eventPath": "/tmp/private/events.jsonl",
"state": {
"sessionId": "session-7",
"runId": "run-9"
}
},
"runtimes": [{
"taskPath": "/tmp/private/tasks.jsonl",
"state": { "status": "running" }
}]
});
let encoded = serialize_agent_runtime_cli_payload(&payload)
.expect("serialize redacted Runtime CLI payload");
assert!(!encoded.contains("/tmp/private"));
let parsed = serde_json::from_str::<serde_json::Value>(&encoded)
.expect("parse redacted Runtime CLI payload");
assert!(parsed.get("sessionPath").is_none());
assert!(parsed["runtime"].get("eventPath").is_none());
assert!(parsed["runtimes"][0].get("taskPath").is_none());
assert_eq!(parsed["runtime"]["state"]["sessionId"], "session-7");
assert_eq!(parsed["runtimes"][0]["state"]["status"], "running");
}
#[test]
fn parses_agent_steer_with_stdin_only_contract() {
let command = parse_cli_command(&[
@@ -202,6 +202,10 @@ struct AgentRuntimeState {
#[serde(default)]
tool_action_budget: u32,
#[serde(default)]
plan_revision: u64,
#[serde(default)]
plan_explanation: String,
#[serde(default)]
plan: Vec<String>,
#[serde(default)]
plan_steps: Vec<AgentRuntimePlanStep>,
@@ -7,6 +7,7 @@ use std::time::{Duration, Instant};
const SWARM_CHAT_POLL_INTERVAL: Duration = Duration::from_millis(250);
const SWARM_CHAT_SETTLE_WINDOW: Duration = Duration::from_millis(1_500);
const SWARM_CHAT_HISTORY_LIMIT: usize = 50;
const SWARM_CHAT_PLAN_STEP_LIMIT: usize = 8;
#[derive(Debug, Eq, PartialEq)]
enum SwarmChatInput {
@@ -694,15 +695,113 @@ fn print_runtime_state<W: Write>(
relation,
state.current_action
)
.map_err(|error| format!("写入终端失败:{error}"))
.map_err(|error| format!("写入终端失败:{error}"))?;
let completed = state
.plan_steps
.iter()
.filter(|step| step.status == "completed")
.count();
let current_step = runtime_current_plan_step(state)
.map(|step| {
format!(
"#{} [{}] {}",
step.index.saturating_add(1),
runtime_cli_value(&step.status),
runtime_cli_value(&step.title)
)
})
.unwrap_or_else(|| "-".to_string());
writeln!(
output,
"[计划] revision={} completed={}/{} current={} | waiting={} | next={}",
state.plan_revision,
completed,
state.plan_steps.len(),
current_step,
runtime_cli_value(&state.waiting_on),
runtime_cli_value(&state.next_step)
)
.map_err(|error| format!("写入终端失败:{error}"))?;
if !state.plan_explanation.trim().is_empty() {
writeln!(
output,
" [计划说明] {}",
runtime_cli_value(&state.plan_explanation)
)
.map_err(|error| format!("写入终端失败:{error}"))?;
}
for step in state.plan_steps.iter().take(SWARM_CHAT_PLAN_STEP_LIMIT) {
writeln!(
output,
" [计划步骤] #{} [{}] {}",
step.index.saturating_add(1),
runtime_cli_value(&step.status),
runtime_cli_value(&step.title)
)
.map_err(|error| format!("写入终端失败:{error}"))?;
}
if state.plan_steps.len() > SWARM_CHAT_PLAN_STEP_LIMIT {
writeln!(
output,
" [计划] 另有 {} 条步骤未显示",
state.plan_steps.len() - SWARM_CHAT_PLAN_STEP_LIMIT
)
.map_err(|error| format!("写入终端失败:{error}"))?;
}
Ok(())
}
fn runtime_current_plan_step(state: &AgentRuntimeState) -> Option<&AgentRuntimePlanStep> {
state
.active_plan_step_index
.and_then(|active_index| {
state
.plan_steps
.iter()
.find(|step| step.index == active_index)
.or_else(|| state.plan_steps.get(active_index as usize))
})
.or_else(|| {
state.plan_steps.iter().find(|step| {
matches!(
step.status.as_str(),
"active" | "in_progress" | "running" | "waiting-for-confirmation"
)
})
})
.or_else(|| {
state
.plan_steps
.iter()
.find(|step| step.status == "pending")
})
}
fn runtime_cli_value(value: &str) -> &str {
let value = value.trim();
if value.is_empty() {
"-"
} else {
value
}
}
fn runtime_state_signature(
state: &AgentRuntimeState,
queue: &AgentRuntimeTaskQueueSummary,
) -> String {
format!(
"{}:{}:{}:{}:{}:{}:{}:{}:{}",
let completed_plan_steps = state
.plan_steps
.iter()
.filter(|step| step.status == "completed")
.count();
let current_plan_step = runtime_current_plan_step(state)
.map(|step| format!("{}:{}:{}", step.index, step.status, step.title))
.unwrap_or_default();
let mut signature = format!(
"{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}:{}",
state.run_id,
state.status,
state.phase,
@@ -711,8 +810,21 @@ fn runtime_state_signature(
queue.pending,
queue.running,
queue.waiting_for_confirmation,
queue.updated_at
)
queue.updated_at,
state.plan_revision,
state
.active_plan_step_index
.map(|index| index.to_string())
.unwrap_or_default(),
completed_plan_steps,
state.plan_steps.len(),
current_plan_step,
state.waiting_on,
state.next_step
);
signature.push(':');
signature.push_str(&state.plan_explanation);
signature
}
fn runtime_event_key(event: &AgentRuntimeEvent) -> String {
@@ -812,6 +924,92 @@ mod tests {
assert!(output.is_empty());
}
#[test]
fn runtime_plan_revision_and_current_step_change_state_signature() {
let mut snapshot = runtime("running", "planning", 0);
snapshot.state.updated_at = 100;
snapshot.task_queue.updated_at = 100;
snapshot.state.plan_revision = 1;
snapshot.state.plan_steps = vec![
AgentRuntimePlanStep {
index: 0,
title: "读取现有 CLI".to_string(),
status: "in_progress".to_string(),
detail: None,
updated_at: 100,
},
AgentRuntimePlanStep {
index: 1,
title: "补充计划展示".to_string(),
status: "pending".to_string(),
detail: None,
updated_at: 100,
},
];
snapshot.state.active_plan_step_index = Some(0);
let initial = runtime_state_signature(&snapshot.state, &snapshot.task_queue);
snapshot.state.plan_revision = 2;
let revised = runtime_state_signature(&snapshot.state, &snapshot.task_queue);
assert_ne!(initial, revised);
snapshot.state.plan_steps[0].title = "核对现有 CLI".to_string();
let current_step_changed = runtime_state_signature(&snapshot.state, &snapshot.task_queue);
assert_ne!(revised, current_step_changed);
snapshot.state.plan_steps[0].status = "completed".to_string();
snapshot.state.plan_steps[1].status = "in_progress".to_string();
snapshot.state.active_plan_step_index = Some(1);
let advanced = runtime_state_signature(&snapshot.state, &snapshot.task_queue);
assert_ne!(current_step_changed, advanced);
}
#[test]
fn runtime_plan_output_is_bounded_and_omits_private_observations() {
let mut snapshot = runtime("running", "planning", 0);
snapshot.state.plan_revision = 7;
snapshot.state.plan_explanation = "已完成读取,进入验证".to_string();
snapshot.state.current_action = "展示持久计划".to_string();
snapshot.state.waiting_on = "开发者确认".to_string();
snapshot.state.next_step = "运行 focused cargo test".to_string();
snapshot.state.observations = vec![
"PRIVATE_OBSERVATION_SENTINEL".to_string(),
"PRIVATE_DETAIL_SENTINEL".to_string(),
];
snapshot.state.plan_steps = (0..10)
.map(|index| AgentRuntimePlanStep {
index,
title: format!("计划步骤 {}", index + 1),
status: match index {
0 | 1 => "completed",
2 => "in_progress",
_ => "pending",
}
.to_string(),
detail: Some(format!("PRIVATE_STEP_DETAIL_{index}")),
updated_at: 100,
})
.collect();
snapshot.state.active_plan_step_index = Some(2);
let mut output = Vec::new();
print_runtime_state(&snapshot.state, &snapshot.task_queue, &mut output)
.expect("print runtime plan progress");
let output = String::from_utf8(output).expect("runtime output is utf-8");
assert!(output.contains(
"[计划] revision=7 completed=2/10 current=#3 [in_progress] 计划步骤 3 | waiting=开发者确认 | next=运行 focused cargo test"
));
assert!(output.contains("[计划说明] 已完成读取,进入验证"));
assert_eq!(output.matches("[计划步骤]").count(), 8);
assert!(output.contains("[计划步骤] #8 [pending] 计划步骤 8"));
assert!(output.contains("另有 2 条步骤未显示"));
assert!(!output.contains("计划步骤 9"));
assert!(!output.contains("PRIVATE_OBSERVATION_SENTINEL"));
assert!(!output.contains("PRIVATE_DETAIL_SENTINEL"));
assert!(!output.contains("PRIVATE_STEP_DETAIL"));
}
#[test]
fn input_channel_preserves_lines_and_eof() {
let (tx, rx) = mpsc::channel();
File diff suppressed because it is too large Load Diff
+282 -45
View File
@@ -270,6 +270,8 @@ interface AgentRuntimeState {
loopIteration?: number;
maxLoopIterations?: number;
toolActionBudget?: number;
planRevision?: number;
planExplanation?: string;
plan: string[];
planSteps?: AgentRuntimePlanStep[];
activePlanStepIndex?: number | null;
@@ -315,11 +317,12 @@ interface AgentRuntimePendingToolActionSummary {
}
interface AgentRuntimePlanStep {
index: number;
title: string;
step?: string;
status: string;
detail: string | null;
updatedAt: number;
index?: number;
title?: string;
detail?: string | null;
updatedAt?: number;
}
interface AgentRuntimeTaskQueueSummary {
@@ -664,6 +667,7 @@ function agentRuntimePlanStepsFromPlan(plan: string[]): AgentRuntimePlanStep[] {
.filter((item) => item.trim().length > 0)
.slice(0, 8)
.map((title, index) => ({
step: title,
index,
title,
status: index === 0 ? 'active' : 'pending',
@@ -672,39 +676,147 @@ function agentRuntimePlanStepsFromPlan(plan: string[]): AgentRuntimePlanStep[] {
}));
}
function agentRuntimePlanStepText(step: AgentRuntimePlanStep) {
return (step.step ?? step.title ?? '').trim();
}
function normalizeAgentRuntimePlanStep(
step: AgentRuntimePlanStep,
fallbackIndex: number,
): AgentRuntimePlanStep | null {
const text = agentRuntimePlanStepText(step);
if (!text) {
return null;
}
const index =
typeof step.index === 'number' &&
Number.isFinite(step.index) &&
step.index >= 0
? Math.trunc(step.index)
: fallbackIndex;
return {
...step,
step: text,
index,
title: step.title?.trim() || text,
status: step.status?.trim() || 'pending',
detail: step.detail ?? null,
updatedAt:
typeof step.updatedAt === 'number' && Number.isFinite(step.updatedAt)
? step.updatedAt
: 0,
};
}
function normalizeAgentRuntimePlanStepList(steps: AgentRuntimePlanStep[]) {
return steps
.map((step, index) => normalizeAgentRuntimePlanStep(step, index))
.filter((step): step is AgentRuntimePlanStep => step !== null);
}
function normalizeAgentRuntimePlanSteps(
state: AgentRuntimeState,
previous?: AgentRuntimeState | null,
) {
if (state.planSteps && state.planSteps.length > 0) {
return state.planSteps;
return normalizeAgentRuntimePlanStepList(state.planSteps);
}
if (state.planSteps !== undefined && state.planRevision !== undefined) {
return [];
}
if (previous?.planSteps && previous.planSteps.length > 0) {
return previous.planSteps;
return normalizeAgentRuntimePlanStepList(previous.planSteps);
}
return agentRuntimePlanStepsFromPlan(state.plan ?? []);
}
function normalizeAgentRuntimeActivePlanStepIndex(
state: AgentRuntimeState,
planSteps: AgentRuntimePlanStep[],
previous?: AgentRuntimeState | null,
) {
if (state.activePlanStepIndex !== undefined) {
return state.activePlanStepIndex;
if (state.activePlanStepIndex === null) {
return null;
}
return Number.isFinite(state.activePlanStepIndex) &&
state.activePlanStepIndex >= 0
? Math.trunc(state.activePlanStepIndex)
: null;
}
if (previous?.activePlanStepIndex !== undefined) {
return previous.activePlanStepIndex;
}
const activeStep = (state.planSteps ?? []).find(
(step) => step.status === 'active',
const activeStepIndex = planSteps.findIndex((step) =>
['active', 'in_progress', 'running'].includes(step.status),
);
return activeStep?.index ?? null;
if (activeStepIndex < 0) {
return null;
}
return planSteps[activeStepIndex]?.index ?? activeStepIndex;
}
function sameAgentRuntimeRun(
state: AgentRuntimeState,
previous?: AgentRuntimeState | null,
) {
return Boolean(
previous &&
previous.agentId === state.agentId &&
previous.sessionId === state.sessionId &&
previous.runId === state.runId,
);
}
function normalizeAgentRuntimePlanRevision(
revision: number | undefined,
previousRevision: number | undefined,
) {
if (
typeof revision === 'number' &&
Number.isFinite(revision) &&
revision >= 0
) {
return Math.trunc(revision);
}
return previousRevision;
}
function shouldReplaceAgentRuntimePlan(
state: AgentRuntimeState,
previous?: AgentRuntimeState | null,
) {
if (!previous) {
return true;
}
const previousRevision = normalizeAgentRuntimePlanRevision(
previous.planRevision,
undefined,
);
if (previousRevision === undefined) {
return true;
}
const nextRevision = normalizeAgentRuntimePlanRevision(
state.planRevision,
undefined,
);
return nextRevision !== undefined && nextRevision >= previousRevision;
}
function normalizeAgentRuntimeState(
state: AgentRuntimeState,
previous?: AgentRuntimeState | null,
): AgentRuntimeState {
const previousPlanState = sameAgentRuntimeRun(state, previous)
? previous
: null;
const replacePlan = shouldReplaceAgentRuntimePlan(state, previousPlanState);
const planState = replacePlan ? state : previousPlanState!;
const planFallbackState = replacePlan ? previousPlanState : null;
const planSteps = normalizeAgentRuntimePlanSteps(
planState,
planFallbackState,
);
return {
...state,
currentGoal: state.currentGoal ?? state.currentTask ?? '',
@@ -714,10 +826,18 @@ function normalizeAgentRuntimeState(
maxLoopIterations:
state.maxLoopIterations ?? previous?.maxLoopIterations ?? 3,
toolActionBudget: state.toolActionBudget ?? previous?.toolActionBudget ?? 3,
planSteps: normalizeAgentRuntimePlanSteps(state, previous),
plan: planState.plan ?? planFallbackState?.plan ?? [],
planRevision: normalizeAgentRuntimePlanRevision(
planState.planRevision,
planFallbackState?.planRevision,
),
planExplanation:
planState.planExplanation ?? planFallbackState?.planExplanation,
planSteps,
activePlanStepIndex: normalizeAgentRuntimeActivePlanStepIndex(
state,
previous,
planState,
planSteps,
planFallbackState,
),
recentToolCalls: state.recentToolCalls ?? previous?.recentToolCalls ?? [],
toolPolicy: state.toolPolicy ??
@@ -745,6 +865,32 @@ function normalizeAgentRuntimeState(
};
}
function mergeAgentRuntimeStateIntoMap(
current: Record<string, AgentRuntimeState | undefined>,
incoming: AgentRuntimeState,
preserveNewerRun: boolean,
) {
const previous = Object.values(current).find(
(runtime) => runtime?.agentId === incoming.agentId,
);
const mergedRuntime =
preserveNewerRun &&
previous &&
!sameAgentRuntimeRun(incoming, previous) &&
previous.updatedAt >= incoming.updatedAt
? previous
: normalizeAgentRuntimeState(incoming, previous);
const next = { ...current };
for (const [key, runtime] of Object.entries(next)) {
if (runtime?.agentId === incoming.agentId) {
delete next[key];
}
}
next[mergedRuntime.agentId] = mergedRuntime;
next[mergedRuntime.taskId] = mergedRuntime;
return next;
}
function agentRuntimeStateFromResult(
result: AgentRuntimeResult,
previous?: AgentRuntimeState | null,
@@ -966,8 +1112,12 @@ function formatAgentRuntimeLoopProgress(
}`;
}
function formatAgentRuntimePlanStep(step: AgentRuntimePlanStep) {
return `#${step.index + 1} ${step.status} · ${step.title}${
function formatAgentRuntimePlanStep(
step: AgentRuntimePlanStep,
fallbackIndex = 0,
) {
const index = step.index ?? fallbackIndex;
return `#${index + 1} ${step.status} · ${agentRuntimePlanStepText(step)}${
step.detail ? ` · ${step.detail}` : ''
}`;
}
@@ -976,9 +1126,20 @@ function agentRuntimeActivePlanStep(
runtime: Pick<AgentRuntimeState, 'planSteps' | 'activePlanStepIndex'>,
) {
const steps = runtime.planSteps ?? [];
const activePlanStepIndex = runtime.activePlanStepIndex;
if (activePlanStepIndex !== null && activePlanStepIndex !== undefined) {
const indexedStep =
steps.find((step) => step.index === activePlanStepIndex) ??
steps[activePlanStepIndex];
if (indexedStep) {
return indexedStep;
}
}
return (
steps.find((step) => step.index === runtime.activePlanStepIndex) ??
steps.find((step) => step.status === 'active') ??
steps.find((step) =>
['active', 'in_progress', 'running'].includes(step.status),
) ??
steps.find((step) => step.status === 'pending') ??
null
);
}
@@ -1089,7 +1250,11 @@ function AgentRuntimeStatusPanel({
).reverse();
const recentToolCalls = (runtime.recentToolCalls ?? []).slice(-3).reverse();
const recentTasks = (runtime.recentTasks ?? []).slice(-3).reverse();
const planSteps = (runtime.planSteps ?? []).slice(0, 5);
const planSteps = (runtime.planSteps ?? []).slice(0, 8);
const hasPersistentPlan =
runtime.planRevision !== undefined ||
Boolean(runtime.planExplanation) ||
planSteps.length > 0;
const toolPolicy = runtime.toolPolicy;
const taskQueueSummary = formatAgentRuntimeTaskQueue(runtime.taskQueue);
const loopProgress = formatAgentRuntimeLoopProgress(runtime);
@@ -1233,14 +1398,20 @@ function AgentRuntimeStatusPanel({
: ''}
</small>
) : null}
{planSteps.length > 0 ? (
{hasPersistentPlan ? (
<div aria-label="Agent 计划进度">
<strong>计划进度</strong>
{planSteps.map((step) => (
{(runtime.planRevision ?? 0) > 0 ? (
<small>{`计划修订号:${runtime.planRevision}`}</small>
) : null}
{runtime.planExplanation ? (
<small>{`计划说明:${runtime.planExplanation}`}</small>
) : null}
{planSteps.map((step, index) => (
<small
key={`${runtime.sessionId}-plan-step-${step.index}-${step.updatedAt}`}
key={`${runtime.sessionId}-plan-step-${step.index ?? index}-${step.updatedAt ?? 0}-${index}`}
>
{formatAgentRuntimePlanStep(step)}
{formatAgentRuntimePlanStep(step, index)}
</small>
))}
</div>
@@ -1796,6 +1967,58 @@ function projectSupervisorRuntimeStatusLabel(
return '分析';
}
function projectSupervisorCollaboratingAgentCount(
supervisorRuntime: AgentRuntimeState | null,
runtimeByAgentId: Record<string, AgentRuntimeState | undefined>,
) {
if (!supervisorRuntime?.runId) {
return 0;
}
const agentIds = new Set<string>();
for (const runtime of Object.values(runtimeByAgentId)) {
if (
!runtime ||
runtime.agentId === PROJECT_SUPERVISOR_AGENT_ID ||
!['agent-delegate', 'agent-delegate-retry'].includes(runtime.source) ||
runtime.parentAgentId !== PROJECT_SUPERVISOR_AGENT_ID ||
runtime.parentRunId !== supervisorRuntime.runId
) {
continue;
}
agentIds.add(runtime.agentId);
}
return agentIds.size;
}
function formatProjectSupervisorCompactProgress(
runtime: AgentRuntimeState,
collaboratingAgentCount: number,
) {
const planSteps = (runtime.planSteps ?? []).filter(
(step) => agentRuntimePlanStepText(step).length > 0,
);
const completedPlanStepCount = planSteps.filter(
(step) => step.status === 'completed',
).length;
const activePlanStep = agentRuntimeActivePlanStep(runtime);
const currentPlanStep = activePlanStep
? agentRuntimePlanStepText(activePlanStep)
: planSteps.length > 0 && completedPlanStepCount === planSteps.length
? '已完成'
: '暂无';
const waitingOn =
runtime.waitingOn?.trim() || agentRuntimeWaitingOnFromPhase(runtime.phase);
const nextStep =
runtime.nextStep?.trim() || agentRuntimeNextStepFromPhase(runtime.phase);
return [
`计划完成:${completedPlanStepCount}/${planSteps.length}`,
`当前步骤:${currentPlanStep}`,
`等待:${waitingOn}`,
`下一步:${nextStep}`,
`专业 Agent 协作:${collaboratingAgentCount}`,
].join(' · ');
}
function isTransientProjectOpenMessage(
message: ChatMessage,
projectPath: string,
@@ -12629,7 +12852,10 @@ export function deriveAgentStatusCards(
runtimeMaxLoopIterations: runtime?.maxLoopIterations ?? null,
runtimeToolActionBudget: runtime?.toolActionBudget ?? null,
runtimeActivePlanStep: activePlanStep
? formatAgentRuntimePlanStep(activePlanStep)
? formatAgentRuntimePlanStep(
activePlanStep,
runtime?.planSteps?.indexOf(activePlanStep) ?? 0,
)
: null,
runtimeTaskQueue: runtime?.taskQueue ?? null,
runtimeRecentTasks: runtime?.recentTasks ?? [],
@@ -14454,9 +14680,14 @@ export function App() {
const nextRuntime = agentRuntimeStateFromResult(payload.runtime);
if (payload.agentId === PROJECT_SUPERVISOR_AGENT_ID) {
const expectedSessionId = projectSupervisorSessionIdRef.current;
const currentRuntime = projectSupervisorRuntimeRef.current;
if (
expectedSessionId !== null &&
nextRuntime.sessionId !== expectedSessionId
nextRuntime.agentId !== PROJECT_SUPERVISOR_AGENT_ID ||
nextRuntime.runId !== payload.runId ||
(expectedSessionId !== null &&
nextRuntime.sessionId !== expectedSessionId) ||
(currentRuntime !== null &&
!sameAgentRuntimeRun(nextRuntime, currentRuntime))
) {
return;
}
@@ -23072,15 +23303,9 @@ export function App() {
if (!runtime) {
return;
}
setAgentRuntimeById((current) => {
const previous = current[runtime.agentId] ?? current[runtime.taskId];
const mergedRuntime = normalizeAgentRuntimeState(runtime, previous);
return {
...current,
[mergedRuntime.agentId]: mergedRuntime,
[mergedRuntime.taskId]: mergedRuntime,
};
});
setAgentRuntimeById((current) =>
mergeAgentRuntimeStateIntoMap(current, runtime, false),
);
}
async function refreshAgentRuntimes(
@@ -23173,13 +23398,16 @@ export function App() {
if (localProjectPathRef.current !== nextProjectPath) {
return;
}
const nextRuntimeById: Record<string, AgentRuntimeState | undefined> = {};
const nextRuntimes: AgentRuntimeState[] = [];
for (const runtimeResult of runtimes) {
const runtime = agentRuntimeStateFromResult(runtimeResult);
nextRuntimeById[runtime.agentId] = runtime;
nextRuntimeById[runtime.taskId] = runtime;
nextRuntimes.push(agentRuntimeStateFromResult(runtimeResult));
}
setAgentRuntimeById(nextRuntimeById);
setAgentRuntimeById((current) =>
nextRuntimes.reduce(
(next, runtime) => mergeAgentRuntimeStateIntoMap(next, runtime, true),
current,
),
);
} catch {
if (localProjectPathRef.current === nextProjectPath) {
setAgentRuntimeById({});
@@ -23389,12 +23617,16 @@ export function App() {
projectSupervisorRuntimeError,
);
const projectSupervisorStatusDetail =
projectSupervisorRuntimeError ||
projectSupervisorRuntime?.error ||
(projectSupervisorStatus && projectSupervisorStatus !== '已完成'
? projectSupervisorRuntime?.currentAction
: '') ||
'';
projectSupervisorRuntimeError || projectSupervisorRuntime?.error || '';
const projectSupervisorCompactProgress = projectSupervisorRuntime
? formatProjectSupervisorCompactProgress(
projectSupervisorRuntime,
projectSupervisorCollaboratingAgentCount(
projectSupervisorRuntime,
agentRuntimeById,
),
)
: '';
const projectSupervisorPendingToolAction =
projectSupervisorRuntime?.pendingToolAction ?? null;
const visibleAgentConversationMessages = latestVisibleItems(
@@ -24026,6 +24258,11 @@ export function App() {
{projectSupervisorStatusDetail ? (
<small>{projectSupervisorStatusDetail}</small>
) : null}
{projectSupervisorCompactProgress ? (
<small aria-label="项目总控 Agent 进度">
{projectSupervisorCompactProgress}
</small>
) : null}
{projectSupervisorPendingToolAction ? (
<div
className="pending-command"
File diff suppressed because it is too large Load Diff
@@ -4563,6 +4563,20 @@
- UI:普通用户只看到总控 Agent 的紧凑状态、等待对象、协作数量、安全确认和唯一最终回复;不展示内部工具计划、原始 observation、动态 child 或开发控制台。Runner 未提供 token delta 时只显示真实状态,不做伪流式。
- 验收:Rust `project_supervisor_` 定向回归覆盖 ID/prompt/config、delivery/claim 幂等、排序锁零部分认领、Agent DB 旁路、未 Observed 门禁、Provider planning 前 durable 等待、parent-wake coalescing/结构性错误投影、重启损坏 barrier、完整身份与迟到 child suppression、delivery `.previous` 恢复、旧 receipt runId 冲突、executing `run_status` 续接与委派 policy 重验;Runner 内部回归覆盖定向 wake 只有目标推进后成功、可重试结果不缓存;Session lane 回归覆盖入队后才通知 Runner。客户端定向回归覆盖 active Supervisor Session、same-run steer、legacy 历史合并、确认/拒绝和唯一终态 assistant。修复后真实 Provider 已证明 design/art 两个专业 Agent 同秒进入 running 并重叠 20 秒,父 run 只写 1 条 waiting、同一 `Observed` claim 认领 2 份回执、首轮恰好 1 条 user / 1 条 assistant、无 reconciliation;同一 Session 第二轮引用上文完成且未新增委派。项目范围精确密钥扫描为 0。V1.15 首轮跑偏和本轮修复前 revision 误伤仍只保留为负向历史,不作为通过证据。
## 2026-07-15 AI 游戏创作 Agent Runtime V1.17 单 Agent 持久计划
- 决策:`submit_agent_tool_plan` 顶层新增 nullable `planUpdate={explanation,steps[{step,status}]}`。strict function arguments 必须出现该字段,无真实变化时传 `null`;结构化更新最多 8 个唯一步骤,状态只允许 `pending / in_progress / completed` 且至多一个 `in_progress`。旧文本协议可缺字段,legacy `plan` 只作 fallback;当前 run 一旦有 `planRevision > 0`,legacy `plan` 不得再覆盖结构化计划。
- 决策:Runtime state 持久化 `planRevision / planExplanation / planSteps / activePlanStepIndex`。有效变化使 revision 单调递增,完全相同的更新幂等不增号,非法更新不改快照;已完成或历史快照中已有的失败终态步骤必须保留,completed 不得回退。外层 run 进入 `failed / budget-exhausted` 时保留最后一次可信计划的 revision、说明、步骤状态和 active index,不把未完成步骤机械改写为失败。结构化计划建立后,工具 action 下标和旧自动步骤 helper 全部失去进度写权限,Agent 必须依据真实 observation 显式更新计划。
- 决策:任一结构化步骤未完成时,空 actions、Provider response 和恢复中的 finalization 都由 `runtime.plan_update` blocker 拦截,不能写 assistant 或 completed。计划更新只属于 Runtime 私有元数据,不是工具 action,不读取或改写项目 policy,不触发 confirm/deny,不推进 project revision 或 verification gate,也不改变待确认动作 fingerprint。
- 审计边界:`thinking_summary` 公共 event 只留正文 SHA-256 与字符数,legacy `plan` event 只留步骤数;`agent.runtime.plan_update` 只留 explanation 哈希与字符数,以及 step 标题哈希、状态和数量。`agent.runtime.tool_plan.repair` 只留尝试计数、协议以及模型输出/调用体预览、解析错误、callId / functionName 的哈希与长度,不落原始正文、错误或 function arguments;仅当前 planning 的私有有界 repair 请求可保留经过过滤的必要上下文。
- 恢复与 steer:context bundle 升级为 `game-creator-runtime-context-bundle.v3` 并保存完整计划快照;v3 revision 或快照与 Runtime state 不一致时失败关闭,损坏 state 进入 `needs-reconciliation`。v2 继续可读,但只能在原身份、task、revision 与 verification gate 校验通过后从当前 state 补齐计划字段,后续 checkpoint 写 v3;v1 仍拒绝。Runner 重启、确认续跑和 stale finalization 不得重建或自动完成计划。same-run steer 丢弃旧 actions / 旧回复但保留终态步骤和 revision,下一版只重审未完成部分。
- Finalization:journal 升级为 `game-creator-runtime-finalization.v2`,在 `prepared` 时绑定最终完整计划快照与 `planSnapshotFingerprint`,并把计划指纹纳入幂等 `finalizationId`。assistant 已落盘而 Runtime state 丢失时,从唯一 task record 与 v2 journal 恢复原 structured plan 后补齐 completed,不请求 Provider、不重放工具;assistant 尚未落盘而 state 丢失时保留 journal 并进入 `needs-reconciliation`,不得只凭 task 或 prepared journal 猜计划并写回复。
- 展示:开发 Agent UI、项目内开发面板、CLI / `agent.run_status` 有界展示 revision、说明和最多 8 个完整步骤;刷新合并只沿用同一 Agent/Session/run。正式用户 Project Supervisor 只显示完成数、当前步骤、等待对象、下一步和专业 Agent 协作数量,不暴露 revision、内部说明、完整步骤、原始 observation、内部动作或动态 child。
- 验收:确定性回归覆盖 schema、native function 显式字段与文本兼容、限制、单调性、外层失败进度保留、终态保留、动作下标零推进、未完成 final 门禁、损坏状态、v3/v2 恢复、finalization v2 state 丢失恢复、公共审计零正文、计划元数据 revision/policy 中立和两类 UI。恢复/steer 专项必须证明同一 run/session、revision 不回退、终态不丢、旧动作零执行和副作用零重放。真实 Provider 必须在无计划/工具配方的 disposable 项目中自行建立并多次更新计划,经历一次 same-run steer 与一次 Runner 重启,最终只在全部步骤 completed 后写唯一 assistant,并由 Runtime state、v3 bundle、task/event/Agent DB/conversation 和副作用计数交叉取证;截至 2026-07-15 尚未记录该专项 PASS。
- 全量回归修正:context bundle v3 为保持计划快照一致,会在每个 action / observation 后同步;repository startup fingerprint 因此不能继续从最新 bundle 读取,否则同一 planning 批次的前置验证改变规范文件后,后续旧写动作会错误放行。pending action 升级为 `game-creator-pending-action.v4`,绑定 Provider planning 实际渲染的 repository fingerprint;同批 actions、确认和恢复统一复核该快照,旧 v1-v3 失败关闭。五类写动作 drift 回归证明旧动作零执行。
- 终审补充:finalization v2 读取边界必须再次要求所有结构化步骤 `completed` 且 active index 为空;仅重算合法 `planSnapshotFingerprint / finalizationId` 的未完成快照也失败关闭。开发 CLI 的 Runtime JSON 只输出状态和安全身份,递归移除 `sessionPath / eventPath / taskPath`,不把项目绝对存储位置写入命令 transcript;Tauri/App 内部结果结构保持不变。
- 验收现状:确定性回归已通过,真实 `gpt-5.5 llm-runtime` 连续三轮均在首个 Provider planning POST 返回前因同一 TLS record-layer failure 失败,未产生 plan/tool/kill/steer 证据;第三轮绝对路径、密钥和诱饵泄漏为 0。V1.17 继续保持未 PASS,Provider 恢复后必须完整重跑。
## 2026-07-13 普通微信支付 V3 退款使用统一观察事务闭环
- 背景:普通微信支付 V3 的退款申请响应、退款结果回调、主动查单和商户平台手工退款发现可能重复、乱序或只出现其中一种;原充值订单只有单一终态,无法表达多次部分退款、权益回收欠款和会员人工处理。
@@ -3171,6 +3171,44 @@
- 验证:Rust 定向回归使用 `project_supervisor_` 前缀,覆盖 delivery/claim 状态机、同 action 幂等、第 4 个新委派拒绝与已预留委派复用/拒绝 suppression、后续 delivery 锁忙时零部分认领、Agent DB 故障后回执仍可重放、未 Observed 阻断 final、Provider planning 前 durable 等待、parent-wake coalescing/结构性错误、重启损坏 barrier、错配和迟到 child、executing `run_status` 续接与 delegate policy 重验;`agent_background_enqueue_notifies_only_after_session_lane_release` 覆盖入队锁序,Runner 内部测试覆盖定向 wake 与不缓存重试。真实 Provider 必须同时证明专业 Agent 时间区间重叠、父 run 仅一次 waiting、同一 Observed claim 认领全部回执、唯一 assistant、第二轮历史引用不新增委派和项目范围密钥扫描为 0。
- 关联:`docs/technical/【技术方案】AI游戏创作Agent Runtime V1.1-2026-07-12.md`、`apps/ai-game-creator-shell/src-tauri/src/delegation.rs`、`agent.rs`、`runner.rs`、`tests.rs`。
## 单 Agent 持久计划不能靠工具下标或恢复猜进度
- 现象:工具 action 1 成功后第二个计划步骤被自动标成完成,模型仍有 pending / in_progress 步骤却写出最终回复;或 Runner 重启、刷新 UI、same-run steer 后 `planRevision` 回退、已完成步骤消失,legacy `plan` 又覆盖新计划。另一类错误是仅更新计划就触发项目 revision 漂移、verification 失效或权限确认。
- 原因:旧 `planSteps` 由短 `plan` 派生,并按 actions 数组下标驱动 `active / completed`,它无法表达跨窗口、恢复和 steer 后的真实任务进度。把 context bundle 当唯一计划事实源、把 v2 缺字段当空计划,或把 `planUpdate` 伪装成受策略工具,也会让 Runtime state、恢复快照和项目副作用门禁互相污染。
- 处理:V1.17 的 `planUpdate` 只接受 1 到 8 个唯一 `pending / in_progress / completed` 步骤,至多一个 `in_progress`;native function 即使无变化也必须显式传 `planUpdate: null`,只有旧文本 JSON 可省略。当前 run 建立结构化计划后,legacy `plan` 和所有按 action 下标推进的 helper 都只能读不能写。`planRevision` 只在有效变化时单调增加;completed 与历史快照中已有的 failed 终态即使被下一版省略也必须合并保留,completed 回退或合并后超过 8 步时整次拒绝。外层 run 进入 `failed / budget-exhausted` 时必须原样保留最后可信进度,不把 pending / in_progress 机械标成 failed。任何非 completed 步骤都阻止成功 final,不能用 response 文本绕过。
- 恢复与 steer:context bundle v3 必须保存并复核 revision、说明、步骤和 active index;与 Runtime state 不一致时失败关闭,不能选“看起来更新”的一份。v2 只能在原身份、task、revision 和 verification gate 校验通过后从当前 state 补齐计划,v1 继续拒绝。same-run steer 只作废旧 Provider actions / 回复并要求重审未完成部分,不能清空终态步骤或重置 revision;Runner 重启同样不得自动勾选。finalization 另以 v2 journal 绑定最终完整计划快照:只有 assistant 已落盘时,state 丢失才可从该快照恢复终态;assistant 未落盘且 state 不可读时必须进入 reconciliation。
- 边界:计划更新是 `.agent/runtime` 私有元数据,不经过项目工具 policy,不推进 project revision 或 verification gate,不改变 pending action fingerprint。开发 UI/CLI 可展示最多 8 步完整计划;普通用户 Supervisor 只能显示完成数、当前步骤、等待、下一步和协作数量,不能把内部 explanation、完整步骤或 currentAction 搬到主聊天。
- 验证:运行 `structured_plan_` 与 `agent_runtime_context_bundle_migrates_v2_and_rejects_v3_plan_mismatch` Rust 定向用例,并用 `appSurface.test.ts` 覆盖刷新恢复和 Supervisor 紧凑摘要。真实 Provider 必须在无计划配方下多次更新计划,完成一步后接受 same-run steer,再经历 Runner 强杀恢复;最终证明 run/session 不变、revision 不回退、终态不丢、旧动作与副作用不重放、未完成时零 assistant、完成后唯一 assistant,计划更新前后 project revision / policy 不变。未运行该门禁时不得写 V1.17 PASS。
### Finalization 只绑定回复会在 Runtime state 丢失后丢计划
- 症状:assistant 已经按稳定 messageId 写入 conversation,进程却在 Runtime completed 投影前退出;重启后 state 文件缺失,系统从 task record 重建出默认或 legacy 计划,最终回复虽然没有重复,结构化计划 revision 和步骤却丢失。反向地,assistant 尚未写入时若也用 journal 单独猜计划,会把过期回复错误提交给用户。
- 原因:task record 不携带完整结构化计划,context bundle 也可能对应 finalization 前的其它 checkpoint;只给 journal 绑定回复和 verification gate,无法证明准备提交时的最终计划快照。
- 处理:`game-creator-runtime-finalization.v2` 在 prepared 时保存完整 `planRevision / planExplanation / plan / planSteps / activePlanStepIndex` 与 `planSnapshotFingerprint`,并将指纹纳入 finalizationId。读取时除校验格式和指纹外,还要再次要求全部结构化步骤 completed 且 active index 为空。assistant 已存在且 task 唯一时,允许从 v2 快照恢复原计划并补齐 Runtime completed;assistant 不存在而 state 缺失或不可读时保留 journal、进入 `needs-reconciliation`,不能自动写回复。已有 state 与 journal 快照冲突时同样失败关闭或让 prepared 回复失效后在同一 run 重规划。
- 验证:`finalization_resume_recovers_persisted_assistant_without_runtime_state` 必须证明无 Provider 重放、assistant 唯一且恢复后的 revision/说明/步骤与 prepared 快照完全一致;`structured_plan_finalization_without_readable_runtime_state_needs_reconciliation` 必须证明 assistant 未落盘时 missing/corrupt state 都零回复、journal 保留且无 completed 审计;`finalization_resume_blocks_internally_consistent_incomplete_plan_snapshot` 必须证明重算合法指纹和 finalizationId 也不能提交未完成计划。
### CLI Runtime JSON 不能暴露项目绝对存储路径
- 症状:真实 E2E 的 task/event/Agent DB/report 均无项目绝对路径,子进程 transcript 扫描却稳定命中 6 次;入队和首次状态读取各返回 3 个路径。
- 原因:开发 CLI 直接序列化 `AgentRuntimeResult`,把仅供 Tauri/App 定位本地 sidecar 的 `sessionPath / eventPath / taskPath` 一并写进 `runtimeJson`。后续 confirm、steer 和 resume 复用同一结果结构,也会重复暴露。
- 处理:保持 Tauri 内部契约不变,只在 CLI JSON 输出视图递归删除三个存储路径;`state`、task queue、events、tasks、run/session/action 身份和 steer 状态继续保留,验收器仍能解析必要证据。不要靠 E2E 忽略 CLI stdout,也不要笼统删除所有 `path` 字段破坏安全相对产物证据。
- 验证:CLI serializer 单测覆盖顶层、嵌套和数组结果;真实 `--agent-runtime-status` 输出对 disposable 项目路径命中为 0,后续完整 `llm-runtime` 报告的 `projectPathTranscriptLeakCount / projectPathReportLeakCount` 必须同时为 0。
### 把模型修复上下文写入公共审计会泄露正文
- 症状:Runtime event 或 `.agent/agent.db` 为了排障直接记录 thinking summary、legacy plan 标题、解析错误、malformed JSON 或 native function arguments;私有任务内容、项目路径或模型调用体因此进入公共审计和 UI 最近事件。
- 原因:格式修复确实需要把上一条输出与错误反馈给同一次 Provider 请求,但“Provider 私有修复上下文”和“持久公共诊断投影”被误当成同一份数据。
- 处理:`thinking_summary` event 只留正文 SHA-256 与字符数,legacy `plan` event 只留步骤数;结构化计划审计只留 explanation 哈希与字符数,以及 step 标题哈希、状态和计数。`agent.runtime.tool_plan.repair` 只留 attempt/maxAttempts、protocol,以及错误、输出或调用体预览、callId/functionName 的哈希与长度。经过过滤和限长的上一条输出与协议错误只可进入当前 planning 的私有 repair 请求,不得落到 event、task 或 Agent DB 正文字段。
- 验证:后台 loop 回归必须断言 thinking event 不含摘要正文、legacy plan event 不含标题;文本与 native repair 回归必须同时证明私有请求仍含足够修复上下文,而 Agent DB 不存在 `protocolError / responsePreview / function arguments / callId / functionName` 原文字段。
### 每次同步 context bundle 时刷新 planning fingerprint,导致同批旧动作越过仓库规范漂移
- 症状:同一 Provider planning 返回多个 actions;前一个验证动作修改了 `AGENTS.md` 或其它启动上下文来源,后一个写动作仍执行成功,下一轮只看到普通 `ok` observation,没有 `repositoryContextDrift=true`。
- 原因:context bundle v3 在 action 激活和 observation 落盘后都会同步完整计划,同时重新扫描 repository startup context。若 drift gate 从最新 bundle 读取 fingerprint,前一个动作造成的漂移会被同步成新基线,后续动作不再与 Provider planning 真正看到的旧规范比较。
- 处理:Provider request builder 必须把实际渲染的 repository fingerprint 和工具计划一起返回,并写入 `game-creator-pending-action.v4` 的 `plannedRepositoryContextFingerprint`;同批自动动作、待确认动作和恢复动作都只复核该持久快照。旧 v1-v3 缺少身份,失败关闭,不能从最新 bundle 猜回。
- 验证:`runtime_v11_closure_repository_context_drift_replans_before_auto_mutations` 必须覆盖 file.write / file.patch / file.delete / project.patchset / project.restore 五种动作,证明前置验证导致规范漂移后旧动作零执行、同 run 收到稳定 drift observation;`legacy_context_and_pending_records_fail_closed` 覆盖 v3 拒绝。
- 关联:`docs/technical/【技术方案】AI游戏创作Agent Runtime V1.1-2026-07-12.md`、`apps/ai-game-creator-shell/src-tauri/src/agent.rs`、`main.rs`、`tests.rs`、`apps/ai-game-creator-shell/src/App.tsx`、`tests/appSurface.test.ts`。
## iOS 退款问询的 result_code 不是 debug 状态
- 现象:为了先观察真实 iOS 退款通知,回调返回 `ErrCode=0 + IosRefundQueryResponse.result_code=1`,并把 evidence 写成“调试阶段不执行自动退款决策”,看起来像安全 ACK,实际已经向微信建议拒绝退款。
@@ -686,8 +686,81 @@ V1.16 把正式用户主聊天从一次性自然语言问答升级为现有 Exte
2026-07-14 修复后真实 Provider 验收已通过。一次性项目 `/tmp/gameagent-supervisor-parallel-e2e-pass-*` 中,父 run `swarm-project-supervisor-1784044475952` 同时创建 `design-director` 与 `art-director` 两个静态委派;两者均在时间戳 `1784044486` 进入 running,分别于 `1784044506`、`1784044536` completed,存在 20 秒真实重叠。父 run 只写入 1 条 `waiting-for-delegate-receipts` task 记录,期间没有继续 Provider 轮询;随后同一 actionId 认领两份 delivery,两个 delivery 均为 `claimed-by-parent`,唯一 claim 为 `Observed` 且 receiptCount=2,父 run 无 reconciliation 并 completed。首轮 Supervisor Session 恰好写入 1 条 user 与 1 条 assistant;同一 Session 的第二轮“基于上文、不要重新委派”请求直接使用历史完成三句回复,delivery 总数仍为 2。项目范围精确 secret 扫描无 API Key、Bearer token 或 `sk-*` 命中。
## V1.17 单 Agent 持久计划
V1.17 把后台单 Agent 每轮临时生成的短 `plan` 升级为同一 run 内可恢复、可单调更新并参与完成门禁的结构化计划。它不新增模型工具、任务系统或项目权限;计划更新仍通过既有 `submit_agent_tool_plan` function arguments 提交,Runtime state 与 context bundle 是持久事实源。
### 工具计划契约
OpenAI Chat / Responses 的 strict function schema 顶层固定为 `thinkingSummary / planUpdate / plan / actions / response`,其中 `planUpdate` 是必填但可为 `null` 的字段:
```json
{
"thinkingSummary": "已完成项目读取,开始实现",
"planUpdate": {
"explanation": "根据真实项目观察推进实现步骤",
"steps": [
{ "step": "读取项目上下文", "status": "completed" },
{ "step": "实现核心玩法", "status": "in_progress" },
{ "step": "运行定向验证", "status": "pending" }
]
},
"plan": [],
"actions": [],
"response": ""
}
```
- `planUpdate.explanation` 必须是非空有界文本;`steps` 必须包含 1 到 8 个标题互不重复的步骤,每个步骤标题也是有界文本。
- Provider 只允许提交 `pending / in_progress / completed`,同一更新至多一个 `in_progress`。持久快照仍可读取历史已有的内部 `failed` 终态以兼容恢复,但它不是 Provider 可提交状态,也不能由外层 run 失败临时生成。
- 复杂任务在首次拆解、真实进度变化、same-run steer 后重审顺序和最终收束时提交 `planUpdate`;没有变化时传 `null`。提交结构化更新时 `plan` 传空数组。
- native function 的 arguments 即使本轮没有计划变化也必须显式包含 `planUpdate: null`;省略字段属于协议错误并进入既有格式修复或失败路径。只有兼容旧实现的文本 JSON 可以省略该字段并继续走 legacy `plan` fallback。
- `plan` 只保留给文本 JSON 兼容协议或旧 Provider 作为 fallback。只要当前 run 已有 `planRevision > 0`,后续 legacy `plan` 不得覆盖结构化计划;同一响应同时带有两者时以 `planUpdate` 为准。
### 单调状态与完成门禁
- `AgentRuntimeState` 持久化 `planRevision / planExplanation / plan / planSteps / activePlanStepIndex`。第一次有效结构化更新把 revision 推到 1;内容或状态真实变化时单调加一,完全相同的幂等更新不增加 revision,拒绝的更新也不改变现有快照。
- 已进入 `completed` 或历史快照中已有 `failed` 的终态步骤必须保留。后续更新即使省略它们,Runtime 也会把终态步骤并回快照;`completed` 不得回退,已有 `failed` 不得由 Provider 改写。合并后仍受 8 步上限约束,超限整次拒绝。
- `activePlanStepIndex` 只对应唯一 `in_progress` 步骤。结构化计划建立后,旧的“按 actions 数组下标激活、完成或重试步骤”辅助逻辑全部失效;工具成功、失败或 action 序号都不能替模型改写结构化进度,Agent 必须根据真实 observation 显式提交下一版 `planUpdate`。
- 只要结构化计划中仍有非 `completed` 步骤,空 actions、非空 response 或恢复中的 prepared finalization 都不能写 assistant、completed 或成功 final。Runtime 返回 `runtime.plan_update` blocker 并在同一 run 继续 planning;普通失败、取消和 reconciliation 仍可按既有失败路径收束,不能伪装成计划成功完成。
- 外层 run 进入 `failed` 或 `budget-exhausted` 时,就结构化计划字段而言,必须原样保留最后一次可信的 `planRevision / planExplanation / plan / planSteps / activePlanStepIndex`。不得把当时的 `pending / in_progress` 机械改写为 `failed`,也不得把 run 失败反推成步骤进度事实。
- `planUpdate` 是 Runtime 控制面元数据,不是白名单工具 action。更新计划不读取或改写 `.agent/policy.json`,不触发 confirm/deny,不推进 project revision,不改变 verification gate,也不改变 pending action fingerprint;真正的文件、命令、验证、委派和提交仍独立经过原有门禁。
### 恢复与 steer
- `.agent/runtime/context-bundles/<agentId>/<runId>.json` schema 升级为 `game-creator-runtime-context-bundle.v3`,新增结构化计划 revision、说明、步骤和 active index 快照。读取 v3 时必须与当前 Runtime state 的完整计划快照一致;revision、步骤、索引或状态不匹配时失败关闭,损坏 Runtime state 投影为 `needs-reconciliation`,不得猜测进度或自动完成。
- `game-creator-runtime-context-bundle.v2` 保持读取兼容:先完成原有身份、task、revision 和 verification gate 校验,再从当前 Runtime state 补入结构化计划字段并按 v3 继续;后续 checkpoint 写 v3。v1 以及缺失既有安全关联的旧记录仍按原规则失败关闭。
- 最终回复 journal 升级为 `game-creator-runtime-finalization.v2`。它在 `prepared` 时绑定最终可完成的完整计划快照与 `planSnapshotFingerprint`,并把该指纹纳入 `finalizationId`;读取 v2 时还必须重新确认结构化步骤全部为 `completed` 且不存在 active index,不能让内部一致但未完成的篡改快照越过完成门禁。恢复不能只凭回复、task record 或 context bundle 猜测最终计划。
- assistant 已按稳定 `messageId` 落入 conversation、但 Runtime state 随后缺失或不可读时,恢复可从唯一 task record 重建同一 run,再用可信的 v2 journal 恢复原 `planRevision / planExplanation / plan / planSteps / activePlanStepIndex` 并补齐 completed 投影,不重新请求 Provider 或重放工具。若 assistant 尚未落盘而 Runtime state 已丢失,则保留 journal 并进入 `needs-reconciliation`,不得仅凭 prepared journal 新写 assistant;已有 state 与 journal 计划快照冲突时同样失败关闭或丢弃过期 prepared 回复回到同 run 重规划。
- Runner 重启、确认续跑和 stale finalization 重规划必须保持同一 Agent/task/Session/run、原 `planRevision`、全部终态步骤和未完成步骤;恢复不能重新从 legacy `plan` 派生进度,也不能因为工具回执已经存在而自动勾选步骤。
- 同一 Provider planning 返回的所有 actions 必须绑定该请求实际渲染的 repository startup fingerprint。pending action schema 升级为 `game-creator-pending-action.v4` 并持久化 `plannedRepositoryContextFingerprint`;后续 action、用户确认和恢复都只用这份 planning 快照做 drift gate,不能从每个 action / observation 后持续刷新的 context bundle 回读。旧 v1-v3 记录缺少该身份,统一失败关闭进入核对。
- same-run steer 继续由 V1.13 丢弃过期 Provider 计划、剩余 actions 或旧最终回复;结构化计划本身不清空。steer observation 明确要求重审未完成步骤,下一版可以调整未完成步骤的标题和顺序,但已完成或失败步骤继续保留,`planRevision` 继续单调递增。
### 展示边界
- 开发 Agent UI、项目内开发面板和 CLI Runtime 状态输出展示有界的完整计划:revision、说明、最多 8 个步骤、每步状态、当前步骤及已有安全 detail。刷新或事件合并只可沿用同一 Agent/Session/run 的上一个快照,切换 run 时不得把旧计划带入新 run。
- 正式用户的 Project Supervisor 聊天只展示紧凑摘要:`已完成数/总数`、当前步骤、等待对象、下一步和当前专业 Agent 协作数量。不得展示 `planRevision`、内部 explanation、完整步骤列表、原始 observation、内部 currentAction 或动态 child 身份。
### 公共审计与私有修复上下文
- `thinking_summary` 公共 Runtime event 只写固定语义摘要以及 `thinkingSummarySha256 / chars`,不写模型摘要正文;legacy `plan` event 只写 `planStepCount`,不写步骤标题。结构化 `plan_update` 的 Agent DB 投影同样只保存 explanation 的 SHA-256 与字符数,以及步骤标题的 SHA-256、状态和数量。
- `agent.runtime.tool_plan.repair` 只保存尝试次数、协议类型,以及协议错误、经过过滤的模型输出或 function call 预览的 SHA-256 与字符数;repair 中的 callId / functionName 也只写哈希。原始模型正文、解析错误、function arguments 或调用体不得进入 event、task 或 Agent DB 公共审计。
- 为了让 Provider 修正格式,同一次 planning 的私有、瞬时 repair 请求可以携带经过统一敏感信息过滤和长度限制的上一条输出预览与协议错误。该上下文只服务当前 Provider 请求,不得反向复制到公共审计或用户可见状态。
### 验收口径
确定性验收至少覆盖:native function 显式 `planUpdate` 与文本 JSON omission 兼容;空说明、重复步骤、未知状态、超过 8 步和多个 `in_progress` 拒绝;幂等更新不增 revision、真实更新单调递增、终态步骤保留和 completed 回退拒绝;工具 action 下标零推进;外层 failed / budget-exhausted 保留最后可信计划;未完成计划阻止 response 与恢复 finalization;损坏 state 进入 reconciliation;context bundle v3 快照一致性、v2 兼容和 v3 mismatch 失败关闭;finalization v2 计划指纹及 assistant 已落盘后的 state 丢失恢复;thinking / legacy plan / repair 公共审计零正文;计划元数据不推进 project revision、verification gate 或权限确认;开发 UI 刷新后仍显示完整 8 步,以及 Supervisor 只显示紧凑摘要。
恢复与 steer 专项必须在同一 run 中先完成至少一个步骤,再分别覆盖 context checkpoint 后强杀 Runner、恢复继续、Provider planning 中接受 steer、旧 actions 零执行和重排未完成步骤。恢复前后 `taskId / sessionId / runId` 必须不变,`planRevision` 不回退,终态步骤不丢失,已有副作用不重放;最后一版所有必要步骤均为 `completed` 后才允许唯一 assistant。
真实 Provider 验收必须使用不提供计划内容、工具顺序或状态迁移配方的 disposable 项目,让模型自行建立至少三步计划、依据真实 observation 更新至少两次、经历一次 same-run steer 和一次 Runner 重启后完成。验收器从 function call arguments、Runtime state、v3 context bundle、task/event/Agent DB、conversation 和副作用计数交叉证明 revision 单调、终态保留、未完成时零 final、最终唯一 assistant、零动作重放、计划元数据零 project revision / policy 变化以及密钥和项目绝对路径零泄漏;开发 UI/CLI 与 Supervisor 摘要另做展示断言。未实际完成这套真实 Provider 门禁前,只能记录“未验收”或外部阻塞,不能把确定性测试外推为 V1.17 PASS。
截至 2026-07-15,确定性门禁已通过:Tauri 单线程全量 690 项中 686 passed / 4 ignored,结构化计划、finalization 恢复和前端竞态定向回归全部通过。真实 `gpt-5.5` `llm-runtime` 连续三轮均在首个 Provider planning POST 返回前因 TLS record-layer failure 进入 failed,尚未产生 function plan、工具动作、Runner kill 或 steer,因此仍未记录 V1.17 真实 Provider PASS。首轮验收器同时发现 CLI `runtimeJson` 暴露 `sessionPath / eventPath / taskPath`;CLI 输出视图移除这三个绝对存储路径后,第三轮项目绝对路径 transcript/report 泄漏计数均为 0。外部请求失败不能替代完整真实门禁,后续 Provider 恢复后必须重跑本节命令。
## 验收命令
- `cargo test --manifest-path apps/ai-game-creator-shell/src-tauri/Cargo.toml structured_plan_ -- --nocapture`
- `cargo test --manifest-path apps/ai-game-creator-shell/src-tauri/Cargo.toml agent_runtime_context_bundle_migrates_v2_and_rejects_v3_plan_mismatch -- --nocapture`
- `npm run ai-game-creator-shell:typecheck`
- `npm run test -- apps/ai-game-creator-shell/tests`
- `cargo test --manifest-path apps/ai-game-creator-shell/src-tauri/Cargo.toml`
File diff suppressed because one or more lines are too long