AstroResearch/docs/agent-optimization-analysis.md
Asfmq 49784739fa feat: Agent 全栈升级——模块化重构、Hooks/Skills/Memory/SubAgent/Team 子系统、审计与任务持久化
架构重构:
  - Agent Runtime 由单文件拆为 runtime/ 目录 12 模块(熔断/流式执行/Token预算/文件缓存/权限等)
  - Agent Tools 由单文件拆为 tools/ 目录 20+ 模块(filesystem/astro/memory/skill/subagent/team 等)
  - 解析器体系重构(common.rs 836行变更),各解析器同步升级
  - Download 服务重构(562行),反爬策略强化
  - LLM 客户端重构(266行),流式调用优化

  新子系统:
  - Hooks 生命周期系统(9种事件类型,PreToolUse/PostToolUse 支持输入输出拦截)
  - Skills 双层加载系统(system-reminder 轻量注入 + LoadSkillTool 按需加载,notify 文件监听热更新)
  - Memory 项目记忆管理(类型/提取/去重/衰减/保活/选择策略/护栏 7 模块)
  - SubAgent 上下文隔离子代理运行器(独立 ReAct 循环 + Hook 管道)
  - Team 多智能体团队协作(文件 inbox 通信、lead/teammate 协调)
  - TaskBoard DAG 任务依赖管理
  - Trajectory 会话轨迹、Terminal 终止信号、Autonomous 自主模式、Background 异步通知

  数据库:
  - agent_tasks 表(DAG 依赖模式,blocked_by JSON 数组)
  - agent_audit_log 表(工具调用审计:名称/状态/耗时/输出预览)
  - agent_identity 迁移(消息/审计/任务的 agent_name 归属,agent_team_members 团队注册表)

  API:
  - GET /chat/metrics 聚合指标端点
  - GET /chat/sessions/:id/audit 会话审计查询
  - GET /chat/questions + POST /chat/answer 人机交互问答

  工程:
  - 新增依赖:serde_yaml、notify、glob、walkdir、lru
  - Skills 目录含 methodology/plotting/presentation 三个初始 SKILL.md
  - CLAUDE.md 完整项目架构文档
2026-06-17 00:14:02 +08:00

21 KiB
Raw Blame History

Agent 架构优化分析

对比 Claude Code 源码 (/home/fmq/program/claudecode/src/) 与 AstroResearch Agent (src/agent/) 基于 2026-06-16 的代码快照。


总体评估

我们的 Agent 已经实现了一个功能完整的 ReAct 研究引擎,涵盖了工具注册/调度、三层上下文压缩、生命周期 Hooks、 多 Agent 团队协作、子代理委托、后台任务、Skills 加载等关键子系统。与 Claude Code 的架构范式高度一致。

以下按影响优先级列出可优化领域。


一、CRITICALStreaming Tool Executor

现状

executor.rs 在 LLM 流式响应完全结束后才通过 join_all 并行执行工具。

// 当前流程LLM stream → 收集所有 tool_use blocks → join_all 执行
let results = futures_util::future::join_all(exec_futs).await;

Claude Code 做法

StreamingToolExecutor 在模型仍在生成 tool_use blocks 时就开始调度执行:

模型输出 tool_use(Read file A) → 立即开始读 A
模型输出 tool_use(Read file B) → 立即开始读 B并发安全
模型输出 tool_use(Bash cmd)   → 排队等待(非并发安全)
模型输出结束                   → 此时 A 和 B 可能已完成

优化方案

/// 流式工具执行器 —— 模型还在输出时就开始执行工具
pub struct StreamingToolExecutor {
    tools: Vec<TrackedTool>,
    tool_registry: Arc<ToolRegistry>,
    tool_context: ToolContext,
    max_concurrency: usize,
}

enum ToolStatus {
    Queued,
    Executing,
    Completed,
    Yielded,
}

struct TrackedTool {
    id: String,
    block: ToolCall,
    status: ToolStatus,
    is_concurrency_safe: bool,
    handle: Option<JoinHandle<ToolOutput>>,
    results: Option<Vec<ToolResultMessage>>,
    pending_progress: Vec<ProgressMessage>,
}

impl StreamingToolExecutor {
    /// 模型每输出一个 tool_use block 就调用此方法
    pub fn add_tool(&mut self, block: ToolCall) {
        let is_safe = self.tool_registry
            .get(&block.name)
            .map(|t| t.is_concurrency_safe(&block.args))
            .unwrap_or(false);

        self.tools.push(TrackedTool {
            id: block.id.clone(),
            block,
            status: ToolStatus::Queued,
            is_concurrency_safe: is_safe,
            handle: None,
            results: None,
            pending_progress: vec![],
        });

        tokio::spawn(async { self.process_queue().await });
    }

    /// 非阻塞获取已完成的工具结果
    pub fn get_completed_results(&mut self) -> Vec<ToolResultMessage> {
        // 按顺序 yield 已完成的结果
        // 非并发安全的工具保持顺序
        // 进度消息立即 yield
    }

    /// 等待所有剩余工具完成
    pub async fn get_remaining_results(&mut self) -> Vec<ToolResultMessage> {
        // 等待 executing 的工具完成
        // 然后 yield 所有结果
    }
}

预期收益:大幅降低端到端延迟,尤其是当模型并行输出多个独立的 Read/Search 类工具调用时。


二、HIGH工具并发分区

现状

executor.rs 对所有工具调用一律使用 join_all 并行执行,不考虑工具的并发安全性。

Claude Code 做法

partitionToolCalls() 将工具调用分区为:

  1. 并发安全批次 — 连续的 isConcurrencySafe=true 工具(如 Read、Grep、WebSearch
  2. 串行批次 — 单个 isConcurrencySafe=false 工具(如 Bash、Edit、Write

并发批次用 all() 并行执行max concurrency = 10串行批次逐个执行。

优化方案

/// 工具特征增加并发安全声明
#[async_trait]
pub trait AgentTool: Send + Sync {
    fn name(&self) -> &str;
    fn description(&self) -> &str;
    fn parameters(&self) -> serde_json::Value;

    /// 工具是否可以与其他并发安全的工具同时执行
    fn is_concurrency_safe(&self, _args: &serde_json::Value) -> bool {
        true // 默认只读工具是并发安全的
    }

    /// 最大并发数(默认无限制)
    fn max_concurrency(&self) -> Option<usize> {
        None
    }

    async fn execute(&self, args: serde_json::Value, ctx: &ToolContext) -> ToolOutput;
}

/// 分区工具调用
fn partition_tool_calls(
    calls: &[PreparedCall],
    registry: &ToolRegistry,
) -> Vec<Batch> {
    let mut batches: Vec<Batch> = vec![];
    for call in calls {
        let is_safe = registry.get(&call.tool_name)
            .map(|t| t.is_concurrency_safe(&call.args))
            .unwrap_or(false);

        if is_safe && batches.last().map_or(false, |b| b.is_concurrency_safe) {
            batches.last_mut().unwrap().calls.push(call.clone());
        } else {
            batches.push(Batch {
                is_concurrency_safe: is_safe,
                calls: vec![call.clone()],
            });
        }
    }
    batches
}

预期收益:避免 Bash/Write 等有副作用的工具与其他工具竞争导致的不确定性。


三、HIGHSibling Abort兄弟中止

现状

当一个工具执行出错时,其他并行执行的工具继续运行,浪费资源。

Claude Code 做法

StreamingToolExecutor 中:

  • 当 Bash 工具出错时,siblingAbortController.abort("sibling_error") 中止所有兄弟 Bash 执行
  • Read/WebFetch 等独立工具的失败不影响其他工具
  • 被中止的工具获得 synthetic error message

优化方案

/// 在并行执行时注入 sibling abort 信号
pub struct SiblingAbortController {
    abort_sender: tokio::sync::broadcast::Sender<SiblingAbortReason>,
    errored_tool_description: Arc<tokio::sync::Mutex<Option<String>>>,
}

enum SiblingAbortReason {
    SiblingError { description: String },
    UserInterrupted,
    StreamingFallback,
}

impl SiblingAbortController {
    /// 当工具出错时调用,如果是 Bash 类工具则通知所有兄弟
    pub fn notify_error(&self, tool_name: &str, tool_desc: &str) {
        if is_bash_like_tool(tool_name) {
            let _ = self.abort_sender.send(SiblingAbortReason::SiblingError {
                description: tool_desc.to_string(),
            });
        }
    }

    /// 每个工具执行前检查是否已被兄弟中止
    pub fn check_aborted(&self, this_tool: &str) -> Option<String> {
        // 如果是本工具报的错,不生成 synthetic error避免重复
    }
}

预期收益:避免无效的后续执行,减少等待时间和 API 调用浪费。


四、HIGHError Recovery Ladder错误恢复阶梯

现状

compact.rs 只有主动压缩(在达到 token 限制前触发。如果压缩不够激进413 prompt_too_long 错误会直接暴露给用户。

Claude Code 做法

query.ts 实现了多层恢复阶梯:

第1层Context Collapse drain便宜commit 已 staged 的 collapse
   ↓ 失败/不可用
第2层Reactive Compactfork agent 摘要整个对话)
   ↓ 失败
第3层Max Output Tokens Escalate临时提升到 64k token cap
   ↓ 再次命中
第4层Multi-turn Recovery注入 meta message继续对话
   ↓ 全部失败
最终Surface the error暴露给用户

每层都有 hasAttempted 守卫防止无限重试autocompact 有 circuit breaker连续 3 次失败后停止)。

优化方案

/// 错误恢复策略枚举
enum RecoveryStrategy {
    /// 尝试更激进的 micro_compact
    AggressiveMicroCompact,
    /// LLM 摘要整个对话历史
    ReactiveCompact,
    /// 提升 max_tokens 上限
    MaxTokensEscalate,
    /// 注入 metacognitive 消息
    MultiTurnRecovery,
    /// 放弃,暴露错误给用户
    Surface,
}

struct RecoveryState {
    attempted_micro_compact: bool,
    attempted_reactive_compact: bool,
    attempted_max_tokens_escalation: bool,
    autocompact_failure_count: u32,
}

const MAX_AUTOCOMPACT_FAILURES: u32 = 3;

impl RecoveryState {
    fn next_strategy(&mut self, error: &ModelError) -> RecoveryStrategy {
        match error {
            ModelError::ContextOverflow(_) => {
                if !self.attempted_micro_compact {
                    self.attempted_micro_compact = true;
                    return RecoveryStrategy::AggressiveMicroCompact;
                }
                if !self.attempted_reactive_compact
                    && self.autocompact_failure_count < MAX_AUTOCOMPACT_FAILURES
                {
                    self.attempted_reactive_compact = true;
                    return RecoveryStrategy::ReactiveCompact;
                }
                if !self.attempted_max_tokens_escalation {
                    self.attempted_max_tokens_escalation = true;
                    return RecoveryStrategy::MaxTokensEscalate;
                }
                RecoveryStrategy::MultiTurnRecovery
            }
            _ => RecoveryStrategy::Surface,
        }
    }
}

预期收益:显著提高长对话的鲁棒性,减少用户遇到的 "context too long" 错误。


五、MEDIUMTime-Based Microcompact

现状

我们的 micro_compact 只基于消息数量/大小触发,不考虑时间因素。

Claude Code 做法

microcompactMessages() 首先检查 evaluateTimeBasedTrigger()

  • 计算距上一条 assistant 消息的时间间隔
  • 如果超过配置阈值(如 5 分钟),服务器的 prompt cache 已经过期
  • 此时直接 content-clear 旧的 tool results保留最近 N 个)
  • 因为 cache 已冷,修改消息内容不会有额外代价

优化方案

pub struct TimeBasedMCConfig {
    pub enabled: bool,
    /// 触发阈值(分钟)
    pub gap_threshold_minutes: u64,
    /// 保留最近 N 个工具结果
    pub keep_recent: usize,
}

impl Default for TimeBasedMCConfig {
    fn default() -> Self {
        Self {
            enabled: true,
            gap_threshold_minutes: 5,
            keep_recent: 4,
        }
    }
}

/// 检查时间触发是否应激活
fn evaluate_time_based_trigger(
    messages: &[ChatMessage],
    config: &TimeBasedMCConfig,
) -> Option<TimeBasedTrigger> {
    let last_assistant = messages.iter()
        .rev()
        .find(|m| m.role == MessageRole::Assistant)?;

    let elapsed = last_assistant.timestamp.elapsed().unwrap_or_default();
    let gap_minutes = elapsed.as_secs() / 60;

    if gap_minutes >= config.gap_threshold_minutes {
        Some(TimeBasedTrigger { gap_minutes })
    } else {
        None
    }
}

预期收益:在长时间闲置后自动清理过期上下文,防止用户回到对话时遇上 context overflow。


六、MEDIUM增强的 Hook 系统

现状

5 个生命周期事件,简单的 HookAction::Continue/Block 二元决策。

Claude Code 做法

27 个 hook 事件4 种 hook 类型command/prompt/agent/http结构化 JSON 协议:

{
  "continue": true,
  "decision": "approve",
  "reason": "...",
  "systemMessage": "...",
  "hookSpecificOutput": {
    "hookEventName": "PreToolUse",
    "permissionDecision": "allow",
    "updatedInput": { ... },
    "additionalContext": "..."
  }
}

优化建议(按价值排序)

  1. PreToolUse 输入修改Hook 可以修改工具参数后再执行(如自动修正 bibcode 格式)
  2. PostToolUse 输出修改Hook 可以后处理工具结果(如自动翻译、格式化)
  3. additionalContext 注入Hook 可以向 LLM 注入附加上下文
  4. SessionStart watchPaths:启动时注册文件监控路径
  5. SubagentStart/Stop:子代理生命周期事件
/// 增强的 PreToolUse 输出
pub struct PreToolUseHookOutput {
    pub decision: HookDecision,
    pub updated_input: Option<serde_json::Value>,
    pub additional_context: Option<String>,
    pub system_message: Option<String>,
}

pub enum HookDecision {
    Allow,
    Deny { reason: String },
    Ask { reason: String },
}

/// 增强的 PostToolUse 输出
pub struct PostToolUseHookOutput {
    pub updated_output: Option<String>,
    pub additional_context: Option<String>,
}

七、MEDIUMPermission Pipeline权限管道

现状

工具没有权限系统。所有工具对 Agent 都同样可用。

Claude Code 做法

多层权限评估管道:

Step 1: Deny rule 匹配        → deny不可覆盖
Step 2: Ask rule 匹配         → ask除非 sandbox override
Step 3: Tool.checkPermissions  → 工具自身逻辑
Step 4: Safety checks          → ask绕过免疫
Step 5: bypassPermissions 模式 → allow
Step 6: Allow rule 匹配        → allow
Step 7: 默认                   → ask

权限规则格式:ToolName(pattern:*),支持多源优先级链。

优化方案

/// 权限行为
pub enum PermissionBehavior {
    Allow,
    Deny,
    Ask,
}

/// 权限规则来源(优先级从高到低)
pub enum RuleSource {
    Policy,      // 企业策略
    UserSettings,
    ProjectSettings,
    LocalSettings,
    CliArg,
    Session,
}

/// 权限规则
pub struct PermissionRule {
    pub source: RuleSource,
    pub behavior: PermissionBehavior,
    pub tool_pattern: String,      // "search_papers" 或 "bash(git *)"
    pub content_pattern: Option<String>,
}

/// 权限检查器
pub struct PermissionChecker {
    rules: Vec<PermissionRule>,
}

impl PermissionChecker {
    pub fn check(
        &self,
        tool_name: &str,
        input: &serde_json::Value,
    ) -> PermissionDecision {
        // 1. 检查 deny 规则(不可覆盖)
        // 2. 检查 ask 规则
        // 3. 工具自身 check_permissions
        // 4. 安全检查
        // 5. bypass 模式
        // 6. allow 规则
        // 7. 默认 ask
    }
}

八、MEDIUMProgress Streaming进度流式传输

现状

工具执行期间没有任何进度反馈,直到执行完成才发送结果。

Claude Code 做法

StreamingToolExecutor 支持工具的进度消息立即 yield即使工具还在执行中。进度消息类型为 "progress",在 UI 中显示为短暂的状态更新。

优化方案

/// 为长时间运行的工具增加进度回调
#[async_trait]
pub trait AgentTool: Send + Sync {
    // ... 现有方法 ...

    /// 带进度回调的执行(默认委托给 execute
    async fn execute_with_progress(
        &self,
        args: serde_json::Value,
        ctx: &ToolContext,
        progress: mpsc::UnboundedSender<ProgressUpdate>,
    ) -> ToolOutput {
        let _ = progress; // 默认忽略
        self.execute(args, ctx).await
    }
}

/// 进度更新
pub struct ProgressUpdate {
    pub tool_call_id: String,
    pub message: String,
    pub percentage: Option<u8>,
}

适用工具download_paper下载进度、parse_paper解析进度、search_papers搜索进度


九、MEDIUMToken Budget Management

现状

没有 token 预算跟踪。Agent 可以无限制地消耗 tokens。

Claude Code 做法

  • budget.total — 用户设定的 token 预算上限
  • budget.spent() — 当前已消耗的 output tokens
  • budget.remaining() — 剩余可用 tokens
  • 硬上限:达到 total 后 agent() 调用会抛错
  • 软上限:在接近限制时注入 nudge 消息提醒模型

优化方案

pub struct TokenBudget {
    total: Option<u64>,
    spent_output_tokens: u64,
    spent_input_tokens: u64,
}

impl TokenBudget {
    pub fn new(total: Option<u64>) -> Self { ... }

    pub fn record_usage(&mut self, input: u64, output: u64) {
        self.spent_input_tokens += input;
        self.spent_output_tokens += output;
    }

    pub fn remaining(&self) -> Option<u64> {
        self.total.map(|t| t.saturating_sub(self.spent_output_tokens))
    }

    /// 在接近限制时生成提醒消息
    pub fn nudge_message(&self) -> Option<String> {
        if let (Some(total), Some(rem)) = (self.total, self.remaining()) {
            if rem < total / 10 {
                Some(format!(
                    "注意token 预算已使用 {:.0}%,剩余约 {} tokens。请尽快给出最终答案。",
                    (self.spent_output_tokens as f64 / total as f64) * 100.0,
                    rem
                ))
            } else {
                None
            }
        } else {
            None
        }
    }
}

十、LOW-MEDIUMInterrupt Behavior中断行为分类

现状

取消信号对所有工具一视同仁。

Claude Code 做法

每个工具声明 interruptBehavior()

  • "cancel" — 用户中断时立即取消(如 Read、Search
  • "block" — 用户中断时继续执行完毕(如 Edit、Write防止文件损坏

优化方案

pub enum InterruptBehavior {
    /// 用户中断时立即取消(默认,适合只读工具)
    Cancel,
    /// 用户中断时继续执行到完成(适合写入工具)
    Block,
}

#[async_trait]
pub trait AgentTool: Send + Sync {
    // ... 现有方法 ...

    fn interrupt_behavior(&self) -> InterruptBehavior {
        InterruptBehavior::Cancel // 默认安全取消
    }
}

十一、LOWCircuit Breaker 模式

现状

compact 失败没有熔断机制,可能无限重试。

Claude Code 做法

  • MAX_CONSECUTIVE_AUTOCOMPACT_FAILURES = 3
  • fallback model 调用遇到 529 Overloaded 时切换到备选模型
  • 所有恢复策略都有 hasAttempted 守卫

优化方案

pub struct CircuitBreaker {
    max_failures: u32,
    failure_count: u32,
    state: CircuitState,
}

enum CircuitState {
    Closed,       // 正常工作
    Open,         // 熔断,拒绝请求
    HalfOpen,     // 试探性恢复
}

impl CircuitBreaker {
    pub fn check(&mut self) -> Result<(), CircuitOpenError> {
        match self.state {
            CircuitState::Open => Err(CircuitOpenError),
            CircuitState::HalfOpen | CircuitState::Closed => Ok(()),
        }
    }

    pub fn record_success(&mut self) {
        self.failure_count = 0;
        self.state = CircuitState::Closed;
    }

    pub fn record_failure(&mut self) {
        self.failure_count += 1;
        if self.failure_count >= self.max_failures {
            self.state = CircuitState::Open;
        }
    }
}

十二、LOWContext Collapse / Projection System架构级

现状

压缩直接修改消息数组,原始上下文永久丢失。

Claude Code 做法

ContextCollapse 采用 commit log + projection 模式:

  1. 不再直接修改消息
  2. 将旧的上下文段替换为摘要 + metadata
  3. 摘要存储在独立的 collapse store 中
  4. 每次查询循环入口通过 projectView() 重放 commit log 重建临时消息视图
  5. Commit 是 staged先暂存再 committedon overflow

适用场景

我们的论文研究场景中Agent 可能会在同一个 session 中研究多篇论文。Context collapse 可以在切换论文时保留之前的研究摘要而不是完全丢弃。

优化方案

/// 上下文段
pub struct ContextSegment {
    pub id: String,
    pub summary: String,
    pub original_message_count: usize,
    pub original_token_estimate: usize,
    pub created_at: chrono::DateTime<chrono::Utc>,
}

/// Collapse 存储
pub struct CollapseStore {
    segments: Vec<ContextSegment>,
    commit_log: Vec<CollapseCommit>,
}

impl CollapseStore {
    /// Stage 一个 collapse还不提交
    pub fn stage(&mut self, segment: ContextSegment) { ... }

    /// Commit 所有 staged collapses
    pub fn commit_staged(&mut self) -> usize { ... }

    /// 重放 commit log生成当前消息视图
    pub fn project_view(
        &self,
        recent_messages: &[ChatMessage],
    ) -> Vec<ChatMessage> { ... }
}

实施优先级建议

优先级 优化项 预计工作量 收益
P0 Streaming Tool Executor 3-5 天 延迟大幅降低
P1 工具并发分区 1-2 天 正确性提升
P1 Error Recovery Ladder 2-3 天 鲁棒性大幅提升
P1 Sibling Abort 1 天 资源浪费减少
P2 Time-Based Microcompact 1 天 长会话体验
P2 Token Budget Management 1-2 天 成本控制
P2 Progress Streaming 1-2 天 UX 提升
P3 增强 Hook 系统 2-3 天 可扩展性
P3 Permission Pipeline 2-3 天 安全性
P3 Interrupt Behavior 0.5 天 可靠性
P4 Circuit Breaker 1 天 稳定性
P4 Context Collapse 5-7 天 长期架构

架构范式已对齐的部分

以下方面我们的实现已经与 Claude Code 的范式高度一致,无需大幅改动:

  1. ReAct Loop 结构Thought → Act → Observe
  2. Tool trait + Registry 模式(虽可细化但已完整)
  3. 三层上下文压缩micro/auto/aggressive
  4. 安全切割点(不切断 tool_call/tool_result 配对)
  5. Transcript 持久化(压缩前保存)
  6. 生命周期 Hooks有基础的 5 事件)
  7. 重复调用检测DuplicateDetector
  8. Skills 两层加载系统
  9. 子代理委托SubAgentRunner
  10. 后台任务系统BgNotificationQueue + mpsc
  11. 多 Agent 团队TeamManager + file inbox
  12. SSE 流式事件到前端
  13. Todo/Task 持久化到 SQLite
  14. Session 生命周期管理