# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Build, Lint & Test Commands ```bash # Build (debug) cargo build # Build (release with optimizations) cargo build --release # Build (release-min profile: size-optimized LTO) cargo build --profile release-min # Run (debug, starts server on http://localhost:8000) cargo run # Run with Obscura in-process browser (no external binaries needed) cargo run --features obscura-inprocess # Run CLI binary cargo run --bin astroresearch_cli # Run health check tool cargo run --bin health_check # read-only scan cargo run --bin health_check -- --fix # auto-repair # Lint cargo clippy # Format cargo fmt # All tests cargo test # Unit tests only cargo test --lib # Run a specific test cargo test test_name # Coverage (requires cargo-llvm-cov) cargo llvm-cov --fail-under-lines 80 # Frontend (cd dashboard first) npm run dev # HMR dev server on :5173, proxies /api to :8000 npm run build # TypeScript check + Vite production build → dashboard/dist/ npm run lint # ESLint npm run lint:fix # ESLint with auto-fix npm run format # Prettier format all source files npm run format:check # Check formatting without modifying files ``` ## Architecture Overview **Stack**: Rust Axum backend (port 8000) + React/Vite/TypeScript frontend (port 5173 in dev). In production, the Rust binary serves the pre-built `dashboard/dist/` via `ServeDir` and `ServeFile` fallback, so there is a single process. ### Source Layer Map ``` src/ ├── main.rs # Axum server entry: logging, DB pool, migrations, vec0 table, │ # client/service construction, route registration, AppState assembly ├── lib.rs # Config struct + from_env() loading from .env ├── api/ # HTTP handlers, AppState, re-exports via handlers namespace │ ├── mod.rs # AppState (shared state), handler re-exports, permission/event types │ ├── agent.rs # SSE chat_agent endpoint, session CRUD, metrics, audit log │ ├── papers.rs # Search, download, parse, translate, embed, citations, library, export │ ├── catalog.rs # VizieR TAP catalog queries + cone search endpoints │ ├── observation.rs # Observation data search/download (LAMOST/Gaia/SDSS/DESI) │ ├── notes.rs # Highlight/note CRUD │ ├── sync.rs # Meta-sync and asset-batch endpoints │ ├── targets.rs # Target query/associate/extract, RAG chat, figure chat │ ├── auth.rs # Login/logout/auth-check (token-based, session memory) │ ├── permissions.rs # Agent tool permission rule management │ ├── search.rs # Search history endpoint │ ├── error.rs # Unified API error type (status code + user-safe message) │ └── helpers.rs # Shared DB helpers, format conversion, path validation (private) ├── agent/ # ReAct-based research agent (LLM-driven tool-use loop) │ ├── tools/ # AgentTool trait, ToolRegistry, tool implementations per domain │ │ ├── mod.rs # AgentTool trait, ToolContext (session_id, sse_tx, silent, etc.) │ │ ├── filesystem/ # read_file, grep_files, write, edit │ │ ├── astro/ # search_papers, download_paper, parse_paper, rag_search, │ │ │ # query_target, save_note, analyze_image, research/* │ │ ├── ask_user.rs # Permission-gated interactive user question │ │ ├── background.rs# bg_task_run, bg_task_check (async slow tasks) │ │ ├── compress.rs # Context compression trigger │ │ ├── memory.rs # save_memory, load_memory │ │ ├── persist.rs # Data persistence (used by memory/compaction) │ │ ├── skill.rs # load_skill │ │ ├── subagent.rs # Context-isolated sub-agent spawner │ │ ├── team.rs # spawn_teammate, send_teammate_message, etc. │ │ └── todo.rs # todo_write (task tracking) │ ├── runtime/ # ReAct loop core │ │ ├── mod.rs # AgentRuntime: session create/resume → context → ReAct → finalize │ │ ├── executor/ # Step executor + helpers (tool dispatch, parallel tool calls) │ │ ├── streaming.rs # SSE AgentStreamEvent types │ │ ├── streaming_executor.rs # SSE-emitting executor wrapper │ │ ├── session.rs # Session CRUD, DB persist, agent_config │ │ ├── token_budget.rs # Soft/hard token limits, tiktoken counting │ │ ├── permission.rs # Deny > Allow > Ask rule chain, pattern matching │ │ ├── permission_profile.rs # Named profiles (research, readonly) │ │ ├── checkpoint.rs # Rewind/retry state snapshots │ │ ├── circuit_breaker.rs# CompactionCircuitBreaker (cooldown rate limiting) │ │ ├── error_recovery.rs # Error classification + retry strategies │ │ ├── hardline.rs # Terminal conditions (too many denials, loops, etc.) │ │ ├── partitioner.rs # Message window partitioning for compaction │ │ ├── file_cache.rs # ReadFileState — file content dedup cache │ │ ├── system_prompt.rs # System prompt assembly (mode, skills, memory, tools) │ │ ├── finalize.rs # Post-loop cleanup: memory extraction, audit log │ │ └── denial_tracker.rs # Consecutive denial detection │ ├── compact/ # Context compression (micro/auto/manual layers + collapse) │ ├── memory/ # Persistent memory manager: extraction, dedup, decay, age, guardrails │ ├── hooks/ # Lifecycle hooks: registry, dispatch, builtins, matcher, traits │ ├── modes/ # Agent session modes (see Agent Session Modes section below) │ ├── skills.rs # SkillRegistry: hot-loads SKILL.md files from skills/ directory │ ├── subagent.rs # Context-isolated sub-agent runner │ ├── team/ # Multi-agent team: file-based inbox, lead/teammate coordination │ ├── background.rs# BgNotificationQueue for async slow-task notifications │ ├── task_board.rs# Persistent task board (agent_tasks table) │ ├── trajectory.rs# Session trajectory recording for audit/debug │ └── terminal.rs # Escape sequence filter for ANSI-heavy tool outputs ├── clients/ # External API wrappers (pure communication, no business logic) │ ├── llm/ # LlmClient (OpenAI-compatible chat + streaming), EmbeddingClient │ ├── ads.rs # NASA ADS API │ ├── arxiv.rs # arXiv Atom XML API │ ├── cds/ # CDS services (sesame.rs — Sesame name resolver; vizier.rs — TAP queries) │ ├── gaia/ # Gaia archive (TAP + DataLink) │ ├── lamost/ # LAMOST archive (ConeSearch + FITS.gz download) │ ├── sdss/ # SDSS archive (Data Lab TAP + SAS bulk download) │ ├── desi/ # DESI archive (Data Lab TAP + HEALPix coadd SAS) │ ├── vo/ # Virtual Observatory generic client │ └── qiniu.rs # Qiniu cloud storage ├── services/ # Business logic │ ├── search.rs # Unified cross-source search (ADS + arXiv dedup) │ ├── download.rs # PDF/HTML download with anti-bot measures and fallback chain │ ├── parser/ # HTML/PDF → Markdown parsers (A&A, IOP, ar5iv, generic, PDF via MinerU) │ ├── paper/ # Paper lifecycle: model, db (CRUD), ingest (ADS/arXiv→DB), reader (markdown) │ ├── translation.rs # LLM bilingual translation with Trie-based astronomy glossary │ ├── rag.rs # Embedding ingest + vector similarity retrieval + LLM answer generation │ ├── cds/ # CDS business layer: target.rs (Sesame name resolver), vizier.rs (TAP + cache) │ ├── observation/ # Cross-source observation data (spectra, light curves, photometry, images) │ │ ├── mod.rs # 双轴正交架构: Source × ProductType → ObservationFetcher → Registry │ │ ├── types.rs # Source/ProductType/ProductSpec/Artifact/ObservationProduct/Candidate │ │ ├── fetcher.rs # ObservationFetcher trait + cone search cache template methods │ │ ├── registry.rs # ObservationRegistry (fetcher registration, capability discovery) │ │ ├── dispatch.rs # search_observation + download_observation (unified entry points) │ │ ├── cache.rs # observation_cache table read/write + file staging │ │ ├── lamost.rs # LamostSpectrumFetcher │ │ ├── gaia.rs # GaiaSpectrumFetcher + GaiaLightCurveFetcher │ │ ├── sdss.rs # SdssSpectrumFetcher + SdssApogeeFetcher │ │ └── desi.rs # DesiSpectrumFetcher │ ├── batch/ # Meta-sync (ADS bulk harvest) and asset-batch processing engines │ ├── citation.rs # Citation network graph data (references + citations) │ ├── session.rs # Agent session listing, metrics, audit log (for API consumption) │ ├── vision.rs # Image analysis via dedicated vision model │ ├── note.rs # Highlight/note storage and retrieval │ ├── pipeline.rs # Synchronous paper processing pipeline (download→parse→translate→embed) │ ├── chunker.rs # Markdown text chunking for embedding │ ├── query_parser.rs # Advanced search query syntax parser │ └── logging.rs # Pretty console + rolling file logger └── bin/ ├── health_check.rs # Library consistency checker and auto-repair ├── cli.rs # CLI interface └── reparse.rs # Re-parse existing library items ``` ### AppState — Central Shared State All handlers access state via `Arc`. Key fields: - `db: SqlitePool` — SQLite connection pool (5 max connections, foreign keys enforced) - `llm: LlmClient` / `medium_llm` / `fast_llm` — OpenAI-compatible LLM clients (multi-tier) - `vision_llm: Option` — dedicated vision model client - `embedding: EmbeddingClient` — embedding model client - `ads: AdsClient` / `arxiv: ArxivClient` — academic search clients - `vizier: VizierClient` — CDS VizieR TAP catalog query client - `lamost` / `gaia` / `sdss` / `desi` — observation data source clients - `observation_registry: Arc` — cross-source observation fetcher registry - `skill_registry: Arc>` — hot-reloaded agent skills - `memory_manager: Arc` — persistent agent memory (MEMORY.md + decay) - `cancelled_runs: Arc>` — agent cancellation tokens - `session_permission_checkers: Arc>` — per-session permission rules - `pending_questions` / `pending_permissions` — interactive agent-user handshake channels - `harvest_status` / `batch_status` — async batch operation status tracking ### Agent System Design The agent (`src/agent/`) implements a **ReAct** (Thought → Action → Observation) loop: 1. **`AgentRuntime`** (`runtime/mod.rs`) orchestrates the loop: session create/resume → context build → ReAct loop → finalize 2. **Streaming**: LLM response is streamed via SSE (`AgentStreamEvent`) — thought, tool_call (with `id`), tool_result (with `tool_call_id`), text_delta, usage, error, done. Tool calls execute in parallel. 3. **Tools**: Each tool implements `AgentTool` trait (name, description, JSON Schema parameters, execute). Core tools: read_file, grep_files, glob_files, run_bash, file_write, file_edit, search_papers, download_paper, parse_paper, get_paper_content, rag_search, query_target, save_note, todo_write, compress_context, load_skill, subagent, ask_user, save_memory, analyze_image (vision model, only when `LLM_VISION_MODEL` set). Plus background tools (bg_task_run, bg_task_check) and team tools (spawn_teammate, send_teammate_message, team_broadcast, check_team_inbox). 4. **Thinking mode**: `enable_thinking` flag propagates from `AgentChatRequest` → `AgentConfig` → `ToolContext` → `LlmClient::chat_stream`. Only enabled for Qwen/DashScope backends; frontend-controlled via the `thinking` request field. 5. **Tool call ID tracking**: LLM may not return tool_call IDs — `LlmClient` generates UUID fallbacks. `ToolCall` and `ToolResult` SSE events carry matching IDs for precise frontend pairing. 6. **ToolContext** (`tools/mod.rs`): Injected into every tool execution — holds `app_state`, `session_id`, `sse_tx` (for intermediate events), `enable_thinking`, `read_file_state` (file cache for dedup), `silent` (sub-agents skip permission prompts). 7. **Context compression**: Four layers — micro (placeholder replacement), snip (old-message truncation), auto (LLM summarization), aggro_micro (aggressive placeholder). Protected by `CompactionCircuitBreaker`. Transcripts persisted in `agent_messages` table, not filesystem snapshots. 8. **Skills** (`skills.rs`): Two-layer loading — system-reminder lists names (~20 tokens each), LLM calls `load_skill` to inject full SKILL.md content. Skills are organized into subdirectories: `methodology/`, `plotting/`, `presentation/`. 9. **Sub-agents** (`subagent.rs`): `subagent` tool spawns a context-isolated sub-agent with its own ReAct loop. Sub-agent messages (system/user/assistant/tool) are persisted to `agent_messages` with `agent_name` identifier. Returns final summary + activity log. SSE progress forwarded to parent via ToolContext. 10. **Memory** (`memory/`): File-based persistent memory (MEMORY.md). `MemoryManager` handles extraction from conversation, dedup, recency decay, age-based pruning, and guardrails. Tools: `save_memory`, `load_memory` (auto-injected in system prompt). 11. **Teams** (`team/`): File-based inbox directory per session for lead/teammate message passing 12. **Background tasks** (`background.rs`): Slow ops (download, parse) can run async; results inject via `BgNotificationQueue` before next LLM call Environment variables for agent tuning: `AGENT_TOKEN_SOFT_LIMIT` (default 80000) / `AGENT_TOKEN_HARD_LIMIT` (default 100000). ### Database SQLite via `sqlx::sqlite`. Migrations in `migrations/` are auto-run on startup (`sqlx::migrate!("./migrations")`). Key tables: `papers`, `citations_references`, `notes`, `agent_sessions`, `agent_messages`, `agent_tasks`, `agent_audit_log`, `paper_chunks_content`, `vizier_cache`, `observation_cache`. Vector embeddings use `sqlite-vec` (`vec_paper_chunks` virtual table, auto-registered before any DB connection). The embedding dimension is controlled by `EMBEDDING_DIM` env var (default 1536). On dimension mismatch, the vec table and chunk content are dropped and recreated. ### Observation Data System (`src/services/observation/`) A cross-source astronomical observation data subsystem with a **dual-axis orthogonal** architecture: **Source** (LAMOST/Gaia/SDSS/DESI) × **ProductType** (Spectrum/LightCurve/Photometry/Image). Each valid combination implements an `ObservationFetcher`, registered in `ObservationRegistry`. New fetchers follow OCP (pure addition, no modification to existing code): 1. Create `{source}.rs`, implement `ObservationFetcher` trait 2. Add `register!(YourFetcher)` in `registry.rs::default()` 3. Add new variants to `Source`/`ProductType` enums in `types.rs` if needed External callers use `dispatch::search_observation()` (metadata only) and `dispatch::download_observation()` (metadata + file download) via `ObservationRequest`. Results are cached in `observation_cache` table with file staging to `library/`. ### Frontend (dashboard/) React 19 + TypeScript + Vite + Tailwind CSS 4. Organized by technical layer (Type-Based): - `pages/` — Page-level panel view components (SearchPanel, LibraryPanel, ReaderPanel, CitationPanel, SyncPanel, ResearchAgentPanel, ObservationPanel, SettingsPanel) - `components/` — Reusable and layout components (sub-folders: agent, reader, sync, layout, dialogs) - `hooks/` — Global and feature-specific custom stateful React Hooks (e.g., useLibrary, useSearch, useNotes, useObservation) - `types/` — Global TypeScript type definitions (types/index.ts) - `utils/` — Common utility helper functions - `assets/` — Static assets and global stylesheet styles Dependencies: `react-markdown` + `rehype-katex` + `remark-math` for Markdown/LaTeX rendering, `framer-motion` for animations, `lucide-react` for icons. ### Obscura In-Process Browser The `obscura-inprocess` feature compiles `obscura-browser` and `obscura-net` from `libs/obscura/` directly into the binary, eliminating the need for external browser binaries. This is used for bypassing Cloudflare/WAF on PDF download. Enabled via `--features obscura-inprocess`. ### Astronomy Glossary (dictionary.txt) A 1MB+ bilingual astronomy terminology file loaded at startup into a Trie tree for longest-match glossary construction, used by the translation service to guide LLM translations with domain-accurate term mappings. ### Agent Session Modes (`src/agent/modes/`) Session-level modes configure the agent's identity, tool access, step limits, thinking, and permissions at creation time. Modes are pure-data constants defined at compile time — adding a new mode requires no logic changes. Three built-in modes (registered in `ModeRegistry::builtins()`): | Mode | ID | Tools | Max Steps | Thinking | Permission Profile | | ------------ | --------------------- | ---------------------------------- | ----------- | --------------- | ------------------ | | 通用科研助手 | `default` | All (unrestricted) | 8 (default) | user-controlled | none | | 深度研究 | `deep-research` | All | 16 | forced on | `research` | | 文献阅读助手 | `literature-reader` | Allowlist (read-only + literature) | 6 | forced off | `readonly` | Key types in `modes/mod.rs`: - **`AgentMode`** — static definition struct: id, name, description, icon, identity/principles overrides, extra sections, `ToolSet`, `ModeConfig` - **`ToolSet`** — `All`, `Allowlist(&[&str])`, or `Except(&[&str])` — constrains which tools are available - **`ModeConfig`** — optional overrides for `max_steps`, `enable_thinking`, `tool_timeout_secs`, `permission_profile` - **`ModeRegistry`** — holds `&'static AgentMode` references; `get(id)` for lookup The mode is stored in the `agent_sessions.mode` column (migration `20260624000000_add_session_mode.sql`, defaults to `'default'`). The frontend `ResearchAgentPanel` exposes mode selection. **Mode vs Skill**: Modes are session-level ("who am I"), Skills are task-level ("how do I do X"). Modes affect initialization; the ReAct loop itself is mode-agnostic. ### Permission System (`src/agent/runtime/permission.rs`) Tool execution is gated by a priority-ordered rule chain: **Deny > Allow > Ask** (first match wins). Rules support content-level pattern matching (e.g., `run_bash(rm *)`). **Permission modes** (set via `AGENT_PERMISSION_MODE` env or mode's `permission_profile`): - `default` — full rule chain evaluation - `accept_edits` — auto-allow `file_write`/`file_edit` within the working directory - `bypass` — skip all Ask checks (Deny rules still enforced) - `dont_ask` — convert all Ask to Deny **Permission profiles** (`src/agent/runtime/permission_profile.rs`) are named presets loaded by `AGENT_PERMISSION_MODE` or a mode's `permission_profile` field. Profiles `research` and `readonly` are used by deep-research and literature-reader modes respectively. **Configuration** (in `.env`): - `AGENT_PERMISSIONS_DENY` — comma-separated deny rules (e.g., `run_bash(rm *),run_bash(sudo *)`) - `AGENT_PERMISSIONS_ALLOW` — comma-separated allow rules - `AGENT_PERMISSIONS_ASK` — comma-separated ask rules - `AGENT_PERMISSION_MODE` — default/accept_edits/bypass/dont_ask ### Multi-Tier LLM Configuration The system supports three LLM tiers with cascade fallback: | Tier | Env Prefix | Purpose | Fallback | | ------- | --------------- | ----------------------------------- | ------------------ | | Primary | `LLM_` | Main agent reasoning | — | | Medium | `LLM_MEDIUM_` | Translation, RAG | Primary LLM config | | Fast | `LLM_FAST_` | Memory extraction, background tasks | Medium LLM config | Each tier has `_API_KEY`, `_API_BASE`, `_MODEL` variants. Unset tiers cascade to the next tier down. **Fallback chain**: `LLM_FALLBACK_CHAIN` (comma-separated model names) provides automatic model rotation on repeated 529 errors. `LLM_FALLBACK_MODEL` is a single backup model tried before the chain. ### Vision Model (`analyze_image` tool) When `LLM_VISION_MODEL` env is set, the `analyze_image` tool (`src/agent/tools/astro/analyze_image.rs`) is registered. It delegates image analysis to a dedicated vision model, enabling the main agent to use a text-only model while still processing images. Supports local paths (relative to library dir) and HTTP(S) URLs. Results stream via SSE `TextDelta` events. Also supports `LLM_VISION_API_KEY` and `LLM_VISION_API_BASE` (fall back to primary LLM config). ### Auto Memory Extraction Controlled by env vars: - `EXTRACT_MEMORY_ENABLED` (default `false`) — enables automatic memory extraction at session end and during compaction - `EXTRACT_MEMORY_THROTTLE_TURNS` (default `3`) — extract every N turns - `EXTRACT_MEMORY_MAX_STEPS` (default `3`) — sub-agent max steps for extraction ### Deployment (`deploy.sh`) The `deploy.sh` script builds the project locally (release-min profile, Composer-based) and deploys to a remote server via rsync + SSH. It handles: source compilation, frontend build, remote directory sync, service restart via docker compose, and health check verification. ## Code Conventions - Use `anyhow` for application errors, `thiserror` for library-style typed errors - Environment variables via `dotenvy` + `std::env::var`, with defaults in `Config::from_env()` - SQL queries use parameterized bindings (`sqlx::query("...").bind(...)`) — never string interpolation - API handlers take `State(Arc)` and return Axum-compatible responses - Agent tools implement `AgentTool` trait; new tools register in `ToolRegistry::new()` - Front-end build is triggered by `build.rs` (auto npm install + build when `dashboard/src/` changes) - New observation fetchers follow OCP: add module + register — don't modify existing fetchers or dispatch - External API clients (`src/clients/`) are pure communication; business logic lives in `src/services/` - Route registration happens in `src/main.rs`; handler re-exports are maintained in `src/api/mod.rs` → `handlers` namespace - Agent hooks (`src/agent/hooks/`) implement lifecycles: PreToolUse, PostToolUse, Stop, etc. — register in `builtins.rs`