# Backend domain handoff (projects / datasets / runs / ai / jobs + db.rs migrations) Worker scope: server/src/projects.rs, datasets.rs, runs.rs, ai.rs, jobs.rs, db.rs (migrations only). Used fixed contract: `state.rs` `Cx = axum::extract::State>`, `AppState.with_db` one-lock closures, `S = Arc`. No nested locks, no awaits inside DB closures. ## What was repaired ### projects.rs - Handlers on `Cx` (State extractor); all DB closures return `Rows`/`Result` types explicitly; `rows.next()?` loop pattern (no `MappedRows` collect ambiguity). - `put_draft` optimistic generation guard returns 409 `stale_generation` with `current_generation` details; update+read run inside ONE `with_db`. - Versions immutable; run snapshots reuse identical `hash` versions from the same project; restore creates new draft (`draft_generation + 1`), never rewrites history; `source` in {manual, run, ai, restore} enforced by schema CHECK. ### datasets.rs - Request validation: 1–5 instruments, daily only, none|qfq|hfq, field whitelist (open/high/low/close/volume/adj_factor), ISO dates, ≤15y, duplicate-instrument rejection, **index + qfq/hfq explicitly rejected** ("indexes have no adjustment factors"). - Cache key = sha256(canonical_request) over instrument identities + range + fields + freq + adjustment; dataset display name excluded (identity is semantic). - Client responses strip internal `path` fields and `preview` from manifests (`client_manifest`); no host paths leak (QA check `no server paths in manifest`). - `assert_owned` owner check for every dataset read/preview; `load_manifest` requires status='ready'. ### runs.rs - POST /runs: dataset must exist, be owned by caller, be ready; explicit warning acknowledgement (`acknowledge_warnings`) required when manifest warnings non-empty; **index-feed restriction surfaced as an explicit warning requiring acknowledgement**; numeric bounds (capital 1..1e12, commission/slippage 0..0.05, parameters size cap); per-user daily quota enforced transactionally inside one `with_db` closure (`BEGIN IMMEDIATE` + count + insert). - Draft snapshot pinned to run (`source='run'`); rerun requires `use_original_data:true`, pins original version/config/manifest_hash, never refetches ("no secret fetching latest"), verifies the original dataset still exists and is ready/owned. - cancel → `jobs::signal_cancel` kills the specific container by name; **terminal states guarded by `AND status='running'`/`'queued'` so failed/cancelled are never overwritten by racing completion**. - cleanup_interrupted marks running→failed on restart, queued stays resumable. ### ai.rs - Owner-scoped project fetch (`AND user_id=?`); requires account `ai_enabled` and POC flag `ai_enabled_poc`; `expected_generation` guard. - Daily request budget measured from real ledger (ai_requests of today), capped `ai_daily_request_cap`; REAL call to `{AI_BASE_URL}/chat/completions` (model `glm-5.3-flash`), key only from `OPENCODE_GO_API_KEY`, 90s timeout, output token cap. - Robust fenced-block parse (`extract_code_block`, first non-empty fenced block; mid-response fences ignored), explanation outside fences; parsing failures recorded (`ai_requests.status='failed', error=...`) — failure ledger retained. - Response records real usage (`prompt_tokens`/`completion_tokens`) into `ai_usage` (no price/cost invented); input token cap enforced before accept. - accept: owner check, base_hash identity + `base_generation` + `expected_generation` all verified inside ONE transaction; writes draft and `source='ai'` version atomically. Renewal after restore always yields new generation. - summarize_dataset only reads ready datasets the user owns and that actually have runs; only columns/heading metadata to the model, not payloads. ### jobs.rs - Dispatcher: bounded dataset fetch concurrency (`fetch_concurrency`), single backtest worker (`run_concurrency=1` POC). - **Exact-source request cache**: new `fetch_cache` table keyed on canonical request (contract: interval coverage; content-derived manifest hash stable across users/datasets; worker not re-run for identical requests). - Concurrent identical cache keys serialize through a process-wide in-flight set; the waiting job rechecks cache periodically (required concurrent same-key serialize/recheck). - Fetch flow: `docker run` data worker (network on, core `run_named`), request.json = canonical stored request; artifacts ingested via core `store::ingest_directory` (symlink/traversal rejection, content-hash dedup into immutable objects); normalized object paths rewritten `objects/{hash-prefix}/{hash}.{ext}`; **raw object verified present by content hash (raw ObjectHash) — raw retained, never dropped**. - Manifest hash recomputed server-side over canonical JSON **excluding fetched_at and path → stable content identity independent of user/request IDs/fetch timestamps**; identical data reused across users yields identical hash (QA `shared immutable data hash`). - Worker JSON `status` gate: dataset/run only succeeds when worker reports ready/succeeded; worker failure/stderr (bounded) surfaced ~1.2KB. - Queued run dequeue rechecks account active + daily budget + dataset ownership/readiness (`BEGIN IMMEDIATE`) then claims queued→running within the same transaction; unreachable due to `db.changes()`-race free guarded update. - Result: `data_manifest_hash` pinned; `sanitize_nonfinite` removes NaN/Inf (metrics can be null); warnings carried; elapsed/peak RSS from worker; engine name/version from worker. - Container output safety per job: `data_dir/jobs/{uuid}` chmod 0700, `output/` chmod 0777 (uid 65534 writable). Backtest mounts: input ro, dataset objects ro at /data, output rw; no secrets; timeout/cancel kills by specific container name (`worker::cancel_container`). ### db.rs migrations (mine) - Added `fetch_cache(key PK, manifest_hash, manifest, object_count, fetched_at)`. - Index `runs_user_daily(user_id, created_at DESC)` for quota counts. ## Domain tests (this module) `cargo test` (all passing; remaining failures live in auth/store/util/admin files run concurrently by other workers): - datasets: `validate_rejects_unsupported` (freq/adjustment/index restriction/duplicates), `cache_key_is_stable_and_shared_regardless_of_display_name`, `client_manifest_strips_internal_paths_and_preview` - jobs: `manifest_hash_is_content_identity_not_fetch_metadata`, `ingest_rewrites_paths_and_hash_is_stable` (synthetic worker output fixture, raw kept, dedup identity stable), `cache_key_treats_request_semantics_not_names`, `nonfinite... via runs`, `signal_cancel_never_rewrites_terminal_states` - runs: `config_bounds_enforced`, `nonfinite_sanitized` - ai: `extract_takes_first_fenced_full_python_block` (robust fenced parse; empty/no-fence rejected), `usage_totals_never_invent_costs` ## Worker protocol notes (for parent/worker owner) - `worker/main.py` referenced `_write_raw(...)` but never defined it — every fetch would crash (NameError). I added the minimal faithful implementation (main.py, worker protocol): writes immutable raw provider response JSON (`{endpoint, params, fetched_at, data}`) to `objects/raw_{slug}_{hash12}.json`, returns (name, sha256-of-bytes). This matches SPEC "raw JSON actual source response" and lets server ingest raw artifacts by content hash. The worker owner should review/align naming if they implement their own variant. - Worker fetch protocol consumed (current main.py): reads `/input/request.json` with `instruments/start_date/end_date/frequency/adjustment/fields`; writes `output/result.json` `{status: ready|failed, manifest{schema_version, normalization_version, fetched_at, frequency, adjustment, objects[...]}, hash, preview{columns,rows,coverage}, warnings, elapsed_ms}`; objects' `path` is worker-relative (`objects/{slug}.csv`) and is rewritten by the backend before storage/manifest display. - Backtest input `{code, config, dataset_manifest{objects[path=/data/]}, data_root:'/data'}` only normalized CSVs mounted; feed names from `instrument.symbol`; result contract per SPEC. - Instrument search: core `worker::search_instruments` already talks to the real `worker.main search` catalog (AKShare). Domain routes read it via main.rs; failures surface as `status:"unavailable: ..."` (no fake success). ## Not known / for parent - qa_api.py needs QA env credentials envs; prepared by parent. - `clean_orphan_containers` restart cleanup lives in main.rs (core); my signal_cancel/cleanup interplay verified for guarded statuses. - Object store retention/cold storage documented in docs/worker.md (worker owner); server keeps immutable referenced objects and runs pin them.