summaryrefslogtreecommitdiff
path: root/docs/backend-domain.md
blob: 3544bc90c49b366765229fc56e89472fb9c2354f (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
# Backend domain handoff (projects / datasets / runs / ai / jobs + db.rs migrations)

Worker scope: server/src/projects.rs, datasets.rs, runs.rs, ai.rs, jobs.rs, db.rs (migrations only).
Used fixed contract: `state.rs` `Cx = axum::extract::State<Arc<AppState>>`, `AppState.with_db` one-lock closures, `S = Arc<AppState>`. No nested locks, no awaits inside DB closures.

## What was repaired

### projects.rs
- Handlers on `Cx` (State extractor); all DB closures return `Rows`/`Result` types explicitly; `rows.next()?` loop pattern (no `MappedRows` collect ambiguity).
- `put_draft` optimistic generation guard returns 409 `stale_generation` with `current_generation` details; update+read run inside ONE `with_db`.
- Versions immutable; run snapshots reuse identical `hash` versions from the same project; restore creates new draft (`draft_generation + 1`), never rewrites history; `source` in {manual, run, ai, restore} enforced by schema CHECK.

### datasets.rs
- Request validation: 1–5 instruments, daily only, none|qfq|hfq, field whitelist (open/high/low/close/volume/adj_factor), ISO dates, ≤15y, duplicate-instrument rejection, **index + qfq/hfq explicitly rejected** ("indexes have no adjustment factors").
- Cache key = sha256(canonical_request) over instrument identities + range + fields + freq + adjustment; dataset display name excluded (identity is semantic).
- Client responses strip internal `path` fields and `preview` from manifests (`client_manifest`); no host paths leak (QA check `no server paths in manifest`).
- `assert_owned` owner check for every dataset read/preview; `load_manifest` requires status='ready'.

### runs.rs
- POST /runs: dataset must exist, be owned by caller, be ready; explicit warning acknowledgement (`acknowledge_warnings`) required when manifest warnings non-empty; **index-feed restriction surfaced as an explicit warning requiring acknowledgement**; numeric bounds (capital 1..1e12, commission/slippage 0..0.05, parameters size cap); per-user daily quota enforced transactionally inside one `with_db` closure (`BEGIN IMMEDIATE` + count + insert).
- Draft snapshot pinned to run (`source='run'`); rerun requires `use_original_data:true`, pins original version/config/manifest_hash, never refetches ("no secret fetching latest"), verifies the original dataset still exists and is ready/owned.
- cancel → `jobs::signal_cancel` kills the specific container by name; **terminal states guarded by `AND status='running'`/`'queued'` so failed/cancelled are never overwritten by racing completion**.
- cleanup_interrupted marks running→failed on restart, queued stays resumable.

### ai.rs
- Owner-scoped project fetch (`AND user_id=?`); requires account `ai_enabled` and POC flag `ai_enabled_poc`; `expected_generation` guard.
- Daily request budget measured from real ledger (ai_requests of today), capped `ai_daily_request_cap`; REAL call to `{AI_BASE_URL}/chat/completions` (model `glm-5.3-flash`), key only from `OPENCODE_GO_API_KEY`, 90s timeout, output token cap.
- Robust fenced-block parse (`extract_code_block`, first non-empty fenced block; mid-response fences ignored), explanation outside fences; parsing failures recorded (`ai_requests.status='failed', error=...`) — failure ledger retained.
- Response records real usage (`prompt_tokens`/`completion_tokens`) into `ai_usage` (no price/cost invented); input token cap enforced before accept.
- accept: owner check, base_hash identity + `base_generation` + `expected_generation` all verified inside ONE transaction; writes draft and `source='ai'` version atomically. Renewal after restore always yields new generation.
- summarize_dataset only reads ready datasets the user owns and that actually have runs; only columns/heading metadata to the model, not payloads.

### jobs.rs
- Dispatcher: bounded dataset fetch concurrency (`fetch_concurrency`), single backtest worker (`run_concurrency=1` POC).
- **Exact-source request cache**: new `fetch_cache` table keyed on canonical request (contract: interval coverage; content-derived manifest hash stable across users/datasets; worker not re-run for identical requests).
- Concurrent identical cache keys serialize through a process-wide in-flight set; the waiting job rechecks cache periodically (required concurrent same-key serialize/recheck).
- Fetch flow: `docker run` data worker (network on, core `run_named`), request.json = canonical stored request; artifacts ingested via core `store::ingest_directory` (symlink/traversal rejection, content-hash dedup into immutable objects); normalized object paths rewritten `objects/{hash-prefix}/{hash}.{ext}`; **raw object verified present by content hash (raw ObjectHash) — raw retained, never dropped**.
- Manifest hash recomputed server-side over canonical JSON **excluding fetched_at and path → stable content identity independent of user/request IDs/fetch timestamps**; identical data reused across users yields identical hash (QA `shared immutable data hash`).
- Worker JSON `status` gate: dataset/run only succeeds when worker reports ready/succeeded; worker failure/stderr (bounded) surfaced ~1.2KB.
- Queued run dequeue rechecks account active + daily budget + dataset ownership/readiness (`BEGIN IMMEDIATE`) then claims queued→running within the same transaction; unreachable due to `db.changes()`-race free guarded update.
- Result: `data_manifest_hash` pinned; `sanitize_nonfinite` removes NaN/Inf (metrics can be null); warnings carried; elapsed/peak RSS from worker; engine name/version from worker.
- Container output safety per job: `data_dir/jobs/{uuid}` chmod 0700, `output/` chmod 0777 (uid 65534 writable). Backtest mounts: input ro, dataset objects ro at /data, output rw; no secrets; timeout/cancel kills by specific container name (`worker::cancel_container`).

### db.rs migrations (mine)
- Added `fetch_cache(key PK, manifest_hash, manifest, object_count, fetched_at)`.
- Index `runs_user_daily(user_id, created_at DESC)` for quota counts.

## Domain tests (this module)
`cargo test` (all passing; remaining failures live in auth/store/util/admin files run concurrently by other workers):
- datasets: `validate_rejects_unsupported` (freq/adjustment/index restriction/duplicates), `cache_key_is_stable_and_shared_regardless_of_display_name`, `client_manifest_strips_internal_paths_and_preview`
- jobs: `manifest_hash_is_content_identity_not_fetch_metadata`, `ingest_rewrites_paths_and_hash_is_stable` (synthetic worker output fixture, raw kept, dedup identity stable), `cache_key_treats_request_semantics_not_names`, `nonfinite... via runs`, `signal_cancel_never_rewrites_terminal_states`
- runs: `config_bounds_enforced`, `nonfinite_sanitized`
- ai: `extract_takes_first_fenced_full_python_block` (robust fenced parse; empty/no-fence rejected), `usage_totals_never_invent_costs`

## Worker protocol notes (for parent/worker owner)
- `worker/main.py` referenced `_write_raw(...)` but never defined it — every fetch would crash (NameError). I added the minimal faithful implementation (main.py, worker protocol): writes immutable raw provider response JSON (`{endpoint, params, fetched_at, data}`) to `objects/raw_{slug}_{hash12}.json`, returns (name, sha256-of-bytes). This matches SPEC "raw JSON actual source response" and lets server ingest raw artifacts by content hash. The worker owner should review/align naming if they implement their own variant.
- Worker fetch protocol consumed (current main.py): reads `/input/request.json` with `instruments/start_date/end_date/frequency/adjustment/fields`; writes `output/result.json` `{status: ready|failed, manifest{schema_version, normalization_version, fetched_at, frequency, adjustment, objects[...]}, hash, preview{columns,rows,coverage}, warnings, elapsed_ms}`; objects' `path` is worker-relative (`objects/{slug}.csv`) and is rewritten by the backend before storage/manifest display.
- Backtest input `{code, config, dataset_manifest{objects[path=/data/<name>]}, data_root:'/data'}` only normalized CSVs mounted; feed names from `instrument.symbol`; result contract per SPEC.
- Instrument search: core `worker::search_instruments` already talks to the real `worker.main search` catalog (AKShare). Domain routes read it via main.rs; failures surface as `status:"unavailable: ..."` (no fake success).

## Not known / for parent
- qa_api.py needs QA env credentials envs; prepared by parent.
- `clean_orphan_containers` restart cleanup lives in main.rs (core); my signal_cancel/cleanup interplay verified for guarded statuses.
- Object store retention/cold storage documented in docs/worker.md (worker owner); server keeps immutable referenced objects and runs pin them.