summaryrefslogtreecommitdiff
path: root/docs/overnight-completion-brief.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/overnight-completion-brief.md')
-rw-r--r--docs/overnight-completion-brief.md24
1 files changed, 24 insertions, 0 deletions
diff --git a/docs/overnight-completion-brief.md b/docs/overnight-completion-brief.md
new file mode 100644
index 0000000..25432e4
--- /dev/null
+++ b/docs/overnight-completion-brief.md
@@ -0,0 +1,24 @@
+# Living Village overnight completion brief
+
+## Authorization and scope
+User explicitly requests continued autonomous development until all tasks complete, with a complete independent game as the acceptance target. There is no next-day deadline; continue until the full acceptance gate passes and do not reduce scope or quality because of the old deadline. Never claim Steam-ready commercial completeness from demo tests. No user intervention for ordinary design choices. Chinese reporting. All coding, test authoring and automation scripts MUST be performed by OpenCode; Hermes dispatches, reviews and exercises artifacts.
+
+Repository /home/somhairle/projects/living-village. Read docs/design/living-village-design.md. OpenCode Web http://127.0.0.1:4096 and https://dev.somhairle.bid. Credentials in /home/somhairle/.local/share/opencode-web/access.private.json, never print. All OpenCode dispatches use `providerID=openai`, `modelID=gpt-5.6-luna`, `variant=max`; the earlier ark-plan setting is historical and must not be used. One writer per worktree. Do not change shared service configuration or restart gateways. Keep sessions publicly retrievable; no homepage redirect workaround.
+
+## Baseline
+HEAD at brief creation 557ae9a; ab0cdf4 fixes relation half-life units and true exponential half-life. Independent Release build passed with 2 warnings, 22 tests passed, --batch 1 1 passed. Larger validation not yet run. scripts/m3_supervise.py exists but needs hardening before unattended trust: no file lock/atomic state, no captured child exit status, weak PID identity, only verdict token checked, current binary may be overwritten by later builds. Do not assume its selftests prove lifecycle correctness. Old mixed-chat changes backed up under .opencode-tmp; preserve evidence.
+
+## Work queue
+1. Start frozen baseline M3 validation in isolated artifacts directory, 30 worlds x20 days then 50x100. Use exact seeds and final total checks. Failure triggers diagnosis and OpenCode fix, never threshold weakening. A baseline validation does not substitute for final release regression if kernel changes.
+2. M4a trade model/inventory/money with emergent negotiated prices, regression tests.
+3. M4b memory/rumour propagation, traceable paths and replay UI; validate 2-5 day propagation with actual headless evidence.
+4. M5 player E interaction, intentions/menu, Q observation, needs panel, chronicle, equivalent NPC/player rules. Verify causal chain returns after 3 intermediaries, not fabricated scripted scenario presented as emergence.
+5. M6 save/load including RNG and events, pause/speed, day-night presentation, consistent licensed placeholder art, coherent main menu/new game/load/settings/controls/exit and complete gameplay flow. Respect original no-failure sandbox and 30 NPC scale. Do not invent a different game to meet deadline.
+6. Package Linux runnable release (self-contained where feasible), launch script, README Chinese controls, release notes, asset licenses, screenshots/video. Do not purchase assets or publish Steam releases.
+7. Independently run full build/tests, deterministic save round-trip, live GUI flows including failure cases, final relevant batch checks, and actual 2 hour stability run. Use autoplay rather than unreliable XTEST on headless SDL. Collect screenshot/frame evidence and logs. Publish source to existing authorized self-hosted remote only after inspecting remotes and exclusions; verify exact remote hash.
+
+## Supervision discipline
+Each cron tick reads this brief and current git/API/log/process state, advances a bounded slice, and writes docs/overnight-status.md with real evidence and active session IDs. Never start duplicate writer or repeat expensive baseline runs. Split OpenCode work into small tasks, require no question tool, short reasoning and build per meaningful change. On idle without verdict investigate last messages; bounded retries then fresh narrowed session. Busy is not progress: inspect last tool/time and file/log growth. No long sleeps in cron. Long test processes need durable PID/start-time/exit-code/logs and pinned binary hash. Use up to modest parallel CPU workers for independent worlds if semantics and totals are preserved. Do not lower tick accuracy or change acceptance population to save time.
+
+## Completion gate and delivery
+All named acceptance criteria must have execution evidence, not only worker claims. Write completion-notice.md only after all verified, with artifact paths, controls, exact test counts, actual remaining limitations. If deadline arrives with incomplete scope, write honest incomplete/blocker report and best runnable build, never fake final. Supervisor output stays local; notifier sends only a new terminal notice to origin Slack. Stop supervisor only after exact terminal state verified. Hermes should wire and test notifier script written by OpenCode, no model credential dependence for notification.