Change 1 replaced the v3 monolith prompt with a thin dispatcher overprompt/*.md+templates/and moved continuity arithmetic intopipeline/story_state.py. The re-based soak shows the refactor doing exactly what it was built to do. The headline result is the carryover queue: all 9 surviving entries matchoriginal_rank × 0.85^ageto within 0.02, with zero entries past the age-5 or rank-2.0 drop rules. On 2026-07-16 the shadow-verify caught v3 keeping an age-6 carryover with non-0.85 arithmetic; moving that arithmetic out of the prompt and into tested Python eliminated the drift class entirely. Hold-queue promotions are equally clean: 8 promotions, every one at exactly a 2-day gap, none late, none MENA.
Two Minor defects surfaced, neither reaching the TUNE bar and neither mapping to any of the three knobs guides/C1_TRIAGE_GUIDE.md permits: a scoring.log schema omission on one day, and one same-entity continuation published as fresh rather than as an Update:. Both are routed to S6 as follow-ups. A third item is bookkeeping, not a soak defect: the live SKILL.md has drifted from its repo mirror.
CHANGE1_ASSESSMENT_PROMPT.mdThis assessment ran interactively and retrospectively, not as the registered one-shot task. The prompt's dates are hardcoded to the original window; all were shifted:
| Prompt says | Actually used | Why |
|---|---|---|
| Soak 2026-07-17 → 07-23 (7 days) | 2026-07-24 → 07-28 (5 days) | Original window was mostly dead — zero briefings 07-17→07-21 (laptop off, operator-confirmed 2026-07-26). Re-based per HANDOFF + Status Log 2026-07-26. |
Verdict file change1-2026-07-24.md | change1-2026-07-30.md | Honest run date. No change1-*.md existed, so the idempotency guard did not trip. |
| "N / 7" denominators | N / 5 | Window length. 5 briefings is exactly the prompt's own <5 → EXTEND floor, so EXTEND does not apply. |
| Steps 6–8 (deploy, Telegram, email) | Step 6 only | Telegram and email exist to notify an absent operator that a scheduled task fired. This ran attended, so only the deploy (a stable URL beside the prior verdicts) was run, at the operator's direction. |
Evidence limitation, stated up front: the three state files are gitignored and are overwritten every run, so only their 2026-07-30 contents survive. Per-day state was reconstructed from story_memory.json (7-day retention, covers 07-23→07-30), scoring.log, synthesis_depth.log, and the briefings' own <!-- Story N: id (provenance) --> comments. One question could not be closed from surviving artifacts — see Spot-check D.
| # | Criterion | Result | Evidence |
|---|---|---|---|
| 1 | Phase 4a+4b regression | ✓ partial | 0 false-positive suppressions · 0 false-positive Updates (11/11 matched_memory_date values resolve to a real same-tag memory entry) · 0 stale carryovers (9/9 decay exact, max age 4, min rank 2.51) · holds surface ≤2 days (8/8 at exactly 2) · state files valid JSON. 1 false-negative match (see M2) — the mirror defect, not enumerated by this criterion. |
| 2 | Markers byte-identical 5/5 | ✓ | Only sanctioned renderings appear; every marker carries class carryover-label and the span count equals the marker count each day (3, 4, 0, 2, 2). 0 markers in the wrong section. 0 promoted badges. imp-red / badge-alert absent 5/5 (S1 rider b holds). ✓ |
| 3 | MENA cap | ✓ | Exactly 2 MENA Top Story cards on all 5 days, 0 violations. Never 0 or 1 — supply exceeds the cap every day, so invariant 6 is load-bearing, not idle. ✓ |
| 4 | commit success 5/5 | ✓ | All 3 files rewritten; memory gains that day's entries on 5/5 days (8, 8, 11, 8, 10); all 64 entries carry exactly 12 fields with is_update / alert / matched_memory_date present. No v3-schema residue. ✓ |
| 5 | Briefing + Telegram 5/5 | ✓ | briefing-*.html + site/briefings/*.html + a site/index.html manifest entry present for all 5 days (manifest now 111 briefings), so deploy.sh Step 6 ran and mutated the manifest each day. Telegram graded on the prompt's own stated proxy: no persistent send log exists, no error trace, no operator complaint. |
| 6 | Force-fill ≤2 days | ✓ 0 days | Zero Carried over from markers across all 5 days; card counts 8, 7, 11, 8, 8 — the 7-story floor was met organically every day, including 07-25 which landed exactly on it. No over-suppression signal. ✓ |
| 7 | scoring.log 5/5, sane shape | ✓ partial | One JSONL line per run day, 5/5, present on heavy days. entries respects the importance ≥ 6 filter 5/5 (0 violations across 216 entries). 07-24 omitted the mandated rank field on all 66 entries — see M1. ! |
| 8 | Overrides rare + reasoned | ✓ | 3 in-window needs_judgment-band pairs (J = 0.43, 0.43, 0.45), all 3 resolved correctly as Updates with sensible matched dates. 1 same-entity case resolved the wrong way (M2). Instrumentation is thin — no worksheet is persisted, so this is graded on surviving run evidence. |
| — | Log interfaces (S2) | ✓ | cluster_position present 5/5 with positions exactly 1–5 each day; legacy cluster_rank absent 5/5 (the 2026-07-16 field-drift defect is fixed); fetch_domain present in 10/10 fetch_failures entries. ✓ |
| Check | Result |
|---|---|
| Test suite | 98 passed ✓ |
prompt/*.md ≤ 4096 B | all 6 (scoring.md 4096, corroboration.md 4093 — zero/near-zero headroom) ✓ |
| REGISTRY block in sync | gen_scoring_table.py --check clean ✓ |
| Live SKILL.md is v4 dispatcher | 3965 B, "Work through the steps in order" ✓ |
Literal $RSS leaked into output | 0 occurrences in all 5 briefings (S1 rider a holds) ✓ |
| Bare-domain links | 0 (invariant 4) ✓ |
| HTML tag balance | 0 errors, 0 unclosed, 5/5 ✓ |
| Day | Cards | MENA | Also Noted | Alert cards | Markers | Promotions | Primary found |
|---|---|---|---|---|---|---|---|
| 2026-07-24 | 8 | 2 | 10 | 1 (consolidated) | 3 × Update | 0 | 4/5 |
| 2026-07-25 | 7 | 2 | 12 | 2 | 4 × Update | 0 | 5/5 |
| 2026-07-26 | 11 | 2 | 10 | 4 | 0 | 4 | 5/5 |
| 2026-07-27 | 8 | 2 | 10 | 3 | 2 × Update | 2 | 5/5 |
| 2026-07-28 | 8 | 2 | 10 | 2 | 1 × Update + 1 × Continued | 2 | 5/5 |
Enrichment across the window: 24/25 top-5 clusters located a primary source (96%), against the Phase 1 baseline of 24/35 (69%). 32 new outlets added. Budgets respected — peak 10 WebSearch (exactly at the invariant-8 cap, 07-28) and peak 8 WebFetch.
| Day | Marker text | Section |
|---|---|---|
| 2026-07-24 | Previously covered 2026-07-23 ×3 | Top Stories |
| 2026-07-25 | Previously covered 2026-07-24 ×2, Previously covered 2026-07-23 ×2 | Top Stories |
| 2026-07-26 | (none — 11 fresh cards) | — |
| 2026-07-27 | Previously covered 2026-07-25, Previously covered 2026-07-26 | Top Stories |
| 2026-07-28 | Previously covered 2026-07-26 | Top Stories |
| 2026-07-28 | Continued from 2026-07-22: | Also Noted |
All six formats match prompt/render.md exactly. Every Update: headline prefix is paired 1:1 with a Previously covered line (3↔3, 4↔4, 2↔2, 1↔1). The single Continued from line is correctly Also-Noted-only. Zero Carried over from markers — no carryover reached Top Stories in the window, consistent with criterion 6.
End-of-window state is not recoverable (overwritten). Current pending_stories.json (2026-07-30) holds 3 entries, all first_seen 2026-07-30 / hold_until 2026-08-01 — none past its hold_until, so no promotion-path miss is visible. All 3 are edtech, ranks 4.0–6.4, and all 3 carry original_rss_url.
The window's promotion ledger, from the briefings' provenance comments:
| Promoted on | Story | First seen | Gap |
|---|---|---|---|
| 07-26 | imagi-vibe-coding-funding-jul24 | 07-24 | 2 d |
| 07-26 | additude-prompt-dependency-jul24 | 07-24 | 2 d |
| 07-26 | ets-mastery-transcript-sale-jul24 | 07-24 | 2 d |
| 07-26 | ethan-mollick-ai-guide-jul24 | 07-24 | 2 d |
| 07-27 | ai-classroom-adoption-uae-jul25 | 07-25 | 2 d |
| 07-27 | k12-market-brief-ny-jul25 | 07-25 | 2 d |
| 07-28 | librarians-avoiding-ai-workshops-jul26 | 07-26 | 2 d |
| 07-28 | uae-schools-optimistic-iran-war-disruption-jul26 | 07-26 | 2 d |
8/8 promotions at exactly the 2-day boundary. 0 late. 0 MENA held (all 8 are ai-edu/edtech) — Phase 4b's "0 MENA entries held" criterion holds under v4. The second-source promotion path also fired at least once: 07-27's iran-ukraine-caspian-sea-attack is annotated (fresh, gained 2nd+ source).
Worth recording: is_promoted_from_hold → NO marker is honoured — the provenance appears only in HTML comments, never in rendered text. That comment trail is undocumented in render.md but was the single most useful artifact in this assessment.
0 violations. 7 alert-flagged clusters reached memory across the window (1, 2, 1, 1, 2), and every alert-scored cluster published same-day. All alert clusters were mena 5/5. No alert=true cluster was held or suppressed at any point; no MENA story entered the hold queue at all. ✓
One apparent discrepancy was chased down and cleared: 07-24 scored 2 alert clusters but rendered 1 alert card. Reading the block directly, that card is a consolidated alert carrying both threads — the Trump/Iran "massive attack" decision and the Houthi strikes on two Saudi tankers plus the UK-bases warning, with all four source links. render.md specifies "every alert=true item in the red card" (singular), so one consolidated card is the spec. Nothing was suppressed or delayed.
In-window: no permanent loss found. Every hold-eligible cluster traced either published, promoted within 2 days, or sits in the current pending list.
What could not be closed: stories held on 07-27 and 07-28 were due to promote on 07-29 and 07-30 respectively. There was no 07-29 run (confirmed: no briefing, no scoring.log line, no memory entries), and the 07-30 briefing shows 0 promotions from hold and no card whose id predates 07-30. Two readings are consistent with the surviving evidence:
rank ≥ 4.0 and rss_sources == 1 and alert == false and not carryover/update/promoted); orpending_stories.json was overwritten by the 07-30 run, so reading (1) cannot be confirmed and (2) cannot be excluded. This is an instrumentation gap, not an established defect — and it is the one thing that would have made this assessment decisive. Remediation is cheap and belongs in S6: append a JSONL event per hold/promote/drop, or snapshot the 3 state files per run under a dated path. Aging arithmetic itself is demonstrably robust to a skipped day — carryover ages are date-derived, not increment-derived, and the 07-26/07-27/07-28 entries correctly read age 4/3/2 as of 07-30 across the missing 07-29.
No degradation to invariants-only on any day: all 5 briefings carry enrichment context, continuity markers, and full templated structure. 0 literal $RSS in any briefing or log (the S1 rider-a expansion holds). 0 HTML tag-balance errors 5/5.
But the live SKILL.md has drifted from its repo mirror. Live is 3965 B (mtime 2026-07-23 19:52); mirror SCHEDULED_TASK_PROMPT.md is 3893 B. Beyond the expected frontmatter/fence differences, the live copy carries an invariant 9 that the mirror never received:
9. UNATTENDED RUN — no GUI, ever. … forbidding browser/preview/dev-server calls to "visually verify" the HTML, because the call sits on an unanswerable permission prompt until it times out (~5 min) and leaves an orphaned dialog that looks like a hung routine. Validate mechanically instead.
That is a correct and valuable fix — it is the canonical unattended-permission-stall failure mode — but it was applied to the live copy only, and appears in neither the mirror nor the Status Log. CLAUDE.md requires dispatcher-level changes to update both. Note the edit landed 07-23 19:52, before the window opens, so the window itself contains no mid-soak prompt change: re-basing removed that confound rather than merely dodging it. The task description also still reads "v3" (a known open operator action from the cutover row).
M1 — scoring.log dropped rank on 2026-07-24. All 66 entries that day omit the mandated rank field; 07-25→07-28 all carry it. render.md specifies rank in the entry schema. Effect is confined to log analysability — the importance ≥ 6 filter and the histogram were correct, and nothing downstream consumes rank today. Root cause is almost certainly emission variance on an unusually long line (66 entries, the window's largest), with no validator to catch it. Fix belongs with the log writer, not a prompt knob.
M2 — one same-entity continuation published as fresh. us-china-ai-openweight-policy-debate (ai-general) appears on 07-25 ("White House Draws New AI Line on China…") and 07-28 ("Amodei says Anthropic doesn't oppose open-weight models…") with identical fingerprint entity and token Jaccard 0.50 — a match on both sufficient conditions of the stated rule (same tag AND (same entity OR J ≥ 0.5)). The 07-28 entry nonetheless committed is_update=False, matched=None and rendered without a marker. This is a false negative: no fabricated prior coverage and nothing wrongly hidden, but the reader loses the continuity signal on a thread already covered 3 days earlier.
Importantly, no permitted knob fixes this. Knob 1 is the needs_judgment lower bound; this pair was already inside the judgment band via same-entity, so widening the band changes nothing. The miss is in how the case was resolved, which is prompt guidance (continuity.md) or an operator call — explicitly not improvised triage per C1_TRIAGE_GUIDE.md. Hence CONTINUE with an S6 follow-up rather than TUNE.
A second instance of the same class sits just outside the window and is worth S6's attention as corroboration: on 07-30, salamanca-sally-robot-paused republished as (fresh; editorial timeline note) with the identical entity string to its 07-26 entry (J = 0.36, inside the judgment band), again committing is_update=False. The model consciously noted the timeline in prose instead of using the Update: rendering. Two instances in six days suggests the boundary case is real rather than a one-off, and that continuity.md's instruction for resolving a same-entity/low-Jaccard pair is the thing to sharpen.
M3 — live/mirror dispatcher drift (bookkeeping). Per Spot-check E: port invariant 9 into SCHEDULED_TASK_PROMPT.md, add a Status Log line for the 07-23 edit, and settle the stale "v3" task description. Not a soak defect; do it before any future shadow-verify or rollback reads the mirror as ground truth.
M4 — hold-queue observability. Per Spot-check D: no per-run state history, so hold/promote/drop cannot be audited retrospectively. This is the assessment's binding constraint, not the pipeline's.
Also carried forward, unchanged and still open from 2026-07-13: carryover_candidates.json has no url field, so carryover-only Also Noted items cannot be cited and are silently omitted. Confirmed still true — pending_stories.json has original_rss_url, carryover does not.
CONTINUE.
The gating question was whether decomposing the monolith and moving continuity arithmetic into Python preserved v3's behaviour without introducing new failure modes. It did more than preserve it — it fixed defects the shadow-verify had attributed to v3. The carryover decay is now exact 9/9 where v3 kept an age-6 entry on non-0.85 arithmetic. Alerts now persist to memory on all 5 days where v3 lost the 07-15 alert entirely. The enrichment log writes cluster_position with no cluster_rank residue, where v3 wrote positions into the wrong field on 07-16. Every hold-eligible cluster in the window either published or promoted at exactly 2 days, where v3 published four hold-eligible clusters outright and permanently lost edtech-screen-time-worth-it. Those are the four concrete v3 defects the refactor was supposed to close, and all four are closed.
Nothing approaches a ROLLBACK trigger. There is no hard-constraint violation: no URGENT alert held or suppressed (7/7 published same-day, the one apparent 07-24 shortfall resolving to a correctly consolidated card), no in-window permanent loss, no state hand-edit, no state corruption. There is no story_state.py logic defect — the 98-test suite is green and every piece of arithmetic I could recompute by hand matched the module's output exactly. TUNE is also wrong, and for a specific reason rather than by default: the two real defects are a log schema omission and a judgment resolution, and neither maps to any of the three knobs the triage guide permits. Reaching for a knob here would be improvisation of exactly the kind C1_TRIAGE_GUIDE.md forbids.
The honest caveats are that this is 5 days, not 7 — the prompt's own floor, and enough to grade every criterion, but the thinnest window that qualifies — and that Spot-check D has an open question I could not close because state files are overwritten. I am calling CONTINUE with that stated rather than hedging to EXTEND: every criterion that is measurable passed, the two defects found are Minor and independently actionable, and a further 7-day soak would re-measure a system whose core arithmetic is now demonstrably deterministic. The residual risk is concentrated in M4 (observability), and the right response to that is instrumentation in S6, not another week of the same blind measurement.
One forward-looking note for the Change 2 gate. M2 is a judgment miss at the fingerprint boundary, and Change 2's corroboration engine adds a third promotion path keyed on the same fingerprint-match function. Sharpening continuity.md's same-entity/low-Jaccard guidance is therefore worth doing before S8b wires that path, not after — otherwise a known boundary weakness gets a second consumer.
[Open on Opus 4.8] The Change 1 assessment returned CONTINUE (assessments/change1-2026-07-30.md).
Execute session S5 from IMPLEMENTATION_PLAN_2026-07.md: update SYNTHESIS_DEPTH_PLAN.md
(decision log + Change 1 status) and the Session Map (S5 done; S6 + S7 unblocked), and
append the Status Log line. Then surface the S6 gate question (plan §5, Q6) to me.
Carry these four Minor findings into S6's scope (all detailed in the verdict file):
M1 scoring.log omitted `rank` on 2026-07-24 (schema drift, 1/5 days; no validator)
M2 one same-entity continuation published as fresh (07-25 -> 07-28,
us-china-ai-openweight-policy-debate, J=0.50); second instance out-of-window on
07-30 (salamanca-sally-robot-paused). NOT knob-tunable — this is continuity.md
guidance for same-entity/low-Jaccard pairs. Sharpen it BEFORE S8b wires the
corroboration promotion path onto the same fingerprint-match function.
M3 live SKILL.md carries an invariant 9 absent from SCHEDULED_TASK_PROMPT.md
(edit of 07-23 19:52, unlogged); task description still says "v3"
M4 no per-run state history -> hold/promote/drop cannot be audited retrospectively;
add a hold-event JSONL or a dated per-run snapshot of the 3 state files
Plus the still-open 2026-07-13 item: carryover_candidates.json has no `url` field.
[Open on Opus 4.8] The Change 1 assessment returned TUNE (assessments/change1-2026-07-30.md). Execute session S5 from IMPLEMENTATION_PLAN_2026-07.md: apply guides/C1_TRIAGE_GUIDE.md for each failed criterion — one knob per cycle, max 2 cycles per criterion; anything the guide doesn't cover comes to me, not an improvised fix. Re-register a 7-day re-assessment (this prompt's shape, dates shifted) before closing.
[Open on Opus 4.8] The Change 1 assessment returned ROLLBACK (assessments/change1-2026-07-30.md). Execute session S5 from IMPLEMENTATION_PLAN_2026-07.md: run guides/C1_TRIAGE_GUIDE.md's 9-step rollback (v3 body from PROMPT_HISTORY.md; state files are shape-compatible both ways). Record the rollback in SYNTHESIS_DEPTH_PLAN's decision log + Status Log (PROMPT_HISTORY stays frozen). If the trigger was a story_state.py logic defect: offline fix + shadow re-verify before any retry.
[Open on Opus 4.8] The Change 1 assessment returned EXTEND (assessments/change1-2026-07-30.md). Not enough data. Register a new one-shot assessment 7 days out reusing CHANGE1_ASSESSMENT_PROMPT.md (shift the dates). If the daily task skipped >=3 days, investigate the task schedule first (the nightly job has silently skipped days before — e.g. 2026-07-14, 2026-07-29).