← All post-mortems
SMAD PickleBot · Post-mortem
Post Mortem: 9/29/26 New members got no welcome DM and no sheet row for 7 days
From 9/22 to 9/29 the daily member sync could not add any new member to the Player sheet, so three new players got no row and no welcome DM. A 9/9 change widened the row the sync writes without widening the range it writes to; Google Sheets refused every insert, and the sync reported success anyway.
Failed
9/22/26 12:20:40 PM
Detected
9/29/26, time not recorded
Fix pushed
9/29/26 11:57:47 AM
Restored
9/29/26 12:13:06 PM
Summary
Once a day, the Sync Group Members workflow compares the SMAD WhatsApp group with the Player sheet. Each new member gets a row on the sheet, in name order, and a welcome DM with the address, parking, court fees, how the weekly poll works and the survey link. Without the row the bot does not know the player: no vote tracking, no charges, no reminders.
On 9/9 at 9:24 PM PT, Join Date, Referrer and Join Source on the Player sheet, backfilled from history and stamped by sync-members added three columns to the sheet and changed how the sync builds a new member's row: the row now ran through every fixed column, including the three new ones. The range the row is written to was left as it was, ending at the win/loss columns. Google Sheets refuses a row that is wider than its range, so from then on every new member's row was refused, and without a row no welcome DM is sent.
Nobody joined for the next 13 days, so nothing showed. The first new member, ND, arrived on 9/22 and was refused at 12:20:40 PM PT. John Stowell (9/27) and cunningham dan (9/28) followed. The sync retried each of them twice a day and failed every time. It logged each failure and then exited as a success, so the workflow watchdog, which alerts on failed runs, saw 14 green runs in a row. Gene found the outage on 9/29 because two new players told him they had heard nothing.
Outage: 7 days, from 9/22 12:20:40 PM PT, the first refused member, to 9/29 12:13:06 PM PT, the last of the three welcome DMs, sent by a sync Snow White ran by hand on the fixed code at Gene's request. The defect itself was live from 9/9 9:24 PM PT, 13 days before any player met it. ND is Andy Tien, "John S" is John Stowell and "cunningham dan" is Dan Cunningham: the sync now takes names from Google Contacts first.
End-user impact: three new members got no welcome and no row. ND waited 7 days, John Stowell 2, cunningham dan 1. They had no location, parking, fee or poll instructions from the bot, and the bot did not know them: poll votes from them were not tracked and they got no vote reminders. Financial impact: none. From the Pickle Poll Log, only Andy Tien voted in the window: "Can't play this week" on the 9/27 poll, on 9/28 at 9:43 AM, before he had a row. His row now carries it. John Stowell and Dan Cunningham did not vote, and none of the three played, so nothing went uncharged. Side effect: the refused inserts left rows with no name in the Player sheet. The logs show 20 refused attempts (ND 14, John Stowell 4, cunningham dan 2), each of which inserted a row before its write was refused. Snow White found 18 such rows, under Chris Lee (2), Jesse David (4) and Nardo Manaloto (12). Each had copied the formulas of the row above it, that player's Mobile included, so the sync saw Nardo's number on 13 rows. Why 18 and not 20 is not established. All 18 were deleted on 9/29 between 12:14 and 12:37 PM PT (time not recorded) with Gene's approval.
Who did what
- Gene: asked for the Join Date and Referrer columns (9/9 4:33 PM PT). Found the outage when two new players had no welcome DM (9/29) and asked for this post-mortem.
- Mr Sandman (cloud session): handed the feature to Snow White with a note that "every positional writer (sync-members, …) must be checked" (9/9). On 9/29 read the run logs, found the cause, wrote the fix and this post-mortem.
- Snow White (desktop session): built the feature in 9428645, added the three columns to the live sheet with
scripts/backfill-join-dates.py, and widened the sync's row. The write range was not widened, and no test ran the insert. On 9/29 ran the fixed sync by hand at Gene's request (restored), confirmed the welcome DMs in GREEN-API's outgoing log, found the leftover rows and why they carried Nardo's number, deleted them with Gene's approval, checked the votes, and added the safeguards in items 9 to 12. Its 9/29 note first gave the deletion time as "about 12:55 PM", a guess made before 12:55; the time was not recorded.
Timeline (PT)
| When | What happened |
| 9/8 12:05:14 PM | Scheduled sync inserts Jerry Farrell and sends his welcome DM (last scheduled success) |
| 9/9 4:33 PM | Gene asks for Join Date and Referrer columns; Mr Sandman hands it to Snow White |
| 9/9 9:01:00 PM | A manual sync inserts Sean Liu and sends his welcome DM: the last successful insert, on the old code |
| 9/9 9:24:50 PM | 9428645 ships: row widened, range not. The backfill adds the three columns to the live sheet the same evening (time not recorded) |
| 9/10 to 9/21 | Every run green; no new member joins, so nothing is inserted |
| 9/22 12:20:40 PM | First refused member: ND. Requested writing within range ['2026 Pickleball'!A58:V58], but tried writing to column [W]. The run reports success |
| 9/22 to 9/26 | ND refused in every run, twice a day, each attempt leaving a nameless row; every run green |
| 9/27 12:23:21 PM | John Stowell joins and is refused too |
| 9/28 2:27:22 PM | cunningham dan joins and is refused too; 3 refused per run from here |
| 9/29, time not recorded | Gene: "Why don't you do daily player sync anymore? I got 2 new players who haven't received welcome dm yet". Detected |
| 9/29 | Mr Sandman reads the 9/28 run log, finds the refused rows and the cause, and fixes it (commits below) |
| 9/29 11:57:47 AM | The fix (a9873cf) is pushed with the incident record (0a68198); its deploys and tests run green. The next scheduled sync would have been the next morning |
| 9/29 12:13:04 to 12:13:06 PM | Restored: Snow White runs sync-members --execute on the fixed code at Gene's request. Andy Tien, John Stowell and Dan Cunningham get their rows and welcome DMs (delivered or read, from GREEN-API's outgoing log) |
| 9/29 12:14 to 12:37 PM, time not recorded | The 18 nameless rows are deleted with Gene's approval, after delete-blank-player-rows was changed to "no first or last name" (66f2c31) |
Why it happened
- The row and its range were computed separately. The insert built its row from one number (every fixed column) and its range from another (the last column it filled). Every other write to the Player sheet goes cell by cell through the column map, so this was the only place where the two could disagree. 9428645 changed one of the numbers and not the other.
- A refused insert did not fail the run.
sync-members logged the error and exited 0. The watchdog alerts on failed runs, so 14 runs that failed at their one job looked healthy. - The insert added the row before writing it, with nothing to undo it. Each refused attempt left a nameless row behind: 20 attempts in the logs, 18 rows found on the sheet.
Why it shipped
- The insert had no test. The commit's test,
scripts/test-sheet-columns.py, checked that the new columns are read as fixed columns and not as game dates. Nothing ran the insert, and no fake Sheets service enforced the rule that a row must fit its range. - The warning was followed by reading, not by running. The hand-off said every positional writer had to be checked. The row width was checked by eye; the range, three lines further down, was not.
- Nothing exercised it for 13 days. The defect needed a new member to show. By then the change was two weeks old and nobody connected it to the refusals.
Action items
| # | Action | Status |
| 1 | A player row and its range come from one column map: ColumnMapper.static_row() and static_range(), used by the insert | Done a9873cf |
| 2 | A member who could not be added fails the run, so the watchdog files an alert | Done a9873cf |
| 3 | tests/test-sync-members.py, the first suite for the insert. Its fake Sheets service refuses a row wider than its range, as Sheets does. It fails on 9428645's code | Done a9873cf |
| 4 | A refused write deletes the row it inserted, so a retry cannot leave a nameless row | Done 0a68198 |
| 5 | The survey sync and jokes steps of sync-members.yml run even when the member sync fails, now that item 2 can fail it | Done 4324176 (Snow White) |
| 6 | Remove the nameless rows the refused inserts left in the Player sheet | Done 18 rows deleted on 9/29 between 12:14 and 12:37 PM PT with Gene's approval, by the rule "no first or last name" (66f2c31, Snow White). The first command, 08cf29b, looked for rows with no value and no formula and found none |
| 7 | Check whether ND, John Stowell or cunningham dan voted or played between 9/22 and 9/29, and charge any game they played | Done (Snow White): only Andy Tien voted, "Can't play"; his row now carries it. Nobody played; nothing owed |
| 8 | Get the three their rows and welcome DMs, and record the restore time | Done Snow White's run at 12:13 PM PT on 9/29; restored 12:13:06 PT |
| 9 | New members are named from Google Contacts first, so a WhatsApp handle like "cunningham dan" is welcomed by name | Done 8f7dfae (Snow White) |
| 10 | A phone number on two player rows fails the run (the copied Mobile formulas put Nardo's number on 13 rows) | Done 66f2c31 (Snow White) |
| 11 | After every real run the sync checks one active row per group member, no shared numbers and no nameless rows, and fails if any check fails | Done 8b92ced (Snow White); the first live check passed, 48 members on 48 rows |
| 12 | A member who leaves the group is told their stats are archived and how to come back (DM + email, Gene's wording); the sync used to archive them silently | Done 66f2c31 (Snow White). Jerry Farrell, archived by the 9/29 sync before this existed, was sent the notice by DM at 12:36 PM PT at Gene's request (no email on file) |
← All post-mortems
SMAD PickleBot · Post-mortem
Post Mortem: 9/27/26 WhatsApp token rotation took WhatsApp down for 4 minutes
On 9/27 a planned rotation of the GREEN-API token took WhatsApp down for 4 minutes. GREEN-API switches to a regenerated token a few minutes after showing it; the rotation script checked too early, stopped, and the old token then expired. Nothing was sent in those minutes, so no player was affected.
Failed
9/27/26 10:40:04 AM
Detected
9/27/26, before 10:40:04 AM (not recorded)
Fix pushed
9/27/26 10:42:36 AM
Restored
9/27/26 10:43:41 AM
Summary
The GREEN-API token is the one key the bot uses for every WhatsApp call. It was being rotated on purpose. The day before, the session had put the token's first half into a test file, which GitGuardian flagged, and had printed the whole token into its own transcript. The rotation script, scripts/rotate-greenapi-token.ps1 (9/26), checks a new token with GREEN-API before writing it anywhere. If GREEN-API refuses it, the script stops and changes nothing.
On 9/27 that safety check caused the outage. GREEN-API switches to a token regenerated in its console a few minutes after the console shows it. The script checked the new token straight away, got a 401, treated it as a bad copy and stopped. The old token was still live at that point, so nothing had broken yet. The session then misread the 401 as "some other GREEN-API key" and shipped a second way to regenerate the token, which also got a 401. A few minutes later GREEN-API retired the old token, and every call from the three functions started failing. Service came back on the fourth run, using a paste-at-the-prompt mode written during the outage.
Outage: 4 minutes, from the first logged failure at 10:40:04 AM PT to the first confirmed send at 10:43:41 AM PT (the incident record's start and end). The session saw 401s shortly before 10:40:04 but did not record the time, so the record starts at the first logged one.
End-user impact: none observed. Nothing tried to send in the window: no poll, reminder or /pb command. Had one arrived, its WhatsApp send would have failed. Financial impact: none. No payment, booking or charge touches GREEN-API.
Who did what
- Gene: regenerated the token in the GREEN-API console and ran the script four times.
- Snow White (desktop session): wrote the script. After run 2 it misdiagnosed the 401 and shipped the
updateApiToken path. It found the outage by checking the old token after run 3, gave a clipboard one-liner that could not work (copying the command replaced the token on the clipboard), then wrote the -Paste mode that restored service, and this write-up.
Timeline (9/27/2026, PT)
| Time | What happened |
| 10:30:04 | Webhook instance-state poll: authorized on the old token |
| not recorded | Run 1. The clipboard held the token already in use (the console had not regenerated). GREEN-API answered 429 (rate limit). The script stopped with nothing written. Correct outcome, confusing message |
| not recorded | Run 2. The console regenerated; the new token starts 259b4d. The script's check got 401 and it stopped with nothing written. The session checked the old token: still authorized |
| not recorded | The session concluded the pasted value was "some other GREEN-API key". Wrong: it was the new token, not yet active |
| 10:37:42 | Pushed rotate-greenapi-token.ps1 regenerates the token itself through GREEN-API's updateApiToken |
| not recorded | Run 3. updateApiToken, called with the old token: 401 |
| not recorded, before 10:40:04 | The session checked the old token: 401 on getStateInstance, getSettings and getWaSettings. Outage identified |
| not recorded | The session gave a one-line .env update that read the clipboard. Copying that command from the session replaced the token on the clipboard, so .env was not updated |
| 10:40:04 | Webhook poll: getStateInstance returned 401 (WARNING, not paged) |
| 10:41:19 – 10:41:41 | Run 4 with -Paste: the new token was accepted (authorized), both secrets were updated and the new revisions went live (smad-picklebot-00523-nkt, whatsapp-message-sender-00397-wbh, smad-whatsapp-webhook-00488-nm8) |
| 10:42:36 | Pushed rotate-greenapi-token.ps1 -Paste: start it, then copy the console's new token |
| 10:43:41 | Restored: the first send on the new token (the reply to Gene's /pb help) was delivered |
Why it happened
- The script assumed GREEN-API switches tokens instantly. On 9/26 a rotation went through on the first check, and the delay was never looked for. A refusal was treated as final, so the script stopped at exactly the moment the switchover was already under way.
- The session diagnosed from one sample. After run 2, the old token still worked, and that was read as proof nothing had been regenerated. "Not active yet" was never considered, and a second, untested mechanism was shipped instead of waiting a few minutes and checking again.
- An instruction that fought the clipboard. The fix-up command read the token from the clipboard, but running it meant copying the command, which replaced the token.
Why it shipped
The script can only be tested by a real rotation, which is itself the risky act. The first real rotation (9/26) happened not to show the delay. Its parse was checked; its behaviour against a slow switchover never was.
Action items
| # | Action | Status |
| 1 | Paste at a prompt: start the script, then copy the token, so nothing overwrites the clipboard | Done df583ba |
| 2 | The check waits out a 401 on a new token, and a 429, every 15 seconds for up to 3 minutes, with the token already saved in .env | Done in the commit that adds this post-mortem |
| 3 | The console-paste flow is the default; updateApiToken moves behind -Api because its one try failed | Done in the commit that adds this post-mortem |
| 4 | Never grep across .env; never build a test fixture from a real secret (why this rotation happened at all) | Done session memory rule, and the redact test's fake token is visibly fake |
← All post-mortems
Post Mortem: 9/20/26 Sunday 8am New Games Poll Broke because of refactor regression
On Sunday 2026-09-20 the 8 AM automation that creates the week's games poll and sends every reminder did not start. The poll reached the group at 5:54 PM, ten hours late. A cost-cutting change to the GitHub Actions workflows the night before caused it.
Summary
PickleBot's daily 8 AM job is a GitHub Actions workflow, Daily Reminder Runner. One of its jobs calls a second workflow, Release Notes, to send the daily digest. GitHub's rule for that arrangement: the calling job must grant every permission the called workflow asks for, or GitHub refuses to start the run at all. The night before, a refactor to cut Actions minutes, Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable (2026-09-19, 10:03 PM PT, Snow White), gave Release Notes actions: write and did not give it to the only job that calls it, daily-reminder-runner.yml's release-digest. At 8:00:04 AM GitHub refused the run (startup_failure, zero jobs). Nothing that rides that job happened: no games poll, no Games This Week report, no Last Call scheduling, no game day, vote, payment or survey reminders, no Venmo sync, no court cache refresh, no digest. The watchdog filed the alert 88 minutes in; the fix waited until 5:43 PM for a session that was awake and allowed to make it; a hand dispatch restored service at 5:55 PM PT. Outage: 9 hours 55 minutes.
The same file had failed the same way on 2026-09-04, and the comment documenting that failure sat six lines above the block that was not updated. The same push carried a second defect: it gated the alert closer on a repository variable that the Actions token cannot write (HTTP 403); the refusal was logged as a WARNING under a step that reported success, so alert auto-closing was silently off for a day.
End-user impact. No player could vote on the new week's games for ten hours: the Sunday poll that normally opens at 8 AM did not exist until 5:54 PM. Every reminder due that morning was also missed. One knock-on found the next day: Shyam's vote-to-court handover did not fire for Tuesday 7 PM. He voted at 7:54 PM Sunday, 47.1 hours before the game, and the handover requires 48 (a member can only book his own free court 48+ hours ahead: Shyam court: skipping 09/22/26 - ... 47.1 hours away in the vote webhook's log). With the poll out at 8 AM he would have voted 11 hours earlier and it would have run. The automation itself is intact; the late poll pushed his vote inside the window.
Financial impact. No direct money loss. Potential loss from players who normally vote on Sunday for the week's games, could not, and may not get around to it later in the week, which means fewer paid slots for those games.
Who did what
Three Claude Code sessions and two pieces of automation acted in this incident. Each is named with what it did right and what it did wrong.
| actor | what it is | part in this incident |
| Snow White | Claude Code session on Gene's desktop; the only session with credentials, run logs, a dispatch token and, by project rule, the right to edit workflows | Caused the regression. Wrote and pushed the change that gave Release Notes a permission its caller did not grant, and the variable gate the workflow token could never open; verified only the top-level path; wrote tests that encoded the same mistake. Was off Sunday morning, so could not repair it for ten hours. Then: dispatched the runner that restored service, removed the variable gate and rebuilt the closer, corrected the incident record, updated the investigator's instructions, wrote this post-mortem. |
| Alert Investigator | Hourly cloud routine: reads open alert issues, triages from the runs API and the code, writes its finding on the issue and in the sessions notes, wakes Mr Sandman; cannot read run logs, dispatch or change code | Triaged correctly and wrote one ambiguous line. Named the cause, the one-line fix and the need for a hand dispatch 46 minutes after detection, and separately spotted that the alert closer's gate had not opened. But reported the Gmail watch as "renewal (day 20, even)": accurate shorthand for "the runner renews on even calendar days and today is the 20th", and open to any reading by someone without the workflow file in front of them. Its live instructions now forbid that shape. |
| Mr Sandman | Claude Code session in a 24/7 cloud sandbox; has the repo and git push; no credentials, no run logs, no dispatch token | Fixed the outage, and added noise to the investigation. Once Gene handed it the workflow edit, it pushed the caller fix and the lint check that catches the whole class within the hour. But it read the investigator's line as "day 20 of a 7-day watch", repeated that in two notes and in the incident record, and told Gene the Gmail watch had a hard deadline that day. It did not verify: the runner's own workflow file states the even-day rule, and the runs API showed the 9/18 renewal step green (the step's log text is out of the sandbox's reach; the file and the step conclusion were not). Its incident entry also said "no monitor caught it", when the watchdog had filed the issue 88 minutes in. Both are corrected. |
| Workflow Watchdog | GitHub Actions workflow, hourly sweep over every workflow's latest run | Detected it. Read the failed conclusion and filed #69 at 9:28 AM. The 8:41 sweep ran 47 minutes late: GitHub's scheduling, not the watchdog's. |
| Cloud Scheduler | GCP job that dispatches the runner at 8:00 AM | Dispatched on time. Its retries cover a failed dispatch call, not a run GitHub accepted and then refused to start. |
| Gene | Owner; the only human in the loop, and the only one who can turn the desktop session on | Away all day, and the decision-maker once back. Spent Sunday driving his son to UCSD and packing, and did not notice the 8 AM poll had not gone out. Around 5 PM, on the drive back from San Diego, opened the Mr Sandman session, saw the outage, and told Mr Sandman what to do. Snow White's desktop had been off since Saturday night. He did not want Mr Sandman making production fixes: it has no live logs and no keys, and it had not caused the regression, so he allowed it only the one-line workflow edit that every weekday run depended on. Home around 6 PM, he turned the desktop on and told Mr Sandman to hand everything else to Snow White. Chose the full dispatch over poll-only, and asked for this post-mortem. |
Timeline (PT)
| when | what |
| 09/19 22:03 | Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable pushed. Its own runs (four deploys, Tests, Lint) green: they exercise release-notes.yml only as a top-level workflow, where its own block sets the token. The called form runs once a day. |
| 09/20 08:00:04 | Cloud Scheduler 8am-daily-runner dispatches; run 35518182853 → startup_failure at 08:00:06. Scheduler retries (2) do not apply: the dispatch itself returned 200. |
| 09:28 | Workflow Watchdog hourly sweep (cron :41; GitHub ran it 47 minutes late) reads the conclusion and files #69. Detection: 88 minutes. |
| 10:14 | Alert Investigator triages on #69 and in the sessions repo: cause, the one-line fix, "dispatch by hand". Correctly notes the Gmail watch renewal was due that day (even calendar day). |
| 10:14 → 17:00 | Nobody who could act was available or allowed to. Snow White's desktop had been off since Saturday night. Mr Sandman cannot dispatch and, by CLAUDE.md, does not write workflows. Gene was on the road to UCSD with his son and did not see the alert. Nearly 7 hours of the outage is this gap. |
| ~17:00 | Gene, driving back from San Diego, opens the Mr Sandman session, sees the outage and directs it. He does not want Mr Sandman making production fixes (no live logs, no keys, not the author of the regression) and allows it only the one-line workflow edit every weekday run depends on. |
| 17:43 | The 8am runner starts again: a caller must grant what the workflow it calls asks for pushed by Mr Sandman: caller fixed; check-workflow-reporting.py check 3; tests/test-workflow-permissions.py; incident entry with restored: null. |
| ~17:50 | Gene, home, turns the desktop on and tells Mr Sandman to hand everything else to Snow White. |
| 17:53 | Snow White dispatches run 35549088614 (reminder_type=all, Gene's choice). |
| 17:54:13 | [OK] Availability poll created in SMAD Pickleball group. |
| 17:55:10 | Run complete, every step green, release digest ran. Restored. |
| 18:01 | Alerts close from the watchdog's hourly sweep; the OPEN_ALERTS gate is gone pushed: defect 2 removed, closer rebuilt, scopes reverted, both incidents recorded. |
| 18:02 | A hand dispatch of the watchdog on the new closer closes #69 with the sweep's own comment. |
What the refactor was
Actions minutes: retire GREEN-API Watch, gate close-on-green on an OPEN_ALERTS variable, one push of 23 files after GitHub's "90% of Actions minutes" alert, did three things:
- Deleted
greenapi-watch.yml, a workflow README already recorded as retired (126 billed minutes that month for a check the keepalive poll makes). Correct, and not involved in the outage. - Gated
close-on-green.yml on a repository variable OPEN_ALERTS, to stop a six-second job from billing a minute 193 times a month. The reporter was to set the variable, the closer to clear it, the watchdog to re-derive it hourly. Defect: GITHUB_TOKEN cannot write repository variables at all; actions: write does not cover it. The gate never opened; the refusal was a printed warning; the step reported success. - To let workflows write that variable, added
actions: write to the permissions block of every workflow containing the string report-failure: fifteen files, including release-notes.yml. Defect: daily-reminder-runner.yml calls release-notes.yml as a reusable workflow (uses: ./.github/workflows/release-notes.yml), does not contain the string, and was not patched. Its release-digest job granted contents: read, issues: write; the callee now asked for actions: write; GitHub refused to start the run.
Defect 3 caused the outage. Defect 2 is why #69 then sat open behind green CI.
Why it shipped
Three failures of method, in order of weight.
- The verification proved the wrong path. I read a green push as proof of the change. The push's runs exercised release notes only as a top-level workflow. The called form runs at 8 AM, once a day, and nothing I did before pushing exercised it.
daily-reminder-runner.yml:333-345, written after the 09/04 outage, says exactly this: "Direct dispatches of release-notes.yml worked all along... The failure only exists in the called form... so the verification runs the night before proved the wrong path." I was editing permissions across the repo and did not read the one file where permissions have a second consumer. - The blast-radius search matched a string, not a relationship. "Who uses
report-failure?" is answerable by grep and I answered it. "Who inherits this job's permissions?" is a graph question with a second edge for every callee, and I never asked it. The static check that now exists (check 3) is the search I should have done by hand. - The tests I added encoded my own predicate. "Every file containing
report-failure grants actions: write" passed on the callee and could not see the caller. docs/LESSONS.md already has a section titled The Author's Tests Inherit the Author's Blind Spot, written on 2026-08-30 for three instances of this shape. I repeated it.
Contributing: the change was framed as a low-risk cost cut and shipped late on a Friday in a single 23-file push; and I declared the variable gate "proven on the first alert", which is to say unproven, with its failure mode set to report success. Both halves of the rule a permissions block is a denylist by omission; a called workflow is bounded by its caller's job were in CLAUDE.md; I applied the first.
The Gmail watch question
Gene asked why the investigator said the Gmail watch renewal "hasn't been done in 20 days". It had not said that, but what it did say was no better for a reader: "Gmail watch renewal (day 20, even)". What that shorthand meant: the runner renews the watch on even-numbered calendar days, 9/20 is one, so the skipped run also skipped a routine renewal. Mr Sandman rendered it as "day 20 of a 7-day watch" in its 5:55 PM note, again in the 6:25 PM relay to Snow White, and in the incident entry's impact line, without checking the runner's workflow file or the 9/18 run, and the number changed meaning on the way from "the 20th" to "20 days old".
The facts from the run logs: the watch was renewed on 9/18 at 8:01 AM PT (Expires: 2026-09-25 15:01:28). The 9/20 attempt was skipped with the rest of the run; the next even day, 9/22, renews it three days before expiry. Nothing was at risk. The incident entry is corrected.
Two lessons, both now written down. The investigator's live instructions (the routine's prompt, and Appendix A of Investigator.md) now require every figure to carry its meaning in the same sentence, and every commit, run and issue to be named by title and link, never a bare hash: "day 20, even" was accurate and useless to anyone without the workflow file open. And a number that crosses a relay must carry its source (docs/LESSONS.md, A Relayed Number Carries Its Source).
A side finding: gmail-watch-renewal.yml ("every 6 days") has no schedule, last ran in March, and failed then. The runner's even-day step is the only live renewal. See action item 9.
Detection and repair
Detection worked: the watchdog's hourly sweep filed #69 88 minutes after the failure, and the investigator had the cause and the fix written within 46 minutes of that. The incident entry as first written said "no monitor caught it"; that was wrong and is corrected.
Repair did not: nearly 8 of the 10 hours were spent waiting for a session that was both awake and allowed to edit a workflow and dispatch a run, and for the one human who could turn that session on, who was on the road all day and saw the alert at 5 PM. The alert reached the right places; it reached nobody who could act. That wait is a policy, not an accident.
Earlier instance: 2026-09-04
The same file failed the same way sixteen days earlier and was never recorded until now. Send the whole day's list, to three chats, at 8am from Cloud Scheduler (09/03) added the call to release-notes.yml with no permissions block on the calling job; the 09/04 8 AM run failed to start (run 33886955002); Let the 8am runner start again: grant the digest job the scopes it calls for fixed it at 9:56 AM and a hand dispatch (run 33898063769) restored service at 9:58 AM. Outage 1 h 58 min. The fix's comment described the trap precisely; a comment is not a check.
Action items
| # | item | status | where |
| 1 | Caller must grant what callee asks: scripts/check-workflow-reporting.py check 3, run by Lint Workflows on every push and by tests/test-workflow-permissions.py in the local suite, so the local suite fails before a push | done | The 8am runner starts again: a caller must grant what the workflow it calls asks for (Mr Sandman) |
| 2 | Remove the variable gate: the closer runs from the watchdog's hourly sweep with no state to write; actions: write removed from all sixteen callers; OPEN_ALERTS deleted; proven on #69 under the workflow token | done | Alerts close from the watchdog's hourly sweep; the OPEN_ALERTS gate is gone (Snow White) |
| 3 | Incident record corrected: 09/20 restored time, Gmail wording, detection credited to the watchdog; 09/04 outage entered with the commits that introduced and repaired it | done | Post-mortem: the Sunday poll never went out (2026-09-20) and the revision that added this section's links |
| 4 | The investigator writes for a reader who has not seen the code: every figure carries its meaning; commits, runs and issues by title and link | done | Live routine trig_019cYNV8RgZ5k9gvmPhvnPCR updated 2026-09-20 6:40 PM PT; Appendix A of Investigator.md in the same revision |
| 5 | Before any permissions: change: grep uses: ./.github/workflows/, read every caller's job, and name which trigger paths the push's own runs will not exercise before calling it proven | done | Snow White's standing rule; docs/LESSONS.md |
| 6 | One push per task, and a permissions change is its own task: never bundled with a deletion and a new gate in one late-evening push | done | One push per task (Gene, 2026-09-19 PT): the project reading of the shared rule |
| 7 | Give the cloud session a repair path when the desktop is off: a committed ops/rerun.txt that a push-triggered workflow reads and dispatches the named workflow with the given inputs. Turns an 8-hour wait into a push. One billed minute per use | done | The rerun lever: a push to ops/rerun.txt dispatches the workflow it names, so the cloud session can repair production when the desktop is off; gmail-watch-renewal.yml is retired (Snow White, approved by Gene 2026-09-22 PT) — rerun.yml, .github/scripts/rerun_lever.py, tests/test-rerun-lever.py |
| 8 | Or remove the edge entirely: have the runner's release-digest job dispatch release-notes.yml through the API instead of calling it as a reusable workflow, so no caller→callee permission inheritance is left to get wrong | declined | Gene, 2026-09-22 PT: check 3 in scripts/check-workflow-reporting.py fails Lint Workflows on the whole class before it can reach main, and tests/test-workflow-permissions.py pins it locally; an API dispatch needs the same actions: write scope, bills the same one job, and turns one run into two to watch |
| 9 | Retire gmail-watch-renewal.yml (no schedule, dead since March, failing) and its README rows; the runner's even-day step is the renewal | done | same commit as item 7: The rerun lever … gmail-watch-renewal.yml is retired — verified from the runs API 2026-09-22 PT: workflow_dispatch only, last run 2026-03-04, failed |
| 10 | A relayed number carries its source: session notes and incident entries quote the log line, not a paraphrase, for any figure with a deadline in it | done | docs/LESSONS.md, A Relayed Number Carries Its Source |