Greenpen · Fourthpen Flows · commissioned by Khoa via Will · 10 Aug 2026
Flows
A flow is a saved route through the room you already have. It adds three
things in front of it — an objective you pick, a check that runs before anything moves,
and a script the Showrunner follows through the existing paces.
Which means the only visible change is the front door. The seats, the rail, the
gate tray and the continuity lane are untouched: this round extends
/lane and /readers, it does not replace them.
Built to the spec on
PR #58
(
docs/FLOWS_SPEC.md, branch
will/flows-spec),
read from the branch at build time — 190 lines, sha256
4a559ab…. Corpus figures are measured and match the corpus
console exactly.
Round 2 — the board.
Khoa reviewed round 1 and could not find
which flow was running or
where the
flows were: the picker was too faint next to the execution surfaces, and the rail showed steps
but not identity. Two changes: the front door now opens onto a
board you pick by shape,
and every running session
names its flow. Everything Khoa did not flag is unchanged.
See the board running →
Round 3 — the Residents model. Khoa-commissioned same-day full redesign, built to
PR #64
(
docs/RESIDENTS_SPEC.md, branch
will/residents-living-spec).
Residents replace the Floor/in-app two-cast model entirely, initiation moves to Keystack, the
old agent drawer becomes the
Tools panel, and the room now runs exactly
five permanent
shapes instead of a growing flow list. This is a same-day supersession of round 1’s
front door and a re-framing of round 2’s board — both stay live below as the record
of what shipped and why, per the pack’s own extend-not-replace rule.
Start at Arrival →
Amended within the hour — and the amendment itself had already been superseded.
§8 of RESIDENTS_SPEC changed its mind
three times the same day: round 3 was built to
“gates stay in 4P” (the spec as dispatched); an amendment then told me that was
superseded by “approvals happen in Keystack as structured taps” (commit
ff4b799); but by the time I read that amendment, Khoa had already ruled a
third time —
4f9d62a, marked final:
conversation is the
control plane, an instruction to a resident
is the authorization, and buttons are an
optional convenience, never a gate. I built to
4f9d62a, not to the
amendment’s own citation, and verified that against the branch rather than trusting the
payload’s summary.
Corrected on already-shipped surfaces: Arrival’s
“needs you — GATE, nothing runs until you tap” framing (exactly the defect
Khoa’s own design test now names) is rewritten as a rendering‑pull invitation; the
Tools panel’s pending lane now distinguishes a still‑open item from one
“decided in Keystack, shown here as status”; and the proposal card’s
footer no longer claims 4P gates anything.
New this pass: the option‑button pattern
in Keystack chat — tap and prose resolving to the identical decided record, and a
too‑big‑for‑chat decision handing off with a deep link.
See the corrected surfaces →
Amendment 2 — the ruling generalized, and the lane inverted. Khoa widened
4f9d62a from “approvals move to Keystack” to
there are no required‑approval moments in creative flows at all. Three of the four
stated impacts were already satisfied by the previous pass;
one was not, and it was the
structural one. The Tools panel’s lane was still a forward‑looking queue with an
orange count and an Open button — furniture that, under this ruling, does not just sit
idle but
misdescribes the model to the person reading it. It is now a backward‑looking
Recent decisions record: what the resident did, and the instruction that authorized it,
inline on every row, because §8 makes the instructing exchange the audit trail. The mark
distinguishes prose from tap without ranking them — a lane that rendered taps as more
official than sentences would quietly contradict the ruling it exists to express.
See the lane →
Round 8 — four refinements from Khoa’s live review.
(1) The catalogue splits by
grain:
Flows run a team over time,
Tools are single
capabilities that return and are done. Only one of them has a cast, which is the tell.
(2) Unfilled seats stay visible and now
name their fix — tap one and it says exactly
one thing, chosen by arithmetic: run the next type down if your roster can fill it, otherwise add
that resident in Keystack. Adding it there fills the seat
everywhere at once, because
casting resolves from your roster at render time. And
Flows are now the parent of Types
rather than a sibling row — types are smaller chips nested under the selected flow.
(3) The concept page is stamped
Spec demo and every shaped tab in the product carries a
visible board glyph, so the in‑tab path is walkable without knowing it is there.
(4) A
drag prototype, concept page only: pick a card up, drop it, and it reorders —
a view list, never the artifact, with an undo and a chip saying write‑back is v2. The
product board still has no drag.
Walk the product: Draft → Board →
or open the spec demo →
Round 7 — the great simplification. Khoa’s principle:
the least surface that
does the job. 4P v1 is four things — the
text, the
board (an alt view of the
text), the
record (what residents did) and the
catalogue (what you can ask for).
Keystack is where you command. Everything else was chrome and it is gone: the composer, the gate
tray, the standalone continuity badge, the “the draft · open” row, the
“addressing” line, the artifact picker, the zoom‑tier selector, two of the three
layout buttons, and markup as a named surface.
History stayed, because versions, compare and
archive/restore have nowhere else to live.
The board became an
alt view inside the tab you are already in — where you are
IS the picker — so Draft + Board gives scene cards and Beats + Board gives
beat cards from one press. A tab whose artifact is not shaped has
no board button at all,
which is round 6’s thesis stated as a product fact rather than an essay. Zoom alone drives the
semantic tier. The record lost its composer and became read‑only collapsed runs that expand to
the whole story,
keeping compare cards because rendering‑pulls are the one thing you
cannot do in a chat window. Every face is now an actual resident — Noa and Wren — and
everything else that posts is
marked machinery rather than given a persona.
Open 4P →
Round 7 · the product
4P — text, board, record, catalogue
The whole surface after the cut. Flip Text/Board inside Draft, Outline or Beats;
the drawer holds the record of what residents did and the catalogue of what you
can ask for, which hands you the sentence to say in Keystack rather than a button to
press — because an instruction to a resident is the authorization.
SPEC DEMO · not the product path
Universal Board View — spec demo
In the product, Board opens from inside each tab, not here. This page exists to show
one component rendering three artifact classes at once, and round 8 adds the
drag prototype so the hand‑feel can be judged before build sequencing.
Kept at full chrome deliberately: a class switcher is the wrong control in a product where you
already stand in a tab, and the right one on a page whose subject is the comparison.
Both draw from board-view.js, so the demo cannot drift from the
product it demonstrates.
Round 6 — flow TYPES, and board view as a property of the content model.
Built to
RESIDENTS_SPEC §7 (commit
efa045c,
the branch head — the spec has not moved since).
Types: a shape no longer owns one
roster. It declares a few named types, each a roster and an order, and the drawer’s type row
reorders and extends the faces to match. The order is still
derived from the registry —
round‑5’s
castOrder() grew a second argument rather than a fork,
and
steps is now computed from the default type, so every older view reads
it unchanged. Khoa settled the one judgement call I flagged last round:
the large focal face is
gone, because the ordered faces already are the flow identity. And the
Writers’ Room
epic is simply the largest declared type — all seats, all eight segments, overnight —
drawn in the same strip with no special casing.
See the type row →
Round 6 delivered the type picker and the universal board. Both stand — derived
face‑orders, dashed placeholders and the pace footprint are unchanged, and the board’s
three content models are the same ones the product now opens in‑tab. Round 7 moved the type
picker into the Catalogue, where browsing belongs, and cut the board’s chrome; nothing about
the derivation changed. Both surfaces are listed once, above.
Round 5 — direction reset on flows, from Khoa. He reviewed the round‑4 flow surfaces
and said plainly he does not get them. What he expects is
this page’s lane, with one tweak:
the drawer faces are the actual residents, in the order the active flow runs them.
The flow IS the face-order. So the lane stays the canonical running view; the drawer renders the
cast as ordered faces with the live worker ringed; and
browsing a flow means flipping face-orders in
that same drawer — not a directory page, not a silhouette. Per his instruction
/browse and
the board come
off the user
path and remain as internal/spec template visualisations only; both stay reachable here because the
pack does not delete its record. Every honesty rule from rounds 1–4 applies unchanged to the
drawer idiom — cast marks, four verdicts, no required‑approval moments.
See the lane →
Round 5 shipped these two surfaces — the lane with the flow as face-order, and the Scene Board
as a zoomable corkboard with semantic zoom, derived positions and no drag. Round 6 extended both
in place rather than adding pages, so they are listed once above: the lane gained the type row,
and the Scene Board became the Universal Board View. Everything round 5 established still holds
— the one thing that changed by ruling is the focal face, which is gone.
Round 4 — hold lifted; Evidence Packs, the flow directory, and the headings ruling.
Built to
RESIDENTS_SPEC §9 (commit
b113445) plus
Will’s round‑4 brief. I re‑checked §8 first: it is
byte‑identical to the
MCP ruling rounds 3a/3b were built to, so nothing from those rounds needed redoing — the earlier
“§8 changed” signal was my own range‑matching artifact, since §9 was renamed
from
Open items to
Evidence Packs.
Start at the Evidence Pack →
Round 3 · the new hero problem
Arrival
You did not open this — a Keystack link put you here, mid‑run. The proposal
card doubles as the landing header; a “needs you” zone answers the one question
that matters in one glance.
Round 3 · best craft this round
The proposal card
Shape, cast, depth, the four‑verdict readiness strip and the brief — the
contract for the run. Three real states, computed against an actual 2‑resident roster:
castable, castable‑but‑queued, and honestly blocked two different ways.
Round 3 · replaces the agent drawer
The Tools panel
Same 392px dock, three states: the catalogue (tools + five shapes, presets nested as
prefills), idle (the maturity ladder), running (the dock‑width board).
Round 2 — what the board changed
Round 2 · call 1
The board is the template, drawn — not a picture of it
The brief said a flow already is nodes and edges, so the board should render the
data model honestly. I took that literally: board.js reads the
stored flow document in the spec’s §3 shape and lays it out. Steps become boxes
labelled with their seat, order becomes arrows, the gate
field becomes a diamond, the loop edge becomes a back-arrow. Nothing is hand-positioned,
so nothing can fall out of step with the template.
Which also settles the read-only question structurally: there is no authoring surface
because there is nothing to author into — the picture is downstream of the data.
Round 2 · call 2
Pick by shape only works if the shapes actually differ
“Choose by shape, not by reading a list” is only true if the three v1 flows
look meaningfully different at a glance. They do, and not by decoration — by their
real structure: Readiness is a single node ending in a report and never writes;
Scene Polish is one node into a gate; First Act is a four-node chain with a
gate and a loop back for blocking annotations.
If a future flow duplicated another’s shape, picking by shape would quietly stop
working. Worth knowing before the flow set grows — flagged below.
Round 2 · call 3 · the guardrail
One state, computed twice — not two states kept in step
The hard guardrail is that the board and the rail must never be two systems. I did not
implement that as discipline. There is one run object, the board draws it, and the rail is
derived from it by deriveRail(), which reads the current
step out of the template and returns the pace, seat, loop state and step count.
The rail cannot contradict the board because the rail does not store a pace — it
asks for one. Both columns on the comparison page are driven by that single object,
and a test steps the run through every position and asserts the two agree at each one.
Round 2 · call 4
“Replace” is the wrong verb, and that is the answer
The board and the rail are not two views of one thing — they are two zoom
levels of one thing. The board says which step of the route; the rail says what is
happening inside that step. Neither is redundant, so replacing one with the other removes
information rather than removing a duplicate.
The honest version of “replaces” is therefore nesting: the rail moves
inside the open node, which is what it always was. I built that as option A and it
is the more elegant idea — but at 300px it costs a level of containment on the
busiest surface for no information gained. I recommend B, and because both are drawn
from the same state, changing later is a layout decision rather than a rewrite.
Round 3 — what changed under Residents
Round 3 · call 1
The two-cast problem does not get solved. It dissolves.
Round 1’s hardest visual problem was making Floor-cast and in-app-cast unmistakable
— a triangle and a dot, never colour alone. Residents-only removes the premise: there
is one agent model, so there is one mark. .cmark.resident keeps the
shape+text rule (a diamond mark, never colour alone — still has to survive greyscale
and a screenshot) but the system it belongs to is simpler by construction, not by discipline.
Round 1/2’s .cmark.app / .castbar.floor
rules are still in flows.css, unedited —
On the rail still renders them. That is honest history, not dead code to clean up.
Round 3 · call 2
The front door was the hero twice. This round it is not the door at all.
Keystack-first means most sessions never touch the front door —
a resident starts the run, and the human’s first frame in 4P is mid-route. The
actual design problem this round is orientation from a cold arrival, not a picker. Arrival answers four questions in one glance — which shape,
which step, who is in which seat, what needs you — and the door becomes the page you
visit when nothing is running, not the page everyone starts from.
Round 3 · call 3
Five shapes need five silhouettes, proven the same way as round 2
DEVELOP/DRAFT/REVISE/TRANSFORM/JUDGE reuse board.js unchanged —
a shape is a flow-document template exactly like round 1’s three flows, so the renderer
needed one small extension (a second lookup table, tried after FLOWS)
rather than a rewrite. The five signatures
(nodes|gate|loop) are 2|T|F, 3|T|F, 4|T|T, 4|T|F, 2|F|F —
asserted distinct the same way round 2 asserted the original three, because round 2 already
flagged that picking-by-shape silently breaks on a collision.
REVISE keeps the round‑1/2 revise‑loop (blocking annotations); TRANSFORM does
not loop — that is the one deliberate difference that keeps their otherwise-identical
node counts (4 and 4) from colliding.
Round 3 · call 4 · best craft this round
The proposal card is not illustrated. It is computed against a real roster.
“The Salt Line” has two residents — Noa writes, Wren edits — and no
Continuity resident. That is not staged for the demo: it is the actual roster
residents.js declares, and resolveCast()
checks every card’s required seats against it. DRAFT reads castable. REVISE and
TRANSFORM both read blocked, honestly, on the same missing seat — but not identically:
REVISE has a declared solo variant (Scene Polish, Editor only) so its fix offers a way to
run today; TRANSFORM has none, so its fix is only add a resident. Two blocks that look
the same until you read them is exactly the kind of flattening round 1’s readiness
report was built to refuse.
Busy is rendered as its own, third fact — not a weaker block. Wren mid‑run
elsewhere makes the DRAFT proposal queue, not fail; the same resolver call, one flag
different, produces a genuinely different verdict.
Round 1 design calls — unchanged
Call 1 · the readiness report
The honesty rules force a fourth verdict
The spec names three: PASS / WARN / BLOCK. But rule two — a check over
nothing says no data yet, never a confident green — cannot be expressed in three,
because a vacuous pass is a pass unless the absence of a measurement has its own
identity on the page.
So there are four. NO DATA is drawn colourless and dashed: not a severity, not a
failure, an absent measurement. It can never be mistaken for a green at a glance,
which is the whole point. The verdict banner also asserts the count first — a
“can run” conclusion is only reachable when a positive number of checks actually
returned something.
And the banner needs a third state, which I only found by testing it. My first build
drew the per-check honesty correctly and then committed the exact sin at the top of the page:
for the television case it printed “This flow can run” in the largest type
on the surface, because nothing was technically blocking — the gating check had
simply never been measured. A vacuous pass, promoted to the headline. The rule that fixes it:
NO DATA on a check that would otherwise gate the run is not a silent yes, so the banner
says “we cannot tell”. Three banner states, not two.
Call 2 · casting
The cast mark rides on the attribution, not the header
The spec puts the cast in the room header. I have done that — and then also put it on
every seat attribution, because a header-only mark fails the moment one note is
quoted, exported, screenshotted or pasted into Slack. At that point the header is gone and a
Floor-agent line reads as the product’s persona, which is exactly what risk 3
forbids.
The mark is shape and text, never colour alone — a triangle and the words
Floor agent, versus a dot and In-app. It survives greyscale print, colour
blindness and a screenshot.
Call 3 · the front door
Both doors are the same size
The fold ruling says Flows is the room’s front door, one user-facing feature. The
risk in building an objective picker is that it quietly becomes the only way in, and
freeform — which the spec says “remains exactly as it is” — degrades
into the fallback nobody chooses.
So the two doors are the same size and sit side by side, and the freeform card says in as
many words that it is unchanged. Flows is offered first by position, not by making the
alternative look like a downgrade.
Call 4 · where the report lives
The report opens onto the stage, not into a 300px dock
A readiness report with per-check fixes and a coverage chart does not fit the dock, and
squeezing it in would produce exactly the readiness theater the spec warns about —
a report too cramped to read is a report nobody reads.
So the dock summarises and the stage expands, which is the pattern the gate
tray and continuity lane already use in v2. No new surface, no new navigation: the same
Open gesture that already ships.
What I found in the corpus while building this
The spec names two blocking gaps — comedy 0 of 66, novels 0 of 38. Both are real; they
match the measured deep-tier table exactly. There is a third, and it fails differently.
| Form | Deep tier | How readiness should answer |
| Screenplays | 22 of 41 | Pass. Enough grammar to bring real precedent. |
| Plays | 29 of 50 | Pass. |
| Comedy | 0 of 66 | Block. The works are ingested; the deep analysis
that turns them into grammar has not run. Our gap, and we can say exactly how big it is. |
| Novels | 0 of 38 | Block. Same shape as comedy. |
| Teleplays | none ingested | No data — not block. There is
nothing to measure, so there is no percentage to draw and no coverage figure to report. |
Comedy and television are not the same failure and should not read the same way. Comedy has
66 works that stopped short of the deep tier; television has none to stop. Television is the
literal case honesty rule two exists for — the check whose set is empty. Both are on
the readiness surface; switch the form to see them.
Open questions — flagged, not guessed
§11 of the FLOWS_SPEC lists what Khoa had not yet decided as of round 1/2; §8 of
RESIDENTS_SPEC lists what is still open this round. Several of those change what I would draw,
and building this round surfaced three more of its own.
| Question | Who | Why it matters to the design |
| Does Flows v1 ship dark in 4P? Spec §11.3, recommended yes. | Khoa |
Decides whether the front door needs a “this is in testing” state at all. I have
not drawn one — if it ships dark to customers the door simply is not there, which
is cleaner than a disabled affordance. |
| The three Floor tokens resolve to zero access. Spec §8 names it as an open
precondition with two paths; the recommendation is re-scoping to project arrays now. |
Khoa · Adam |
The casting display shows grant state per seat, so the run cannot look ready when it is
not. But a cast that can never bind has no useful design state beyond “blocked”
— if path (b) is chosen instead I would want to redraw this, because an account-wide
grant should look wrong in the UI, not normal. |
| Is there a fourth verdict? My call 1 adds NO DATA to the spec’s three. |
Will |
It is forced by honesty rule two as written, but it is an addition to the spec’s
vocabulary and the build will encode it. Worth ratifying before Gab implements three. |
| Can a WARN be dismissed permanently, or only for this run? The spec says warns are
“visible and dismissible at the human’s choice” but not for how long. |
Will |
I have drawn per-run dismissal. If dismissal persists, a project with no synopsis stops
warning forever, which quietly defeats the check. |
| Does the readiness report get saved? It is pure read, so nothing forces it to
persist — but the wrap minutes are supposed to show the route taken. | Will |
If a run’s readiness result is not recorded, the minutes cannot honestly say what was
known at the start, and a dismissed warning becomes invisible after the fact. |
| docs/LIBRARY_SPLIT_EPIC.md is cited but not on the branch.
Spec §3 cites it for the “53 untyped collections” measurement. | Will |
Not a design issue — but the storage-typing rule leans on that measurement, and I
could not read it to check the flows collection would be the 54th. |
| Board replaces the rail, or sits above it? Both built, side by side at real
width. I recommend B (sits above). | Khoa |
The one Khoa asked to judge by eye. A is more elegant and stays viable if the dock widens
or the board moves to the stage; B is cheaper, keeps the shipped rail untouched, and is
easier to reverse. |
| Does picking-by-shape survive more flows? The three v1 shapes are genuinely
distinct. The backlog lists six more. | Will |
Competing Takes and Table Read would both be short chains and could read as the same
shape. If the set grows past ~6, shape stops being a sufficient discriminator and the
board needs a second cue — better decided before the set grows than after. |
| Where does the board live during a run — dock or stage? I have it in the
dock at 300px. | Khoa · Will |
A four-node flow fits. A longer one will not, and will either scroll horizontally or
shrink below legibility. The readiness report already opens onto the stage for this exact
reason; the board may need the same escape hatch. |
Should TRANSFORM declare a solo/reduced variant too?
Answered by the round‑6 types ruling. | Will |
“Reduced variant” now has a name and a home: it is a declared type.
TRANSFORM ships Basic (Writer → Editor) beside Full, so its
missing‑Continuity block on the proposal card offers the same
two‑way fix REVISE always had. Better than the answer I was asking for: the proposal card
no longer hunts for a preset with a hand‑written solo field, it
computes the largest declared type this roster can actually fill. |
Does the Tools panel’s pending-approvals badge need its
own surface? Answered by amendment 2. | Will |
Dissolved rather than decided: with no required-approval moments, there is no queue to
open. The lane became a Recent decisions record instead, and a record needs no
second surface — it links to the transcript that already holds the detail. |
| What does the human actually see in Keystack chat? The spec is explicit that
Keystack confirms the proposal and 4P renders it, but the confirm UI itself is outside 4P by
design. | Will · Khoa |
I drew the 4P side of the contract (the proposal card) in
full and treated the Keystack side as out of scope for this pack — but the card’s
credibility as “the contract for the run” depends on the two sides actually
agreeing, which I cannot verify without seeing the chat surface. |
| Does the no-required-approvals ruling retro-apply to the shipped v1/v2 gate tray?
I audited the older pages rather than assuming, and did not change them. |
Will · Khoa |
My read is that nothing there is a defect, for two distinct reasons worth separating.
The gate tray is per-tap hunk review — the canonical
rendering‑pull case §8 explicitly keeps in 4P, so it survives as a
convenience surface. The “blocking annotation” language is not an
approval requirement at all: it is the flow engine’s own loop condition (unresolved
contradictions keep REVISE looping), which no human tap resolves. What would be a
defect is copy claiming a run is halted until someone taps in 4P — I found none in
v1/v2, and fixed the two instances I had introduced in round 3. Flagging rather than
silently rewriting, since those pages are the record of what shipped. |
Scope and grounding
Design only. No build, no backend, no template authoring. Every colour, radius and
dimension is inherited from room.css, which is itself read from
fd-room.tsx, fd-agent-dock.tsx and
globals.css. Round 1/2 additions: a colourless slate for the
NO DATA verdict, and one provenance hue for the (now-historical) Floor-cast mark. Round 3
addition: one further provenance hue, --resident, for the single
mark type residents-only replaces it with — still used only for provenance, never for
status.
flows.css loads after room.css and
adds only; the v1 and v2 pages render identically with it absent, and round 3’s
additions render identically with round 1/2’s pages, verified the same way. Round 3 adds
one new script, residents.js — the roster and the casting
resolver every card on the proposal page and
Arrival actually computes against, not illustrative copy.
QA risk tier: T3 — non-executable mocks, no product behaviour change.
Names, pages and figures are illustrative. The verdicts, the fix actions, the route
semantics and the casting rules are the design. Corpus coverage numbers are measured and real.