# Flows: the one automation engine (DECISIONS 20b, 24, 25, 42)

Design input for turning Flows into a real workflow builder and folding every other engine into it, 2026-09-29. Read from the
app at `376fdca1` (the code is the authority), `docs/DECISIONS.md`, `plans/BOARD.md` and `docs/reminders-spec.md`. Nothing was run
and prod was not read: row counts below are quoted from the board and are not verified. Paths: `engine` =
`apps/api/src/flows/flow-engine.service.ts`, `types` = `packages/shared/src/flows/types.ts`, `schema` = `packages/db/prisma/schema.prisma`.

## 1. Today, in brief

| Part | What it does now | Evidence |
|---|---|---|
| Graph | `{ entryNodeId, triggers[], nodes[] (max 200), edges[], layout }`; an edge leaves a named port | `types:954-971` |
| Steps | message (text, media, card, gallery, a question, a `dynamic` block that fetches its text), template, condition, actions, randomizer, smart delay, start flow, AI hand-over, comment note; any other kind runs as a logged no-op | `types:867-935`, `engine:1007-1131` |
| Actions | set or clear a field, add or remove a tag, web request, start flow, LeadRat (save, update, note, set status) | `types:674-725` |
| Triggers wired | `dm` on WhatsApp (`basic-flows/basic-flow-engine.ts:276-282`), Instagram (`instagram/ig-dm.service.ts:362`), Messenger (`facebook/fb-dm.service.ts:180`); `post_comment` on Instagram and Facebook (`instagram/comment-pipeline.ts:2061-2063`) | grep of `surface:` callers |
| Triggers declared only | `story_reply`, `live_comment`, `contact_event`: no caller passes their surface; `contact_event` appears once, `flow-trigger.service.ts:231`. The picker says so (`trigger-catalog.ts:51,91,100`). The website chat cannot even be chosen (`types:506`) | `types:642-648` |
| Data between steps | Contact attributes and tags only. A question writes `contact.attributes` (`engine:2298,2330-2337`); a web request maps its response into attributes (`engine:2584-2594`); inserts read attributes (`engine:1362-1367`). `FlowSession.variables` holds engine internals only (`_retry_*`, `_callStack`, `_commentId`, `_test`, `_handoverResumeNodeId`) | `schema:1352` |
| Inserts | Two resolvers: `interpolateFields` with `{{x\|fallback}}` (`interpolate.ts:27-49`) and a private regex in request bodies that ignores the fallback (`engine:2605-2613`) | |
| Run | `FlowSession`: `contactId` required (`schema:1346`), at most one `active` per contact (partial unique index, `schema:1335-1339`); a second start joins the first (`engine:484-492`) | |
| Where it runs | Inside the api, on the inbound webhook (`flow-trigger.service.ts:564-566`). Waits, answer timeouts and send retries are BullMQ jobs on `flow-delay` (`flow-delay-queue.provider.ts:17`) that call `/internal/flows/*` (`internal-flows.controller.ts:28-60`); a 15-minute sweep expires runs and resumes hand-overs (`engine:782-854`). Loop guard: 100 steps (`engine:156,997`) | |
| Versions | One draft saved in place; each publish appends a version, newest wins; a running session keeps its version | `schema:1311-1318`, `flows-admin.controller.ts:224-227` |
| Gate and test | `Flow.state` off, shadow, live (`schema:260-264`). A builder test runs the draft on the tester's test contact, sends held (`flow-test-run.service.ts:83-106`, `engine:2833-2852`) | |
| Run log | `FlowRunEvent`: outcome, error, `detail` JSON per event (`schema:1372-1387`, `engine:2854-2872`). No step input or output: a web request logs its status only (`engine:2596-2598`) | |
| Errors | Message and template steps have an `error` port (`engine:1039-1058`) and retry a send Meta refused with "try later" (`engine:1542-1569`). Actions have none: a failed request or CRM write is logged and the run takes the default edge (`engine:2595-2601,2553-2554,1094-1104`) | |
| Sub-flow | Start flow jumps inside the same session with a call stack and returns to the caller (`engine:2678-2727,2778-2809`); nothing is passed in or out | |
| AI | `ai_handover` hands the thread to the agent and resumes after `sessionTtlHours` through the sweep (`engine:2740-2770,795-839`) | |

The other engines:

| Engine | What it is | Written by | Read by | Overlap with Flows |
|---|---|---|---|---|
| FlowRule (`flow_rules`, `schema:1132-1143`) | keyword or any-message trigger, then send template or add tag (`basic-flows/flow-rules.service.ts:6-29`) | `/api/flow-rules` (`flow-rules.controller.ts:17-44`, no web caller); one per QR code (`campaigns/qr-campaigns.service.ts:80-99`) | `BasicFlowEngine.dispatch`, WhatsApp only, and only when no flow claimed the message and the WhatsApp AI gate is off (`basic-flow-engine.ts:221-256`) | Complete: a WhatsApp message trigger with a template step and an add-tag action |
| BasicFlowEngine (`apps/api/src/basic-flows/`) | The WhatsApp inbound router: opt keywords, pause and opt-out hold, flows, AI, FlowRules, default reply (`basic-flow-engine.ts:171-264`) | code | the WhatsApp webhook | Only its FlowRule and default-reply legs are automation |
| AutomationRule "engine v2" (`schema:1145-1197`) | schedule (cron) or event trigger, AND/OR conditions, actions send template, add tag, start flow, notify, stay silent (`packages/shared/src/engine/types.ts:40-119`) | `/api/automation/rules` (`automation/automation.controller.ts:21-65`); no screen calls it (`routes/Automation.tsx:6-24`) | `dispatchEvent` (`automation.service.ts:334-347`) on events from `crm.adapter.ts:34`, `campaigns.service.ts:590`, `leadrat/internal-crm.controller.ts:39` and the quiet scan (`automation.service.ts:366-434`) | Its events are copied into Flows' contact events (`types:557-571`), but the bus never reaches a flow. Nothing fires a schedule rule (no `"schedule"` in `apps/api/src` or `apps/worker/src` outside tests). Its start flow is a no-op (`automation-effects.service.ts:42-45`) |
| Reminders | Removed 2026-09-28; rebuild plan `docs/reminders-spec.md` section 13 | | | |
| Campaigns (`schema:970-1010`) | One template to a segment: now, scheduled, or by API key | Campaigns screen | worker | Emits `campaign.finished`; not a rules engine |

Prod, per `plans/BOARD.md:23,74` and `docs/reminders-spec.md` section 14 (not verified): 2 flows, both off; 0 FlowRule rows;
AutomationRule rows `reminder-new` and `reminder-bulk`, due for deletion (`BOARD.md:133`).

## 2. Gaps against an n8n, Make or Activepieces class builder

| Capability | Today | Needs |
|---|---|---|
| Named step outputs | None; only contact attributes (section 1) | A stable `key` per step; every step's output stored and readable as `{{ steps.<key>.output.<path> }}` |
| Runs with no contact | `contactId` NOT NULL (`schema:1346`); one active run per contact across all flows | Nullable contact, a `subject` (a lead id), and runs that do not hold the conversation, so one staff member can have five |
| Schedule trigger | None in Flows; AutomationRule declares cron but nothing ticks it | A cron trigger in UAE time, one repeatable BullMQ job per published trigger, rebuilt on publish and on worker start |
| Inbound web request trigger | None | `POST /api/hooks/flows/:triggerId` with a secret stored hashed (the `Campaign.apiKeyHash` pattern, `schema:987-991`); the body becomes `trigger.*` |
| Contact events: created, tag added, field changed | Declared (`types:557-571`); no emitter (`contacts/tags.service.ts` emits nothing) | Emit on the existing bus (`automation/automation-events.ts:90-115`) from tags, field writes and contact creation; one flow dispatcher subscribes |
| LeadRat status changed | `crm.state_changed` is emitted (`internal-crm.controller.ts:39`) to engine v2 only | The same dispatcher; a LeadRat status read and a reassign action (`reminders-spec.md` section 13) |
| Story reply, story mention, reaction | Story reply is declared and handled as an ordinary message (`trigger-catalog.ts:91`); mention and reaction are not in the schema | Route the Instagram webhook's story and reaction fields to their own surfaces |
| Error routes and retries | Message and template only (section 1) | An `error` port and a retry policy (count, backoff) on every step that can fail; the error text readable by the next step |
| Per-step history | Outcome strings only | One step-run row per execution with input, output, error, attempt and duration, shown in the Runs tab and the test panel |
| Credentials | An encrypted `connections` table exists (`schema:416-428`, `connections/connections.service.ts:40-43`) but a step cannot use it; request headers sit in the graph, masked for non-admins (`flows/header-secrets.ts:3-18`) | A step names a connection; secrets never enter the graph or the run log |
| Branching, looping, merge | Branching by port (condition, randomizer, buttons); one edge per port (`graph.ts:114-124`); merge by wiring two edges into one step; no parallel paths, no loop over a list | A "for each item" step over a list output; parallel paths are not proposed (section 3) |
| Sub-flows | Start flow with a call stack, no inputs or outputs | Inputs mapped on the call; the sub-flow's last output returned as the step's output |
| Test run with sample data | Test on the draft with a test contact and held sends; the trigger is a typed message only | Pin a sample trigger payload (or one taken from a past run) and re-run one step alone with its stored input |

## 3. The model

| Concept | Meaning | Stored in |
|---|---|---|
| Flow | The named automation, its gate (off, shadow, live) and folder | `flows` (unchanged) |
| Version | An immutable graph once published; one editable draft | `flow_versions` (unchanged) |
| Trigger | What starts a run: a message, comment, event, schedule, web request, topic, or another flow | inside the graph, as today |
| Step | One node; has a `key`, typed input settings and named output ports | inside the graph; `key` is new and optional (defaults to the id) |
| Run | One execution; holds the conversation or runs in the background | `flow_sessions` (extended) |
| Step run | One execution of one step: input, output, error, attempt | `flow_step_runs` (new) |
| Waitpoint | Why a run is paused and what resumes it | `flow_waitpoints` (new) |
| Connection | A named, encrypted credential a step may use | `connections` (reused) |
| Piece | One integration as a typed package of triggers and actions (LeadRat, Meta, HTTP, AI) | code, `packages/shared/src/flows/pieces/` |

Data, additive only (no column dropped, no row rewritten):
- `flow_sessions`: add `holds_thread boolean NOT NULL DEFAULT true`, `trigger_kind text`, `trigger_payload jsonb`, `subject jsonb`,
  `parent_session_id uuid`. Relax `contact_id` to nullable. Rebuild the one-active-per-contact index as
  `WHERE status='active' AND holds_thread`, so every existing row keeps today's behaviour and a background run never mutes the AI.
- `flow_step_runs`: `id, session_id (cascade), flow_version_id, step_id, step_key, iteration, status (running, ok, failed, skipped,
  waiting), input jsonb, output jsonb, error, attempt, started_at, ended_at`; unique `(session_id, step_id, iteration)`. Outputs
  over a seeded size are cut and marked; retention is a seeded setting, as `automation.retention` is today.
- `flow_waitpoints`: `id, session_id, step_id, kind (reply, delay, answer_timeout, agent, web_request, event), resume_at, match jsonb,
  resolved_at`; the BullMQ job id is the waitpoint id, which keeps today's one-job-per-arming rule (`flow-delay-queue.provider.ts:78`).
- `ConnProvider`: add `http` (an enum value is additive). `qr_campaigns`: add nullable `flow_id`.
- `FlowRunEvent` stays as the human-readable log; the step-run row is the data.

The engine loop:
1. A trigger creates the run. A message or comment still claims synchronously in the api, because the webhook must know at once
   whether a flow owns the message (`flow-trigger.service.ts:319-381`); its first steps run inline, as now, for chat latency.
2. Every other start (schedule, event, web request, topic) and every resume adds a job to a new `flow-run` queue. The worker stays a
   thin shim calling `/internal/flows/run` (the `flow-delay` pattern, `apps/worker/src/processors/flow-delay.processor.ts:18-27`).
3. Replay-and-skip: the engine loads the run's step runs, walks from the cursor, and never re-executes a step whose row for that
   iteration is `ok`; it reuses the stored output. A crash between two steps resends nothing.
4. A step that must wait writes a waitpoint and returns. A reply, a timer, the agent finishing, an inbound web request or an event
   resolves it; an event waitpoint also cancels ("the lead wrote to us" ends a reminder run waiting in a delay).
5. Failure: retry per the step's policy, then the `error` port, else the run fails with the step run showing why.

References: `{{ steps.<key>.output.<path> }}`, `{{ trigger.<path> }}`, `{{ contact.<field> }}`, and bare `{{field}}` meaning a contact
field, as today, with `|fallback` everywhere. One resolver, `interpolate.ts`, replaces the private regex in the engine. Publish
refuses a reference to a step that is not upstream. A Transform step evaluates JSONata over `{ trigger, steps, contact }` with a
time limit. A Code step runs JavaScript in isolated-vm: no network, 1 s, 32 MB (fork 3).

The AI as a step: `ai_handover` stays as it is (the AI takes the conversation). A new AI task step asks the one agent (ruling 15)
for a structured answer (classify, extract, draft) through the agent's own model connection, as the intent check does
(`flow-trigger.service.ts:521-523`), under a seeded hourly ceiling like `aiIntentMaxChecksPerHour` (fork 4). It never messages
the customer by itself; a message step does.

A Topic's destination as a flow: `LeadDestination.kind` is `leadrat` or `recruitment` (`packages/shared/src/agent/agent-rows.ts:193-197`)
and nothing reads topics yet (`agent-rows.ts:207`; no reader in `apps/api/src`). Add kind `flow` with a flow id. When the AI files an
enquiry under that topic, a background run starts with `trigger.topic` (topic id, summary, captured fields); it saves the lead,
assigns and notifies, while the AI keeps the conversation.

## 4. Folding the other engines

| Engine | Becomes | Existing rows | Order |
|---|---|---|---|
| AutomationRule | Event rule: a flow with a contact-event trigger; schedule rule: a schedule trigger; notify: a "notify staff" step; send template, add tag: the same steps; stay silent: dropped | Converter with a dry run that prints each row and the flow it would make; flows land `off` (gate flips are his); old rows set inactive, never deleted; `automation_runs` kept to its retention. The two reminder rows are not converted | 1st, after stage 3 |
| FlowRule and QR codes | A flow with a WhatsApp message trigger (keyword to `contains`, any message to an empty `all` group; matching is case-insensitive, `trigger-match.ts:77-78`, harmless for QR tokens) and template and add-tag steps. QR create makes a flow and fills `qr_campaigns.flow_id` | Same converter; `flow_rules` rows kept inactive | 2nd |
| BasicFlowEngine | Stays as the WhatsApp router for opt keywords, pause and AI routing (compliance, never switchable by a flow edit). Its FlowRule leg is removed after the fold; the default reply becomes a flow trigger "message nobody answered" | The `messages.autoReplies` setting's default reply becomes one flow, `off` until he switches it | 3rd |
| Reminders | Section 13 of the spec: "new lead nudge" and "older lead nudge" flows on a LeadRat status trigger, background runs with the lead as subject, a "send to a staff member" step, UAE time-of-day waits, event waitpoints for "lead wrote", "status moved", "marked done"; the digest on a schedule trigger | None (`pending_followups` had 0 rows, spec section 14) | 4th, needs stages 2 to 4 |
| Campaigns | Unchanged as a sending screen; flows gain "replied to a campaign" and `campaign.finished` triggers (fork 1) | None | with stage 3 |

## 5. Stages, each shippable

| Stage | What ships | Proof |
|---|---|---|
| 1. Step runs and named outputs (the foundation) | `flow_step_runs`; the engine writes input and output for every step (a web request's response body included); step `key`; `steps.*` and `trigger.*` in the one resolver; the Runs tab and test panel show each step's input and output | Full api and shared suites green; a new case in `apps/api/src/flows/flow-engine.e2e.test.ts`: a web request to a stub returns `{ name }`, the next message says `{{ steps.fetch.output.name }}`, and the held test message carries the name |
| 2. Background runs, schedule and web request triggers | Nullable contact, `holds_thread`, `subject`, the `flow-run` queue, cron and inbound web request triggers | A shadow flow on a 5-minute schedule writes a run with no contact; a `curl` to its hook URL with the secret starts a run whose `trigger.*` shows in the step input, and a wrong secret gets 401 |
| 3. Event triggers | Dispatcher on the event bus; tag, field and contact-created emitters; LeadRat status; story reply routed; Topic destination `flow` | Tagging a test contact starts a run; a `crm.state_changed` post to the internal endpoint starts a run with the status in `trigger.*` |
| 4. Error ports, retries, connections, waitpoints | Error port and retry policy on every failing step; `http` connections; `flow_waitpoints` replacing the implicit pauses | A request to a stub answering 500 retries N times then takes the error port; a marketing read of the graph holds no secret |
| 5. The fold | Converters (dry run first), the reminder flows, removal of the FlowRule leg and the unused `/api/automation` and `/api/flow-rules` routes after he confirms | Dry run lists every row and its flow; after, no active `automation_rules` or `flow_rules` row, and the new flows' runs match the old run counts over a week in shadow |
| 6. Builder depth | Transform (JSONata), Code (isolated-vm), for each, sub-flow inputs and outputs, AI task step, pieces for LeadRat, Meta and HTTP | Per step, a unit suite plus one test run on the canvas |

## 6. Forks only the owner can decide

1. Campaigns stay their own screen and gain flow triggers, rather than becoming a flow step. Recommend: stay; a bulk send with pacing, stats and an API key is one step, not a workflow, and ruling 20b would otherwise delete a working screen.
2. Ruling 24 says the engine never composes a conversational reply, yet flows answer messages by trigger (ruling 25). Recommend: rewrite 24 to "a flow speaks only where a trigger he set claims the message; the AI owns every other inbound".
3. A Code step runs staff-written JavaScript on the server. Recommend: admins only, isolated-vm with no network, 1 s and 32 MB; the capability n8n and Activepieces users expect, with the risk contained.
4. An AI task step spends model calls on every run. Recommend: allow it under a seeded hourly ceiling shown on the step, as the AI-intent trigger already does.

Decided here (main's call, reversible): extend `flow_sessions` rather than add a second run table; conversational starts stay
inline; converted flows land `off`; opt-in and opt-out confirmations stay outside flows; no parallel paths; our own engine and piece
SDK on the Activepieces pattern (its `LICENSE`: MIT outside `packages/ee/` and `packages/server/api/src/app/ee`; its
`packages/web` pins `@xyflow/react` 12.3.5 and `packages/server/api` `bullmq` 5.61.0, read 2026-09-29), keeping our
`@xyflow/react` ^12.11.0 (`apps/web/package.json:18`). No n8n, AGPL or Inngest code (their licences were not re-read here).
Stale text found: `types:20-23` still names a questionnaire graph and the n8n `Flow` registry (dropped, `schema:1268-1271`);
`types:550-556` says flows and rules share one bus, which no code does.
