All work
PLAN ZINNIA / AI AGENTS & INTERNAL OPERATIONS

Building AI agents
for Plan Zinnia.

I built four internal agents to scaffold vendor categories, audit new work, monitor platform health, and draft vendor upsells—with review and evidence built into the workflow.

MY ROLE

Founder & Product Lead / Agent design and implementation

CONTEXT

Two-sided wedding marketplace

SCOPE

Developer workflows, platform monitoring, and revenue operations

PREVIOUS WORKFLOW~Half a day

Coordinated across engineering, design, and ops

NEW WORKFLOWOne command

Followed by review

BUILT-IN VISIBILITYPreview + QA

Screenshots, QA records, and scheduled reports

DEVELOPER WORKFLOWS / BUILDER + QA
npm run builder add-vendor <type> --qa
01 / BUILDER AGENT

Scaffold the category

Branch checkpoint → site + Figma context → bounded generation → migration + forms

Preview screenshot + Builder log
02 / QA AGENT

Audit the preview

Runtime + network checks → responsive overflow → accessibility audit → severity summary

Screenshot + QA record
Human review Inspect the scaffold and its QA findings before approval.
PLATFORM OPERATIONS / HEALTH + UPSELL
03 / HEALTH AGENT

Monitor the running platform

Daily checks → explicit severity → email policy → admin + leadership report

Health report + history record
04 / UPSELL AGENT

Prepare vendor nudges

Monday snapshot → vendor-specific drafts → cooldown + cap → admin review

Draft email + human-sent nudge
Operator judgment People decide how to resolve findings and which vendor messages to send.
Four agents support two workflows. Builder and QA handle category setup; Health and Upsell surface ongoing operational work for human review.
01 / THE PROBLEM

Adding a category meant
coordinating several systems.

Expanding beyond venues to DJs, hair and makeup, or officiants required more than another marketplace listing.

Each new vendor category touched a Supabase enum, three React form directories, admin approval tables, a Notion tracker, and a QA pass. That was roughly half a day of coordinated work split across engineering, design, and ops.

I first built the Builder and QA agents to collapse that sequence into a single CLI command with automated checks. I then added Health and Upsell agents to monitor the running platform and prepare weekly vendor outreach. The goal was to reduce repeat setup work while keeping the team able to inspect what changed.

THE DESIGN CHALLENGE

Make category expansion easier without losing design consistency, recovery paths, or ops visibility.

02 / THE BUILDER AGENT

Turn a vendor name
into a reviewable scaffold.

A Node CLI composes six specialized steps, with explicit responsibilities and visible fallbacks.

The whole chain can run unattended. The LLM handles a constrained generation task; branch protection, database scaffolding, preview capture, and reporting remain deterministic.

01

Protect the working tree

The CLI refuses to run outside the dedicated builder-dev branch. A labeled checkpoint precedes any write, so a failed run can roll back the working tree.

02

Read a sibling experience

A site reader fetches the integration environment’s /venues route. Cheerio extracts headings, labels, and buttons as a structural reference for the new category.

03

Bring in the design system

The Figma REST API supplies tokens, components, and frame metadata. The agent writes reports/design-map.json and includes it as context for generation.

04

Generate a bounded set of artifacts

GPT-4.1-mini receives one structured prompt and returns JSON containing a Supabase migration, Form.tsx, index.tsx, and a typed servicesConfig.ts. Missing credentials or an unparseable response trigger deterministic templates.

05

Scaffold the database changes

The migration adds a vendor type to the vendor_type enum, a jsonb metadata column, a vendor_types row, and an admin_approvals entry. Idempotent guards protect repeated runs.

06

Capture the preview and log the run

Puppeteer captures the vendor’s integration preview. A Builder tracking database in Notion records the vendor name, branch, migration path, result, and screenshot for ops review.

03 / THE QA AGENT

Check what the Builder ships
before a reviewer opens it.

The QA Agent runs automatically with --qa after a Builder run, or independently against any vendor route. Headless Puppeteer loads the integration URL and collects the signals a reviewer needs.

Runtime and network health

Collect console errors, uncaught pageerror events, and failed network requests during page load.

Responsive overflow

Compare scrollWidth with innerWidth on the page and main content to flag horizontal overflow.

Accessibility

Run @axe-core/puppeteer to collect accessibility violations for review.

Severity and evidence

Combine accessibility violations, runtime errors, and network failures through a severity heuristic. Log the screenshot, counts, severity, viewport, and mode.

A thin Playwright layer in qa/tests/smoke.spec.ts runs scheduled checks against the public homepage. Builder-triggered audits and scheduled smoke tests write the same record format into a dedicated QA Notion database.

The Builder log and QA findings give ops and admin reviewers a visible trail from the scaffold to its audit results.

Builder and QA run example for a videographer category, showing generated files, preview capture, accessibility findings, runtime checks, and the structured QA summary
Builder run and QA audit example for a videographer category. Scroll inside the preview to follow the scaffold through its audit and structured summary. Open the full image.
04 / THE HEALTH AGENT

Make silent failures
visible to the team.

In September, several systems had quietly broken: Prerender had lapsed for three months, five edge functions had never been deployed to production, and new-message emails had stopped going out.

Nothing was consistently checking the running platform. I built report-platform-health, a Supabase edge function that pg_cron runs daily, to email a structured report to admins and leadership.

01

Check the platform in one pass

Compare deployed functions with the repo and migration state with production. Check Prerender render health, Stripe subscription integrity, critical bot-fetch targets, and vendor activation funnels.

02

Assign severity with explicit rules

Classify each check as fail, warn, info, or ok. Critical-function allowlists make severity a defined rule rather than an impression.

03

Make the email policy part of the monitor

FAIL sends an email every run until it clears. WARN sends on change or on Mondays. Every Monday sends a full summary, even when all checks are clear, so a quiet monitor cannot disappear indefinitely without a missing expected report.

WHY THE ALL-CLEAR MATTERS

A silent monitor can look exactly like a dead one. The Monday summary gives the team a recurring signal that the checks are still running.

The agent only persists its own history record. Database reads go through a Postgres function declared STABLE, which rejects writes within that function.

Sample Plan Zinnia health report showing one failing check, two warnings, resolved issues, twelve passing checks, and suggested follow-up actions
Sample Health Agent report design: severity, changes since the previous run, check details, and suggested follow-up actions in one digest. Scroll inside the preview to see the full report. The displayed findings illustrate the report format. Open the full image.
05 / THE UPSELL AGENT

Prepare the weekly outreach
for an admin to review.

The next operational task was deciding which vendors to nudge and what to say. I built draft-vendor-upsells to do that preparation every Monday.

The agent reads each Inquiry-plan vendor’s own data through upsell_draft_snapshot(): couples who reached out in the last 30 days, upgrade-intent signals, and time since the last nudge. This snapshot function is also declared STABLE.

Ground the draft in actual usage

Use the vendor’s own inquiry counts and wording aligned with the Inquiry plan they saw at signup. A draft might mention “6 couples reached out to you in the last 30 days”—an illustrative example, not a reported outcome.

Limit repeated outreach

A 14-day nudge cooldown and a weekly draft cap prevent an unattended review queue from producing a flood.

Keep the send with a person

Email drafts to admins, with leadership copied only in production. A person reviews, edits, and sends from Admin > Usage > Send nudge.

Record the human action

The send records nudge_sent. Subsequent weekly reporting can use that record to show whether nudging is working.

A week with no drafts produces no email. A failed snapshot still sends an email, so a broken job cannot pass as a quiet week.

Sample weekly Upsell Agent digest with four vendor-specific drafts for human review, inquiry statistics, cooldown information, and follow-up reporting
Sample Upsell Agent digest design. Scroll inside the preview to see the vendor-specific drafts and follow-up reporting. The displayed figures illustrate the format. Open the full image.
06 / ARCHITECTURE DECISIONS

Give the model a bounded job
inside an auditable workflow.

I designed the four agents around bounded responsibilities, named integrations, and explicit side effects.

The model generates boilerplate from strong structural references: the enum name, design tokens, and the shape of an existing vendor experience. Git state, migrations, screenshots, and Notion logging sit around that generation step as deterministic operations.

Use existing product patterns

Live page structure and Figma context guide new scaffolds toward the established experience.

Make recovery explicit

A checkpoint and automatic working-tree rollback limit the cost of a failed run. Database changes use idempotent guards.

Keep fallbacks predictable

Deterministic templates let scaffolding continue when model credentials are missing or generated JSON cannot be parsed. Fallback behavior keeps dependencies from becoming a single point of failure.

Preserve human visibility

Preview captures, logs, and structured QA findings make the output inspectable before admin review.

This architecture let me put category setup behind a one-line CLI without the ops team losing visibility into what it produced.

Apply the same boundaries to scheduled operations

Health and Upsell run on pg_cron schedules with service-role credentials. Their database reads use STABLE functions; their only persistent writes are their own history records. Reports and drafts go to people who can act on them.

The database restriction applies inside each read function. The surrounding workflow keeps operational actions with admins: Health surfaces problems, and Upsell prepares messages for a human to send.

Both agents also define what email silence means. Health sends recurring summaries and repeated failure reports. Upsell stays quiet when there are no drafts, but reports a failed snapshot.

Normal Builder run with the LLM available, generating vendor artifacts and logging a successful scaffoldFallback Builder run without an OpenAI key, generating deterministic template artifacts and logging the scaffold for review
07 / OUTCOMES

Less coordinated setup,
with evidence attached.

  • New-vendor-category onboarding moved from roughly half a day of cross-functional work to a single command plus review.
  • Every scaffold has visible QA findings before admin review. Builder-triggered audits and scheduled smoke tests use a shared QA record format.
  • Deterministic fallbacks help the pipeline continue through API outages, missing keys, or stale design context.
  • Platform checks became a scheduled daily report, with a full Monday summary and repeated notifications for unresolved failures.
  • Weekly vendor outreach preparation became a Monday inbox item with usage-based drafts, cooldowns, and human control over sending.

The time comparison describes the onboarding workflow. It does not imply a measured end-to-end runtime or remove the need for review. Health and Upsell outcomes describe the operating workflow; no measured revenue uplift or reduction in incidents is claimed.

RELATED WORK

See how I designed and built
the marketplace these agents support.

Explore Plan Zinnia