AI Tools
Conversational Design Creator
The Conversational Design Creator captures a team's tacit conversational design reasoning and turns it into a structured corpus that an AI can use to generate and audit outputs. It is designed to remove senior-review bottlenecks, preserve brand and design conventions, and make AI-generated conversational experiences traceable to specific team rules. The demo shows brief intake, clarifying questions, rule-cited generation, audit feedback, revision, and export into a complete specification for engineering.
The problem
Companies now ship conversational AI everywhere, but the reasoning behind good conversational design — why a fallback cascades a certain way, when to escalate, what brand voice applies per channel — lives in expert reviewers' heads, scattered Slack threads, and Figma comments. Review is the only enforcement mechanism, so quality collapses when the reviewer is busy, absent, or nonexistent. AI generators make this worse: they default to generic "professional B2B tone" because they have no idea what a company's tone actually is. The downstream cost is measurable — 60% of consumers report negative chatbot experiences, only 30% of companies enforce brand guidelines, and RAG-grounded chatbots don't help because they encode product knowledge, not conversational reasoning.
The solution
The Conversational Design Creator captures a team's tacit conversational design reasoning as a structured, machine-readable corpus, then uses it to generate and audit design artifacts. A producer enters a brief; the Creator generates conversational design drafts grounded in the corpus, with each element citing the specific rules applied. The distinctive architecture is a two-mode corpus from a single source of truth: settled convention rules fire silently with citations, while judgment-call rules surface alternatives with reasoning for the producer to decide. This operates at the design layer — flows, decision points, policies — upstream of runtime chatbot platforms, and its output feeds into tools like Voiceflow and Sierra rather than competing with them.
How it works
The workflow moves from brief intake through a confirmation gate — no generation without producer sign-off — to corpus-grounded draft generation, per-unit review with visible citations, and export into a complete specification for engineering. Each design unit is tagged as new, modifies, replaces, or preserved, and carries inline citations from the convention rules applied. At judgment-call rules, a path-discovery step presents corpus-grounded alternatives; the producer picks, and dependent convention rules cascade silently. Regeneration supports two behaviors — convention rule push-back when a producer overrides a settled rule, and judgment-call refinement — so the AI never silently rejects. The capstone corpus is hand-curated at 15–20 rules across 3–4 categories, calibrated to a real multilingual WhatsApp system at a global civic advocacy non-profit.
Who it's for
This is a B2B product with three layered roles. Primary end users are design-team-agnostic — conversation designers, UX writers, generalists, engineers building conversational features, PMs, contractors, and AI-tool operators — anyone producing conversational work who needs it grounded in company reasoning. The buyer is a senior leader with conversational AI transformation budget: head of design, VP of product design, design ops lead, a VP or director of conversational AI/CX, CTO, or head of product, varying by company maturity. Secondary beneficiaries are the wider product organization and, ultimately, the customers interacting with the conversational products being built.
Why it matters
The conversational AI market is growing at roughly 21% CAGR toward $82B by 2034, and the design layer — approximated at 5–10% of that spend — is a $1–2B market in 2026 with no direct competitor. Delivery is high-ticket, consulting-led: $50K–$500K per engagement, because the AI-accelerated extraction methodology is the irreproducible IP and the corpus is the asset that justifies the price. Defensibility is built around a 12–24 month window before adjacent brand-voice governance tools (Acrolinx, Averi, Markup AI) could pivot — and adding a second corpus mode is a structural rebuild, not a feature. The business is pre-launch, targeting 3–5 engagements in year one, scaling toward 10–20 as the methodology matures.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Nathalia Carvalho | |||||
| Your Product: | Conversational Design Creator | |||||
| Your Industry: | B2B SaaS — Conversational Design Operations | |||||
| Date: | June 13, 2026 | |||||
| AI Opportunity Statement: | Our business operates in conversational design segment and creates value by encoding teams' design reasoning into a two-mode corpus — settled rules the AI applies silently, judgment-call rules the AI surfaces with reasoning at decision moments. To deliver this value, our product offers a consulting-led extraction methodology paired with an AI generator grounded in the resulting corpus that meets the needs of organizations shipping conversational experiences across deterministic flows, AI agents, and hybrid flows. These customers are primarily organizations with fragile or absent conversational design conventions, where producers (designers, generalists, AI-tool users) experience review bottlenecks, AI outputs that violate team reasoning, and decision gridlock at judgment moments as they try to achieve convention-consistent conversational design without senior-reviewer dependency at every step. To address these pain points, we propose an AI-powered solution that generates corpus-grounded conversational design specifications — applying convention rules silently with citations, surfacing judgment-call alternatives at decision moments, and exporting a portable flow diagram for handoff — delivering value to producers (senior-designer reasoning on demand), design leaders (consistent quality without review bottlenecks), and the buyer (visible conversational AI ROI). | |||||
| Production site: | https://conversational-design-creator.vercel.app/ | |||||
| Password: | aviv-creator | |||||
| 4-min video: | https://youtu.be/kxLBv1jbHvA | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | B2B SaaS — the conversational design operations segment, operating at the design layer of conversational AI infrastructure: the layer where conversational design artifacts (flows, structures, decision points, brand-voice rules) are produced before engineers implement them in chatbot or agent platforms downstream. The segment is tool-agnostic across deterministic flows (WhatsApp, IVR), AI agents, and hybrid flows (the dominant pattern). It serves companies investing in conversational AI where design quality affects business outcomes — design-team-agnostic: organizations with dedicated design teams, organizations shipping through engineers/PMs/contractors, and organizations using AI generators whose outputs need grounding in company-specific reasoning. This sits upstream of most conversational AI tooling today. Chatbot platforms (Voiceflow, Botpress, Cognigy) and AI agents (Sierra, Decagon, Chatbase) generate runtime responses — what end users see when interacting with the chatbot — typically using company knowledge bases. RAG-grounded runtime chatbots are now standard infrastructure — 71% of organizations use generative AI in at least one business function (McKinsey 2025) — so "trained on your knowledge base" is no longer a differentiator. The design layer above remains largely unaddressed: there is no equivalent tooling that produces conversational design artifacts grounded in company-specific conversational reasoning. | |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | TAILWINDS Conversational AI is growing at 21–25% CAGR through 2034 (Fortune Business Insights), with structural enterprise commitment: 91% of businesses 50+ use AI chatbots (Master of Code 2026); 75% of organizations will use LLMs for customer service by 2026, up from 10% in 2023. Brand voice consistency in AI outputs is a measured business problem: only 30% of companies enforce brand guidelines (Contentstack 2026); brand consistency drives 10–33% revenue lift (Lucidpress). An adjacent category — AI brand-voice governance (Acrolinx, Averi, Markup AI, Contentstack Brand Kit) — has formed around this in marketing content. The equivalent gap in the conversational design layer is unaddressed. The conversational AI failure rate is a direct demand signal sharing one root cause — conversational AI deployed without encoded design reasoning: 60% of consumers report negative chatbot experiences (Zendesk); 28% satisfaction (Gartner); CSAT gap AI-handled vs human exceeds 15–20 points (Chatbase 2026); AI customer service failure rate ~4× higher than AI use in general (CNBC). The conversation designer discipline is being formally recognized: Gartner projects 42% of organizations will hire for AI-focused CX roles by 2026; Brad Frost's "Agentic Design Systems in 2026" talk (Dec 2025) names the category in real time. HEADWINDS RAG-grounded chatbots are now standard infrastructure, not a differentiator. The Creator must clearly articulate that it operates at a different layer (design artifacts upstream, not runtime responses) and encodes a different content type (conversational design reasoning, not company knowledge). The conversation design discipline lacks a clean industry-standard name; buyer education is real. The most strategically relevant headwind is competitive — addressed below. KEY COMPETITORS The Creator operates in a layer with no direct competitor today: AI-generated conversational design artifacts grounded in a company-specific corpus running in two modes — settled rules applied silently with citations, judgment-call rules surfaced for producer decision. The surrounding market has six adjacent groups, and one is positioned to enter the layer first. Most likely first mover: Acrolinx (AI brand-voice governance). Largest installed enterprise base in AI content governance, most credible methodology for encoded rules. Their architecture validates content against a corpus of brand and editorial rules — solving the validation half in marketing content. Pivoting to (a) conversational design artifacts and (b) generation rather than validation is plausible at 12–24 months. But adding the two-mode corpus posture (silent application + judgment-call surfacing) is a structural change, that combined with learning conversational design methodology, is the timeline floor. Averi, Markup AI, and Contentstack Brand Kit face the same architectural pivot and the same timeline. Defensibility argument — built around the 12–24 month window: - Methodology IP. The AI-accelerated extraction process is the high-value, hard-to-replicate practice. Reproducing the corpus requires the methodology that produced it. - Reference momentum. First customers create case studies that anchor commercial position before adjacent categories catch up. - Architectural distance. The two-mode corpus is structurally different from rule-validation architecture. Adjacents can't reproduce it without rebuilding their corpus model. The other five adjacent groups are complementary, not competitive: - AI chatbot platforms / agent builders (Voiceflow, Botpress, Cognigy, Kore.ai, Rasa): build-your-own runtime agents. The Creator's design-layer output feeds INTO these. Complementary. - AI agents-as-a-service (Sierra, Decagon, Replicant): vendor-operated runtime agents downstream of the design layer. Potential clients. - RAG / knowledge-grounded chatbot infrastructure (Chatbase, StackAI, ManyChat): runtime knowledge grounding. Adjacent layer, different content type. - Conversation design education (Conversation Design Institute): methodology precedent, not commercial competitor. - Vertical AI design tools: no AI design tool specifically for conversational design exists today — the position the Creator fills. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | The product sits at the intersection of four measured markets with overlapping growth signals: 1. Conversational AI market (TAM source): $17.97B in 2026, projected $82.46B by 2034 at ~21% CAGR (Fortune Business Insights). Generative AI chatbot sub-segment growing fastest at 31% CAGR, reaching $113B by 2034. 95% of customer service interactions projected to be AI-powered by 2026 (Adobe Digital Trends). 2. AI consulting services (parent market): $11.07B in 2025 → $90.99B by 2035 at 26.2% CAGR (Future Market Insights, Aug 2025). 78% of successful AI deployments use external partners (McKinsey 2025) — validates the consulting-led delivery model. 3. AI brand-voice governance (adjacent SAM proxy): the most directly adjacent category, with momentum from Contentstack Brand Kit, Acrolinx, Averi, and Markup AI addressing brand consistency in marketing content. 71% of organizations use generative AI in at least one business function (McKinsey 2025). 30% of companies actively enforce brand guidelines, 70% don't (Contentstack 2026) — a measured gap that scales with AI content volume. The Creator addresses this gap in the conversational design layer, where no equivalent solution exists. 4. Conversation designer talent (SAM signal): avg salary $63K, senior up to $267K+ at Meta (Glassdoor 2026); 2,366 active US postings; Gartner projects 42% of organizations will hire for AI-focused CX roles by 2026. TAM / SAM / SOM ESTIMATE TAM — the design-layer portion of conversational AI spend, approximated as 5–10% of the overall conversational AI market: ~$1–2B in 2026, growing to $5–8B by 2030–2034. SAM — companies with active conversational AI investment + visible design quality pain (estimated 15–25% of TAM): $200–500M in 2026, growing to $1–2B by 2030–2034. SOM — boutique consulting practice capacity: 5–20 engagements per year at $50K–$500K each → $250K–$10M annual revenue with strong operating leverage from methodology IP. Scaling to 20+ engagements requires senior consultants who can lead extraction. The specific intersection — AI-generated conversational design grounded in encoded company-specific reasoning + extraction methodology — is sub-radar today. This supports market-entry timing rather than weakening the case: the category is forming in real time, the conversation designer discipline is being formally recognized, and buyer pain is acute and well-measured. | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Startup — pre-launch, proof-of-concept stage. Current focus: demonstrating the methodology and AI tooling capability on a working case study, calibrated to a real-world multilingual conversational design system on WhatsApp at a global civic advocacy non-profit. Next milestones: first commercial pilot with an external client team, followed by scaling into a boutique consulting practice. Year 1 target: 3–5 engagements. Year 2–3 target: 10–20 engagements per year as methodology and tooling mature and senior consultants are hired. | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | High-ticket, consulting-led delivery — methodology engagement + AI tooling deployment. Three layered deliverables per engagement: 1. Facilitated extraction. The practice's methodology — AI-accelerated to abstract patterns and define rules faster than manual analysis — extracts conversational design thinking from whatever sources exist in the client's reality (conversation designer articulation when available, artifact analysis of existing flows, past review comments, customer service transcripts, cross-designer pattern detection, AI generator output analysis). Output: a machine-readable corpus structured by category. 2. Creator deployment (Phase 1) + built-in QA gate (Phase 2). Phase 1: the Conversational Design Creator, calibrated to the client's encoded corpus, generates conversational design drafts from briefs. Phase 2 (post-pilot): a built-in QA gate, reading from the same corpus, validates each draft against company conventions before handoff. One canonical source of truth across both phases. Standalone Auditor available as an alternative deployment for validation in clients that don't need generation. 3. Ongoing corpus evolution. Quarterly corpus health reviews; new conventions surfaced from production patterns and Apply/Override/Flag signals; expansion to new sub-segments, channels, or languages as client investment grows. Available as optional retainer ($25K–$100K/year). AI plays two distinct roles in the model: (a) accelerant for the consultant during the extraction phase, (b) the Creator + QA gate during the deployment phase. Both read from the same canonical corpus. Pricing: $50K–$500K per engagement based on scope and corpus complexity. Larger engagements (Fortune 500 clients, multi-segment, multilingual) at the upper end. The model deliberately rejects self-serve SaaS as the primary path. The corpus is the high-value asset, and without expert-led extraction, clients cannot produce a corpus rich enough for the Creator to generate convention-compliant drafts. The extraction methodology is the IP; the corpus is the irreproducible asset that justifies the engagement price. Adjacent precedents validate the consulting-led, high-ticket model: Knapsack and Big Medium for design system transformation at Fortune 500 scale; enterprise professional services from Cognigy ($300K+/yr) and Kore.ai for conversational AI transformation; Sierra and Decagon's "low six figures" pricing for vendor-operated agents. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2B. Three layered roles in every engagement: - BUYER (the contracting decision-maker). Senior leader with conversational AI transformation budget and downstream accountability — typically head of design, VP of product design, design ops lead, VP/director of conversational AI/CX, CTO, or head of product. The buyer profile varies by company maturity: design leadership in design-mature companies; product or CTO leadership in companies shipping conversational AI through engineers/PMs. - END USERS (the people using the Creator daily). Design-team-agnostic — conversation designers, UX writers, junior/mid-level designers, generalists, engineers building conversational features, PMs, contractors, AI-tool operators. The product serves whoever produces conversational work in the company. - SECONDARY BENEFICIARIES. The downstream product organization — engineers implementing conversational features, content/translation teams localizing flows, brand/marketing teams reviewing voice consistency — and ultimately the end users of the conversational products themselves (customers, members, patients, citizens). When conversational design quality is encoded into AI-generated work, the cost of poor experiences declines: 60% of consumers report negative chatbot experiences (Zendesk); 40% prefer no chatbot at all over a bad one; brand consistency drives 10–33% revenue lift (Lucidpress). The dual-stakeholder pattern (buyer ≠ user) is what makes this B2B rather than B2B2C — the buyer is accountable for outcomes the end users produce; both populations are inside the same organization. The end users of the conversational products being built (customers, members, patients, citizens) are downstream beneficiaries but never paying customers of this practice. | |||||
| Differentiators | What are the key differentiators for your company? | Four differentiators distinguish the Conversational Design Creator from adjacent alternatives: 1. Encodes conversational design thinking, not just flows. Most conversational design tools (Voiceflow, Botpress, Chatbase, Rasa, Cognigy, Kore.ai) encode the what — flows, intents, scripts, integrations. The Creator generates from the why — the reasoning behind why a button decision must handle unexpected input, why fallback cascades the way it does, when escalation must trigger, how multi-turn context should be maintained, what brand voice rules apply per channel. RAG-grounded runtime chatbots (Chatbase, StackAI, Voiceflow with knowledge base) train on company documents — product docs, FAQs, policy PDFs. The Creator's corpus encodes a different layer: the team's conversational reasoning, which for most teams lives in heads, review comments, and scattered Slack threads, not in document repositories. The category — encoded conversational design thinking as machine-readable corpus driving an AI generator — does not exist as a named market today. 2. Design layer, not runtime layer. Enterprise chatbot platforms (Voiceflow, Botpress, Cognigy, Kore.ai, Rasa) and AI agent products (Sierra, Decagon, Replicant) generate runtime responses — what the end user sees when interacting with the chatbot. The Creator generates design artifacts — flows, conversation structures, decision points, policies — that engineers and builders implement in those platforms. This is one to two layers upstream. The Creator produces design artifacts as structured specifications, not as positioned implementation in any specific tool — keeping the architecture tool-agnostic across Figma, Voiceflow, Bird, and other downstream platforms. Design artifacts are also human-readable by non-developer stakeholders — making conversational reasoning visible to PMs, brand teams, customer service leads, translation teams, and anyone accountable for the conversational experience. When conversational design lives only in platform code or config (the default for engineer-led shops), it is invisible to everyone who isn't a developer; design artifacts make decisions inspectable and improvable by the people accountable for outcomes. The Creator's output feeds INTO Voiceflow / Botpress / Sierra / Decagon, not against them. Closest existing precedent for encoded conversational reasoning is Decagon's Agent Operating Procedures (AOPs) — "natural language instructions that compile into structured logic" — but Decagon applies this to runtime agent behavior, not to design artifacts. Sierra's context engineering philosophy ("when a strong model makes a poor decision, the issue is usually insufficient context") validates the thesis but operates at the runtime layer. 3. Single integrated architecture: generation and validation from one corpus. The Creator (Phase 1) and the built-in QA gate (Phase 2) share the same canonical corpusThe Creator generates a draft; the QA gate validates that draft against the same corpus before handoff. One source of truth, two AI behaviors within one integrated product. This is structurally different from off-the-shelf AI generators (which don't follow team conventions) and from standalone validators (which can only catch violations after the fact). The integrated architecture works across all three conversational sub-segments — deterministic flows, AI agents, hybrid flows — because the corpus structure is consistent even as the downstream artifact format differs per client. A future Conventions Assistant for self-service corpus maintenance extends the same one-corpus-multiple-AI-readers architecture without adding a separate product. 4. The moat is architectural, not just commercial. Three defensibility layers compound within the 12–24 month window before adjacent categories can catch up. Architectural distance. The corpus runs in two modes from one source of truth — settled rules applied silently with citations, judgment-call rules surfaced for producer decision. Brand-voice governance adjacents (Acrolinx, Averi, Markup AI) have only mode one. Adding mode two is not a feature addition — it requires redesigning their corpus structure from single-mode rule validation to two-mode rule + decision-framework. That structural change is the timeline floor. Methodology IP. Each company's encoded conventions are unique — capturing specific reasoning patterns, brand voice, customer context, and edge cases. The AI-accelerated extraction methodology — pattern abstraction across designer articulation, existing artifacts, review comments, and CS transcripts — is the hard-to-replicate practice. The product is deliberately delivered as high-ticket consulting rather than subscription, because the extraction is the work. Reference momentum. First customers anchor commercial position before adjacents catch up. The methodology produces usable corpora across the maturity spectrum — from dedicated design teams to companies shipping AI chatbots through engineers and PMs — making the practice commercially viable from engagement one. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | N/A - New Product from 0 to 1 | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | N/A - New Product from 0 to 1 | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | N/A - New Product from 0 to 1 | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | External users (B2B). Three layers: PRIMARY END USERS — the people producing conversational work in companies investing in conversational AI experiences. Design-team-agnostic: conversation designers, UX writers, junior and mid-level designers, generalists covering conversational projects, engineers building conversational features, PMs and product owners shipping conversational AI, contractors and agency staff without team context, and AI-tool operators whose conversational outputs need grounding in company-specific reasoning. The persona deliberately decenters the conversation design steward (when one exists): companies with a strong lead have a workaround in that lead's tacit knowledge; companies without one are in real pain and represent the larger commercial market. BUYER — head of design, VP of product design, design ops lead, VP/director of conversational AI/CX, CTO, or head of product. Accountable for conversational AI adoption outcomes (CSAT, ROI, escalation rates, brand voice consistency) and the downstream quality of conversational work shipping into the broader product (customer service, support, sales, member engagement). Buyer profile varies by company maturity: design leadership in design-mature companies, product or CTO leadership in engineer-led shops. SECONDARY BENEFICIARIES — the entire product organization (PMs, engineers, content and translation teams, brand/marketing teams concerned with voice consistency) and the end users of the conversational products being built (customers, members, patients, citizens). When conversational design lives in code or platform config, non-developer stakeholders cannot inspect, audit, or improve it — the Creator's design artifacts solve this legibility problem. When conversational conventions are encoded and applied to AI generation, the cost of poor experiences declines: 60% of consumers report negative chatbot experiences (Zendesk); 40% prefer no chatbot at all over a bad one; only 30% of companies enforce brand guidelines (Contentstack 2026); brand consistency drives 10–33% revenue lift (Lucidpress). The conversational products themselves — and the customers who interact with them — are the largest population reached. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | The topic label "Journey Map (current-state)" and the question text ("...when they are using your product / service...") point in different directions. I've interpreted this cell as the current-state journey, since the Design phase has a separate Workflow (future) cell for the target-state journey with the product. Please flag if a different interpretation is expected and I'll revise. Current-state happy path for people producing conversational work in companies without structured conversational design convention enforcement. Producer-agnostic — the actor at each step may be a dedicated conversation designer, generalist, engineer, PM, contractor, or AI-tool operator. 1. BRIEF INTAKE. Producer receives a conversational design task — a chatbot flow, conversational feature, voice agent path, WhatsApp/SMS journey, or new language for an existing flow. 2. REFERENCE HUNTING. Producer searches existing artifacts — Figma files, Notion pages, Slack threads, customer service transcripts, conversation design playbooks (if any exist), existing flows in Voiceflow/Botpress/Bird — for similar patterns. Happy path: finds good references. More common: inconsistent or incomplete references requiring interpretation. Producers without team context (contractors, new generalists) often find nothing relevant. Producers in engineer-led shops often have no design artifacts to reference — only existing platform configurations or code. 3. CONVENTION INFERENCE. Producer infers company conventions from artifacts (entry-point naming, fallback handling, escalation triggers, tone-of-voice patterns, edge-case behavior, button-decision conventions). Undocumented conventions are guessed at or asked about in Slack — typically to a conversation designer or design lead when one exists. In companies without a design lead, producer guesses, makes stylistic choices that may or may not match prior work, or copies patterns from generic AI conversational design best practices that don't reflect the company's actual reasoning. 4. AUTHORING. Producer builds the conversational design in Figma, Voiceflow, Botpress, Bird, or equivalent — or, in engineer-led shops, directly in platform configuration or code (skipping the design artifact entirely). AI conversational generators (Claude, ChatGPT, Gemini, voice agent platforms) may produce initial drafts the producer then refines. AI outputs default to generic conversational design patterns, not company-specific reasoning — producer either rewrites significantly, accepts off-brand outputs, or ships outputs that look right but aren't right. 5. REVIEW REQUEST. Producer pings a senior conversation designer, design lead, or brand owner for review (Slack, Figma comments, sprint review). Happy path: reviewer available. Common: busy or unavailable. In companies without a designated reviewer, this step is skipped and producer self-checks. In engineer-led shops where conversational design lives in code, reviewers (if any) can't easily inspect the work — the legibility problem makes review impractical. 6. CONVENTION ENFORCEMENT (the heavy lift). Reviewer manually checks the design against their internalized understanding of conventions. Identifies violations, leaves comments with corrective examples, engages in back-and-forth on novel cases (edge-case handling, fallbacks, escalation, unexpected inputs). When no reviewer exists, violations are caught later — by customer complaints, brand audits, or engineering pushback during implementation. When work lives in code/config, violations may not be caught at all until customers experience them. 7. PRODUCER CORRECTION LOOP. Producer corrects the design, re-pings the reviewer (if any), who re-reviews. Loop continues until compliant — or, in unreviewed contexts, until producer is out of time and ships. 8. HANDOFF AND ONGOING MAINTENANCE. Design moves to handoff (engineering implementation, translation/localization for multilingual flows, content review). Producers and engineers surface convention questions during implementation. New conventions emerge from edge cases discovered in production, and may or may not get documented anywhere. When work lives only in code/platform config, downstream teams (translation, brand, customer service) can't trace decisions back to design intent. Even in the happy path, this journey is friction-heavy: dependency on expert reviewer availability, time spent on reference hunting and convention inference, manual enforcement, review cycles. In realistic cases — incomplete references, busy/absent reviewer, AI tools generating non-compliant work, conventions never documented, design that lives in code — the journey breaks down significantly. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | Pain points cluster into three layers: LAYER 1 — STRUCTURAL PROBLEMS (predates AI; foundation-layer) L1.1 - Conversational conventions live in heads, not documents. The reasoning behind why a button decision must handle unexpected input, why fallback cascades the way it does, when to escalate, how to maintain multi-turn context, what brand voice rules apply — most of this lives in expert reviewers' heads, individual experience, and tacit knowledge. When documented, the documentation is narrative, scattered (Notion, Slack threads, Figma comments), often out of date, and not enforceable by any system. L1.2 - Review is the only enforcement mechanism. Convention compliance happens because a reviewer catches violations during review (Step 6 of the journey). There is no system, no validator, no automation. The reviewer is the gate — or there is no gate at all. L1.3 - Expert reviewer time is the bottleneck (when one exists). Time spent on convention enforcement is time not spent on higher-value work. Senior IC time is the most expensive and constrained resource on design-mature teams. L1.4 - Convention drift on absence. If the expert is on PTO, on leave, busy, or has left the team, convention enforcement collapses. In companies without a senior conversation designer, this is the default state, not the exception. L1.5 - Knowledge dispersion. Conventions accumulate across years of sprint corrections, Slack threads, customer service transcripts, and review comments. No single person has the full picture. Critical context is lost when channels are archived or team members leave. L1.6 - Legibility: conversational design that lives in code or platform config is invisible to non-developer stakeholders. In engineer-led shops where conversational AI is built directly in Bird, Voiceflow, Botpress, or custom code without an intermediate design artifact, the conversational logic lives inside the implementation. PMs cannot see what decisions were made. Brand teams cannot audit voice consistency. Customer service leads cannot trace why a conversation fails. Translation teams need engineers to walk them through every flow. New team members cannot learn the system without reading code. Conversational design becomes invisible to every stakeholder who is not a developer — including the people accountable for the business outcomes (CSAT, brand voice, escalation rates). For these companies, the absence of a design artifact is the foundational pain, not a secondary one. L1.7 - Cross-functional handoffs amplify the problem. Engineers implementing conversational features, content and translation teams localizing flows, brand teams reviewing for voice consistency — all depend on conventions being legible. Without an encoded corpus, they all depend on the expert reviewer. LAYER 2 — AI-AMPLIFIED PAIN (acute, recent, fastest-growing) L2.1 - AI conversational generators produce off-brand, non-compliant work. Producers using Voiceflow AI, Botpress LLM, ChatGPT, Claude, Gemini, Sierra/Decagon agent platforms generate drafts that default to generic conversational design best practices, not the company's specific reasoning. "AI gives you 'professional B2B tone' because it has no idea what YOUR tone is" (Averi). Outputs look right but aren't right. The reviewer now has to review AI-generated work in addition to human work, multiplying review load — or, in companies without a reviewer, the off-brand work ships. L2.2 - RAG-grounded chatbots are NOT a solution. 71% of organizations use generative AI in at least one business function (McKinsey 2025); most enterprise chatbot platforms now offer "train on your knowledge base." But this grounds the chatbot in product docs and FAQs, not in conversational design reasoning. The chatbot knows the company's product; it doesn't know how the company talks to customers. L2.3 - Hybrid flows compound the problem. In hybrid flows (deterministic logic + AI nodes), the AI node generates conversational outputs inline that violate the surrounding deterministic flow's conventions. The hybrid pattern is increasingly dominant; the pain is growing fastest. L2.4 - AI tools cannot be trained on conventions that don't exist in machine-readable form. Even when companies want to make their conversational AI tools "company-aware," they have nothing structured to feed in. The advertised "train on your knowledge base" assumes a structured corpus that for most companies doesn't exist. L2.5 - AI generators amplify the legibility problem. When engineering teams use AI to ship conversational features directly into platform config or code (skipping any design artifact), the conversational logic is now both AI-generated AND invisible to non-developers. The combination is acute: nobody outside engineering can see what the AI generated, nobody can audit whether it reflects company reasoning, and nobody can improve it without going through a developer. L2.6 - Buyers are spending on conversational AI that fails to deliver ROI. Companies adopt AI conversational tools expecting velocity and quality gains; the lack of foundation produces poor outputs, increased review work, customer complaints, and escalation rates. LAYER 3 — DOWNSTREAM BUSINESS COST (the leadership-visible narrative) L3.1 - Poor customer experience at scale. 60% of consumers report negative chatbot experiences (Zendesk); 40% prefer no chatbot at all over a bad one; only 28% report satisfaction (Gartner). CSAT gap AI-handled vs human exceeds 15–20 points (Chatbase 2026). Translates directly to brand erosion and customer churn. L3.2 - Brand voice inconsistency. Only 30% of companies actively enforce brand guidelines (Contentstack 2026); brand consistency drives 10–33% revenue lift (Lucidpress). "If your chatbot sounds like a different company from your marketing copy, you have a brand coherence problem no amount of good product design can fix" (Atom Writer). L3.3 - Customer service cost overruns and escalation rates. When conversational AI fails to resolve issues, escalation rises, customer service teams handle higher volumes, and the cost savings AI was supposed to deliver evaporate. CNBC reports AI customer service failure rate is nearly 4× higher than AI use in general. L3.4 - Slow development cycles. The review-correction loop adds days to every conversational feature. Engineering waits on design clarification. Translation waits on convention sign-off. Velocity drops. L3.5 - Failed AI ROI undermines broader AI strategy. Leadership trust in AI investment erodes — affecting future AI budget allocations. MOST FREQUENT AND SEVERE PAIN POINTS, RANKED: L2.1 - AI-generated conversational work that doesn't reflect company-specific reasoning (Step 4 onward). Daily occurrence, growing fastest, especially in hybrid flows and in companies shipping AI without dedicated design coverage. The single largest growing tax on conversational AI ROI. L1.3 - Expert reviewer bottleneck (Step 6, when a reviewer exists) OR convention drift in non-design-mature companies (no reviewer at all). Either way, the convention enforcement layer is broken — by overload or by absence. L1.6 - Legibility — conversational design that lives in code or platform config is invisible to non-developer stakeholders. Foundational in engineer-led shops; compounded when AI generates the work directly into the platform. Audits, incidents, or cross-team handoffs force this pain into leadership visibility. L3.1 - Downstream customer experience and brand cost. Persistent, leadership-visible, hardest to ignore — the pain that gets the buyer's attention reliably. L1.5 - Knowledge dispersion across scattered sources. Persistent, gets worse over time, blocks AI tool integration attempts. L3.4 - Slow velocity in shipping conversational experiences. + L1.7 - Cross-functional handoffs amplify the problem. Cumulative cost across every feature, every language, every channel. | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | The pain analysis identified five clusters of pain points. AI is well-suited to address four of them, in different ways. Ranked from most severe and frequent to least: Opportunity 1: AI-generated conversational design grounded in a company-specific corpus. Addresses: Cluster A (L2.1, L2.3, L2.6) — AI generator misalignment. AI's specific value: LLMs can generate structured conversational design (flows, decision points, brand-voice-compliant content) when given a structured corpus as grounding context, with each generated element citing the corpus rule(s) applied. Works across all three sub-segments (deterministic, AI agents, hybrid). Highest priority: the only AI opportunity that hits all three stakeholders directly (company ROI, team velocity, user experience) and addresses the fastest-growing pain (L2.3 hybrid flows). The wedge into the design-team-agnostic market. Opportunity 2: AI validation of conversational design artifacts against company conventions. Addresses: Cluster B (L1.2, L1.3, L1.4) — Convention enforcement breakdown. AI's specific value: LLMs can read a corpus + a conversational design artifact, apply each corpus rule, and produce structured findings (correctly applied / potential issue / violation) with explanations and suggested corrections — replacing the reviewer-as-gate pattern. Input-agnostic: validates human-authored work, drafts from the conversational design generator, or drafts from external AI generators. Second-highest priority: solves the universal pain (review bottleneck in design-mature; absence-of-review in engineer-led). Throughput multiplier. Critical companion to Opportunity 1 — without validation, generation alone can't guarantee compliance. Opportunity 3: AI-accelerated extraction of conversational design thinking into a structured corpus. Addresses: Cluster C (L1.1, L1.5, L2.4) — Knowledge structure gap. AI's specific value: LLMs can read scattered sources (existing flows, past review comments, customer service transcripts, AI generator outputs) and abstract patterns across them, proposing draft conventions for steward/team review. Source-agnostic — works with whatever inputs exist in the company's reality. Foundational prerequisite — Opportunities 1 and 2 don't work without a corpus. But this AI is used by the consultant during the engagement's extraction phase, not by the end user in the deployed product. AI-accelerated extraction tooling is part of the broader methodology and is not built as a separate capstone prototype. Opportunity 4: AI runtime drift detection — comparing production conversational AI behavior against the corpus. Addresses: Cluster D (L1.6, L1.7, L2.5) — Legibility in engineer-led shops. AI's specific value: LLMs can compare a runtime system's actual outputs (sampled from production traffic) against the corpus, flagging drift from the company's specific reasoning. Surfaces conversational design that's invisible inside platform code or config. Highly valuable for engineer-led shops where design lives in code; lower frequency overall (concentrated in one customer profile). Closes the post-deployment legibility gap that a design artifact alone doesn't address. Opportunity 5: AI-powered corpus health and ROI analytics for buyers. Addresses: Cluster E (L3.1–L3.5) — Downstream business cost / failed AI ROI signaling. AI's specific value: LLMs can analyze Apply/Override/Flag patterns from product usage, surface corpus coverage gaps, and generate executive-level summaries of conversational AI investment performance — making ROI visible to the buyer. Indirect value — Cluster E outcomes auto-improve when Opportunities 1 and 2 are deployed. Strengthens buyer renewal and budget defensibility. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | HMW 1 — Opportunity 1: Conversational design generation grounded in a company-specific corpus. How might we make conversational design AI-generation actually reflect a company's specific reasoning? H1.1 - Conversational Design Creator (web app). Producer enters a brief → AI generates a complete conversational flow grounded in corpus → outputs structured JSON exportable to Figma, Voiceflow, Bird, etc. Each generated element cites the corpus rule(s) applied. HiTL: producer reviews, refines, regenerates. H1.2 - Brief-to-flow Figma plugin. Same generation capability runs inline in Figma. Producer writes the brief in a Figma node; plugin generates the flow as Figma nodes the producer can edit directly. No tool-switching. H1.3 - AI conversational style transfer. Take any conversational draft (off-brand AI output, generic template, draft from another tool) → AI rewrites it in the company's voice using the corpus, with a diff view showing what changed and why. Grammarly for conversational brand voice. H1.4 - Corpus-aware AI tool wrapper. Browser extension or platform plugin that wraps ChatGPT, Claude, Voiceflow AI, or Botpress LLM. When the producer prompts the underlying AI, the wrapper auto-injects relevant corpus context. The AI generator becomes corpus-aware without the producer remembering to include conventions. H1.5 - Conversational template library generator. Producer specifies a use case category (fundraising opener, onboarding step, opt-out path, donation receipt) → AI generates a library of corpus-grounded templates the team uses as starting points. Library auto-updates as the corpus evolves. HMW 2 — Opportunity 2: Validation of conversational design against company conventions. How might we automate convention enforcement so quality doesn't depend on a senior IC's availability? H2.1 - Built-in QA gate (integrated with the Creator). AI validates the Creator's output against the corpus, surfacing findings with explanations and suggested corrections. HiTL: producer Applies / Overrides (with reason) / Flags for corpus update per finding. H2.2 - Standalone Conversational Design Auditor. Validates any draft — human-authored, Creator-generated, or output from an external AI generator — against the corpus. Used as a pre-handoff check, accessible via web app or API. Input-agnostic. H2.3 - Asynchronous QA gate. Runs continuously across all conversational design artifacts in the team's workspace (Figma files, Voiceflow projects, Bird configs). Flags violations and routes findings to producers via Slack/email — no manual triggering. Surfaces drift between formal reviews. H2.4 - CI/CD linter for conversational design. Integration in the engineering deployment pipeline. Before any Bird/Voiceflow/Botpress config goes to production, AI validates against the corpus. Blocks deployment on critical violations; warns on minor ones. Engineering-side gate. H2.5 - Handoff validation gate. Mandatory check before conversational design moves from design to engineering. AI produces a compliance report attached to the handoff package. Engineering won't accept work that hasn't passed. Makes the handoff itself the enforcement moment. HMW 3 — Opportunity 3: Extraction of conversational design thinking into a structured corpus. How might we extract tacit conversational reasoning from scattered sources into a structured corpus faster than manual analysis? H3.1 - AI-led extraction interviewer. AI conducts structured interviews with conversation designers/team leads, capturing tacit conventions through directed conversation. Transcripts converted to draft corpus rules for steward review. H3.2 - Auto-documentation extractor. AI reads scattered sources (Slack threads, Figma comments, past review feedback, Notion docs, customer service transcripts) → abstracts patterns → drafts initial corpus rules for refinement. H3.3 - Customer service transcript analyzer. AI surfaces conversational patterns from real user interactions (failure modes, escalation triggers, unexpected inputs) → proposes corpus rules grounded in production reality, not just tacit team knowledge. H3.4 - AI generator output reverse-engineer. AI reads existing AI-generated conversational content the company has shipped, infers what conventions were applied (or violated), proposes rules to formalize the implicit ones and correct the inconsistent ones. H3.5 - Cross-designer pattern detector. AI compares conversational flows across multiple designers' work in the team's repository, identifies consistent patterns vs inconsistencies, surfaces both: consistencies → candidate corpus rules; inconsistencies → resolution needed at corpus-definition time. HMW 4 — Opportunity 4: Runtime drift detection in production conversational AI. How might we detect when deployed conversational AI drifts from the company's encoded reasoning in production? H4.1 - Inline runtime drift monitor. Real-time watcher on production conversational AI traffic. Each AI response compared against the corpus; violations flagged live to a dashboard. Like Datadog APM for conversational compliance. H4.2 - Sampled production audit. Periodically samples production conversation transcripts → AI validates each against the corpus → produces a compliance report (weekly/monthly). Less expensive than live monitoring; still catches drift over time. H4.3 - Hybrid flow diff view. For hybrid flows (deterministic + AI nodes), AI compares each AI node's actual outputs against the surrounding deterministic flow's conventions, highlighting drift. Surfaces problem nodes to engineering for review. H4.4 - Brand voice health score. Continuous metric tracking the % of production conversational AI outputs aligned with the corpus's brand voice rules. Aggregated across channels, languages, segments. Buyer-facing dashboard. H4.5 - Customer-complaint cross-reference. AI cross-references customer complaints/escalations with the specific conversational AI outputs that triggered them, then validates those outputs against the corpus. Connects observable user pain to specific convention violations. HMW 5 — Opportunity 5: Corpus health and ROI analytics for buyers. How might we make conversational AI investment ROI visible and defensible to the buyer? H5.1 - Quarterly corpus health report. AI analyzes Apply/Override/Flag patterns from product usage and generates an executive summary: corpus coverage gaps, convention drift trends, AI generator alignment rate, recommendations. Buyer-facing document. H5.2 - AI investment ROI dashboard. Tracks conversational AI ROI against three baseline metrics — CSAT delta vs pre-AI baseline, escalation rate vs human baseline, deflection rate vs target. AI generates trend explanations and intervention recommendations. H5.3 - Corpus coverage heatmap. Visualization showing which corpus rules are well-enforced vs poorly-enforced across the team's work. AI identifies systemic gaps for training, corpus updates, or workflow changes — not just individual violations. H5.4 - Conversational AI vendor scorecard. AI assesses how well the buyer's chosen conversational AI vendor (Voiceflow, Sierra, Decagon, Cognigy) is performing against the company's corpus. Helps buyers decide on vendor switching or doubling down. H5.5 - Predictive ROI forecaster. AI uses corpus maturity, team adoption metrics, and production drift data to forecast next-quarter conversational AI ROI. Helps the buyer defend budget allocations in advance, not retrospectively. | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | Ranked the 25 Diverge ideas by Impact (severity and frequency of pain addressed, per the pain analysis) and Feasibility (prerequisite dependencies, development complexity, prep work). TOP THREE BY COMBINED SCORE: 1. AI-led extraction interviewer (3.1) — 9/10 2. Conversational Design Creator (1.1) — 8/10 3. Standalone Conversational Design Auditor (2.2) — 8/10 SELECTED CAPSTONE BUILD: Conversational Design Creator. The Creator generates conversational design drafts from a producer brief, grounded in a hand-curated corpus extracted from the case study (a real-world multilingual conversational design system on WhatsApp at a global civic advocacy non-profit). Each generated element cites the corpus rule(s) applied. Producer reviews with HiTL actions: Approve / Regenerate. REASONING: - Highest impact among capstone-eligible options — addresses Cluster A (AI generator misalignment), the main pain cluster from the analysis. - Distinctive demo — two-mode corpus-grounded generation (silent convention rules + surfaced judgment-call alternatives) is a fresh AI category, distinct from off-the-shelf generators and from validators (a more crowded space). - IP-protected — methodology stays proprietary as engagement value; product showcases what the methodology produces. This is why the top-ranked extraction interviewer (3.1) was ruled out as the capstone showcase. - Anchors the broader product vision — Creator establishes the two-mode corpus-grounded AI architecture that QA gate, Auditor, runtime drift detection, and buyer analytics all extend in Future Phases. CORPUS SCOPE: 15–20 rules across 3–4 categories (brand voice, message structure, platform conventions, compliance). Single conversational domain to keep scope tractable. Time-boxed at 6 hours of prep work. DEFERRED TO FUTURE PHASES (documented in PRD, not built): - Built-in QA gate (Phase 2 — completes the validation half of the original Creator+QA vision) - Standalone Auditor (alternative deployment mode for validation) - AI-led extraction tooling (methodology-side, proprietary) - Runtime drift detection (closes Cluster D legibility gap) - Buyer-facing analytics (corpus health, ROI) | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | VISUAL HERE TWO-MODE CORPUS POSTURE — central architecture. The corpus runs in two modes from a single source of truth: Convention rules fire silently with citations — language gates, payload patterns, button limits, voice principles. AI applies them without asking. Producer sees the audit trail as citation chips. Judgment-call rules surface alternatives with reasoning at the moment the producer needs to decide — path selection, audience-context, decision frameworks. AI asks; producer picks; dependent convention rules cascade silently after the decision resolves. Every workflow step below operates within this architecture. WORKFLOW STEPS 1. PRODUCER PROVIDES BRIEF [manual]. Producer types, pastes URL, or uploads document. Brief includes use case, channel, audience, constraints, scope. [Corpus defines: brief-parsing rules — glossary, scope-reference patterns, required fields.] 2. CREATOR READS CORPUS + EXISTING STATE + PARSES BRIEF [AI, background]. Loads corpus, reads existing state filtered by brief scope, parses brief. Existing-state summary surfaces visibly so grounding context is transparent, not hidden. [Corpus defines: source of existing state, scope filters, what's in/out of scope.] 3. CREATOR PRESENTS INTERPRETATION + CONFIRMATION GATE [AI + user choice]. Interpretation summary: what was understood, how it relates to existing state (new / modifies / replaces / preserved), and any remaining questions. Multi-round: producer signs off OR refines (corrects scope, adds context). AI re-interprets. Loop until aligned. No generation without sign-off. 4. CREATOR GENERATES DRAFT [AI, foreground]. Structured by corpus-defined design unit. Each unit tagged as new / modifies / replaces / preserved. Each unit carries inline citations from the convention rules silently applied. Path-discovery brainstorm fires automatically at any judgment-call rule the corpus surfaces — AI presents corpus-grounded alternatives with reasoning; producer picks; dependent convention rules cascade silently (visible as citations growing on the unit). [Corpus defines: the design unit, structural conventions, voice/tone rules, naming, platform/channel constraints.] 5. PRODUCER REVIEWS DRAFT BY DESIGN UNIT [manual]. Master-detail review. Citations visible per unit. Existing-state context (new / modifies / replaces / preserved) shown. 6. PRODUCER ACTS PER DESIGN UNIT [user choice — HiTL]: - APPROVE — move on. - REGENERATE — multi-turn chat with AI. Two distinct behaviors: Convention Rule Push-back when producer overrides a settled rule (AI checks corpus envelope, applies within range OR surfaces alternatives at hard limits — e.g., button-cap §6.1); Judgment-call refinement when producer wants different options at a decision moment. No silent rejections. - Manual edits happen in the downstream tool after export. 7. PRODUCER EXPORTS DELIVERABLE [manual]. Portable specification — flow diagram + unit content + citations as audit trail. [Corpus defines: output format, handoff items, optional direct-write delivery per engagement.] END STATE The corpus-grounded, existing-state-aware specification is exported with citations showing which conventions were applied at each decision point. Both corpus modes are auditable: silent rules visible as citation chips, judgment-call resolutions visible as cascade growth on the units they affected. BRANCHES AI's interpretation wrong at the gate → Producer corrects → AI re-interprets → Loop. Producer wants to refine brief mid-flow → Brief revisitable; corpus + existing-state reads re-run. Generated unit doesn't match intent → Producer regenerates with feedback (Convention Push-back or Judgment-call refinement). Producer notices a corpus gap → Flags for methodology-side corpus review. Producer abandons partway → Draft state saved; resumable. WHERE AI ADDS VALUE Summarize — Step 2 (existing state) + Step 3 (brief interpretation). Create — Step 4 (corpus-grounded draft generation). Explain — Step 4 onward (citations make AI reasoning auditable). Suggest — Step 3 (clarifying questions) + Step 6 (Convention Push-back alternatives + Judgment-call options). WHAT REMAINS MANUAL Brief authoring (Step 1) — producer intent originates with the producer. Confirmation at the gate (Step 3) — no generation without sign-off. Unit-by-unit review and approval (Step 5–6) — human judgment irreplaceable. Export action (Step 7) — producer owns the deliverable handoff. WHAT'S A USER CHOICE Brief content, scope, input method (Step 1). Confirmation gate: sign off OR refine (Step 3). Per-unit actions: Approve / Regenerate (Step 6). Whether to flag corpus gaps (branch). When to export (Step 7). | |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | VISUAL HERE Four main screens covering the producer's path from brief to handoff, plus four state variations that fire at specific decision moments. Units reviewed as text — not as rendered WhatsApp screens — matching how conversational designers actually work. NAVIGATION Brief Input → [AI processes — State A] → Confirmation Gate → [AI generates] → Draft Review ↔ Unit detail panel → [diagram renders] → Export A persistent 4-step stepper (BRIEF · CONFIRM · REVIEW · EXPORT) anchors the user. AI processing between steps is a visible state with progress steps and ETA. STATE VARIATIONS - State A — AI Processing overlay between Brief and Gate. Steps with checkmarks + progress bar. - State B — Refining (multi-turn). Fires from Refine at the Gate. AI walks producer through the interpretation; producer corrects scope or asks questions. - State C — Convention Rule Push-back (LIVE API badge). Fires when producer overrides a settled rule (e.g., 4 buttons when corpus caps at 3 per §6.1). AI checks corpus envelope and surfaces 3 corpus-grounded alternatives with reasoning. Runs against the real model. - State D — Path-Discovery Brainstorm (multi-turn). Fires at a judgment-call rule during generation. AI surfaces corpus-grounded alternatives; producer picks; dependent convention rules cascade on the affected unit after resolution (visible as citation chips growing in count). KEY DECISION POINTS - Brief: the brief itself. Scope and reference state move to the gate. - Gate: Generate (primary, green), or Refine to talk it through with AI (opens State B). - Per unit: Approve (green) or Regenerate (purple — opens State C or D depending on context). Manual edits in downstream tool after export. - Export: Export diagram (primary), or Refine to step back to Review. KEY UI ELEMENTS - Top bar with breadcrumb + 4-step stepper. - Brief: single textarea, attach/URL options, primary action right-aligned. - Gate: AI-detected scope + reference state cards (coral), coral interpretation card with proposed-changes pills, Generate / Refine side-by-side. - Review: master-detail split. Left = compact unit list with status dots and Approved/Pending pills. Right = detail panel with text content (coral), citation chips, HiTL actions. Two side-by-side panels in the wireframe demonstrate progressive cascade state — first panel pre-brainstorm (8 chips on unit 2.2, 2 pending), second panel post-resolution (10 chips with 2 new, 1 pending). - KPI cards above the review split: Total, Pending, Approved. - Chat bubbles in States B/C/D: coral for AI, dark for the producer, with small role labels. - Chat input row: free-text input + dark send button + green ✓ accept button for explicit "commit AI's last proposal" action. - Quick-nudge tags above the chat input ("More paths", "Different scope", etc.) for fast common feedback. - Export: vector flowchart preview (Entry → Language decision → EN/non-EN branches → Welcome message + Navigation array + button pills → Existing paths) with AI-generated nodes in coral and preserved nodes in gray. Export diagram / Refine buttons side-by-side. - Buttons by intent: primary dark, generate green, approve green, regenerate purple, live-API badge red. HOW THE LAYOUT ACCOMMODATES AI FEATURES - Coral marks AI content throughout — human/AI boundary visible at a glance. - Units reviewed as text — message content, button labels, structural choices — not rendered screens. - AI Processing visible with progress and ETA — not a hidden spinner. - The Gate consolidates everything AI derived so confirmation is one deliberate act. - Citations live inline with each unit, never hidden — every output traceable to corpus rules. - Two-mode posture visible at the UI layer: convention rule citations accumulate as chips silently; judgment-call brainstorms (State D) fire as multi-turn dialogues that resolve into citation growth on the affected unit. The cascade is the visual proof that the architecture works. - Convention Rule Push-back (State C, LIVE API badge) is structurally distinct from Path-Discovery Brainstorm — same chat UI, but corpus-envelope check + alternatives at hard limits, not surfaced options from a judgment rule. - HiTL controls are three explicit intents across surfaces: Approve, Regenerate, and ✓ accept-AI-proposal (in chat states). - Vector diagram on Export is the deliverable artifact, not just a screen — portable spec with decisions, content, and convention rules per node. | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | PROTOTYPE HERE The prototype is a single-file HTML build (~2,100 lines, deployable to GitHub Pages) demonstrating the Creator's critical AI moments. Typewriter animation on chat; vector flowchart on Export. Built against the wireframes. Critical AI moment: the path-discovery brainstorm cascade. AI surfaces alternatives at a judgment-call rule, producer picks, dependent convention rules cascade silently. Visible in 30 seconds: producer reviews unit 2.2 Navigation array with 8 citation chips → opens brainstorm via in-body link → AI surfaces 4 corpus-grounded paths → producer picks Newsletter signup → unit 2.2 approved, citation count grows 8 → 10 with two new chips animating in (Newsletter chain §7.6, PAD chain §7.2). Silent convention rules + surfaced judgment-call rules + dependent cascade in one continuous flow. LIVE API moment. State C — Convention Rule Push-back — runs a real Anthropic API call. Producer overrides a convention rule (asks for 4 buttons when corpus caps at 3 per §6.1); AI checks the corpus envelope and surfaces 3 corpus-grounded alternatives with reasoning. Captured live once with the real corpus loaded, output pasted verbatim into the prototype with a visible LIVE API badge. Honest framing: real model output, captured live, used in the recording. Screens in scope (built & clickable): Brief Input → AI Processing overlay → Confirmation Gate → Draft Review (master-detail, 4 units, Session #042 Welcome step scope) → Export Loading overlay → Export with SVG diagram. State variations: A (AI Processing), B (Refining), C (Convention Rule Push-back · LIVE API), D (Path-Discovery Brainstorm). Out of scope (future phase): dashboard, history, auth, settings, corpus management UI, manual Editing state. Input → AI → Output → HiTL pattern. Input: brief textarea + chat in states B/C/D. AI processing: overlays with progress steps + ETA; "✓ AI complete" badges. Output: coral content surfaces (interpretation, unit body, citation chips). HiTL: Approve / Regenerate per unit, Generate / Refine at the gate, ✓ accept button across chat states for explicit commit-to-AI-proposal. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | # Conversational Design Creator — Master Prompt # Who you are You are the Conversational Design Creator — an AI design partner working inside one team's encoded design system. Your job is to take a producer's brief and turn it into a corpus-grounded draft made of design units — messages, decision points, structural choices, journey steps, whatever the corpus defines as the unit of design. You are not a general-purpose conversational designer. You operate inside the corpus this team has explicitly encoded. You generate from it, you cite it, and you stay inside its boundaries. # What you produce, and for whom Your output is the design specification, not the finished implementation. Producers take your accepted units to their build tool — Figma, Voiceflow, Bird, a spec document — and implement from there. Write unit content that reads cleanly to both a human reviewer and a downstream consumer. Do not produce platform-specific markup, frame coordinates, or visual layout — those belong to the implementation tool. # How the corpus operates: two modes The corpus runs in two modes from a single source of truth. Every corpus entry is tagged class: 1 (convention rule) or class: 2 (judgment-call rule). Your job is to honor the class. ## Class 1 — convention rules. Settled rules with a single valid answer. Apply silently and cite (language gates, payload patterns, button limits, voice principles, structural conventions). The producer sees the audit trail in grounded. ## Class 2 — judgment-call rules. Rules where the corpus surfaces multiple valid options for producer decision (path selection, audience-context, decision frameworks). Don't pick silently. Ask using the alignment-question pattern below. After the producer resolves the judgment call, any dependent Class 1 rules cascade — apply them silently and add their citations to grounded. # How you sound Warm, direct, decisive — like a senior conversational designer on the producer's team. Brief, never chatty. You use "I" sparingly and never apologize unnecessarily. You think out loud when context is fuzzy. You ask sharp targeted questions when the corpus gives you clear options. Your signature move is the corpus-grounded alignment question. When the producer's input is fuzzy or you hit a Class 2 judgment-call rule, name the situation and surface the corpus options. ## Two patterns: ### When the producer raises a structural choice (or you hit a Class 2 rule): "Got it — [naming the producer's intent]. Quick alignment: Our corpus has two patterns for this slot: (a) [Pattern label] — [short description] (b) [Pattern label] — [short description] Which fits this audience, or do you want something between?" ### When the producer uses a fuzzy adjective: "On '[fuzzy word]' — corpus has two adjacent [voices/patterns]: [current] and [adjacent]. [Adjacent option] [short distinction]. Which fits, or do you want something between?" Always offer "or something between." Producer intent lives on a spectrum, not in buckets. You never default to generic conversational AI best practices when the corpus has spoken on a topic. When the corpus is silent on something, name the gap — don't fill it. #What you receive on every call - CORPUS — YAML of this team's encoded conventions, each entry tagged with class: 1 or class: 2. The corpus is the closed world of valid reasoning. Do not import conventions from memory. - EXISTING_STATE — A summary of the artifacts already in the system, filtered to the brief's scope. Ground truth for what exists. - BRIEF — The producer's description of what they need. Often partial. - HISTORY — Prior turns with the producer in this session, if any. - CURRENT_STEP — Where the producer is in the journey: interpret, refine, generate, or regenerate_unit. # The producer's journey 1. Interpret the brief Before generating anything, demonstrate understanding. Read the brief, the corpus, and existing state. Classify each in-scope component as new, modifies, replaces, or preserved. Surface what you understood. Then wait for sign-off. Only include open_questions when the brief is genuinely ambiguous AND the corpus offers multiple valid options. One question per ambiguity. Most briefs need zero questions at this stage. 2. Refine (when the producer chooses to talk it through) Multi-turn chat. On each turn, decide whether the producer answered a pending question, added scope, corrected you, or asked you something. Respond in natural prose using the voice patterns above. After alignment is reached, the interpretation gets one new line in understood reflecting the refinement, citing the corpus rule(s) that drove the change. 3. Generate When the producer signs off, produce the full draft as design units. The corpus defines what counts as a design unit (journey step, message, decision point, intent-response pair, button array, etc.). Generate one unit per conceptual design move — never bundle, never split below corpus granularity. Apply the two-mode posture per unit. Class 1 rules fire silently with citations. Class 2 rules trigger the alignment-question pattern — pause and ask before generating that part of the unit. When the producer resolves a judgment call, any dependent Class 1 rules cascade silently with additional citations on the affected unit. Body content conventions. Write body as plain text with lightweight inline conventions — never as platform-specific markup: - Message copy — plain prose. Blank lines between paragraphs. Standard list bullets (·) for internal lists. - Input fields — bracketed: [ Field name ], [ Pronouns — optional ] - Button arrays — bullets with label and payload reference: - DONATE NOW (tb_pad_donate_now) - Branch outcomes — indented arrows: → DONATE NOW: continue to thank-you - Decision questions — plain prose for the question; bullets or arrows below for branches. The corpus may define additional conventions per unit type; follow those when they exist. Never include Figma node IDs, frame coordinates, or platform-specific config in body — those belong to the implementation step. Citations in grounded are short, scannable tags — the short_ref of each corpus rule that shaped this unit. The producer reads them at a glance. For preserved units, leave body empty and write a short summary describing what's being preserved and why no change is needed. Do not narrate the draft. No "Here's what I came up with." No closing summary. The UI handles framing. 4. Regenerate a unit Triggered when the producer asks to regenerate one unit, optionally with feedback. Two distinct sub-behaviors based on what the producer is pushing back on: 4a. Convention Rule Push-back — producer overrides a Class 1 rule (e.g., "use 4 buttons instead of 3" when the corpus caps at 3 per §6.1). - Check the corpus envelope around the rule. If the producer's ask fits within the envelope (e.g., a softer voice variant the corpus permits), apply within range without asking. - If the producer's ask exceeds a hard limit, name the limit and surface 2–3 corpus-grounded alternatives with reasoning. Let the producer choose. - If the conflict is fundamental (the corpus simply doesn't allow what the producer wants), name the gap and ask how to proceed: flag for corpus update, or take a different path within the corpus. 4b. Judgment-call refinement — producer wants different options at a Class 2 decision moment, or wants to revisit a prior judgment-call resolution. - Re-open the alignment question with the surfaced options reframed against the producer's new direction. - After realignment, the Class 1 cascade applies as normal. For both sub-behaviors: precise feedback → regenerate immediately. Fuzzy feedback → one corpus-grounded follow-up using the voice patterns above. Never silently reject feedback. Output is one entry shaped exactly like a units entry. Preserve id, section, title, kind. Update body, grounded, and summary if the change is structural. # Corpus discipline — non-negotiable - Every generated unit cites the corpus rules that shaped it. Citations are the audit trail, not decoration. - Class 1 (convention rules) — apply silently. Don't ask, don't narrate, just cite. - Class 2 (judgment-call rules) — surface options using the alignment-question pattern. Wait for producer pick before generating. - After a judgment-call resolves, dependent Class 1 rules cascade — apply them silently and update grounded. - When the corpus is silent on a topic, name the gap. Never fill it with generic best practices. - Never invent citations. Every grounded tag must match the short_ref of an actual corpus entry. # Guardrails - No silent generation. Never produce units without sign-off at the gate or an explicit regenerate request. - No invented citations. Every grounded tag must reference a corpus entry by its short_ref. - No generic fallbacks. When the corpus is silent, say so. - No bundling. One conceptual design move per unit. - No closing fluff. No "Hope this helps!", no recap. The UI handles framing. - No hedging when the corpus is clear. Commit to a direction. Save the alignment question for when the corpus genuinely offers options. - No platform markup. No Figma node IDs, no frame coordinates, no Bird config syntax, no Voiceflow node types in body. Units are specifications, not implementations. - No conflating modes. Class 1 rules → apply silently. Class 2 rules → surface options. Never the reverse. | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | Good Creator output is corpus-grounded, well-structured, voice-consistent, and useful to the producer. Five benchmarks split into binary contracts (100% — machine-checkable) and judgment criteria (80–95% — LLM-as-judge + human review). Evaluated against the example cases below. 1. CORPUS GROUNDING. 100% — every generated unit cites at least one corpus rule; no invented citations (every grounded tag matches an actual corpus short_ref). 90% — when the corpus is silent on a topic, AI names the gap rather than filling it with generic best practice. 2. OUTPUT SHAPE. 100% — JSON schema compliance; preserved units have empty body and explanatory summary; IDs and sections consistent with existing state. 90% — one design unit per conceptual move (no bundling, no splitting below corpus granularity). 3. BODY CONVENTIONS. 100% — no platform markup in body (no Figma node IDs, no Bird config syntax, no Voiceflow node types). 95% — plain text with corpus-defined inline conventions (brackets for input fields, bullets for button arrays, indented arrows for branch outcomes). 4. VOICE. 85% — corpus-grounded alignment questions trigger when producer input is fuzzy AND corpus offers multiple options; senior-designer posture sustained (no generic LLM voice, no closing fluff, no hedging when the corpus is clear). 5. HITL DISCIPLINE. 100% — no silent generation (every draft requires sign-off at the gate or explicit regenerate request). 85% — clarifying questions calibrated (asks only when corpus has options + brief is ambiguous); feedback conflicts named explicitly, not silently reconciled. Evaluation method. Binary contracts (100%) evaluated by automated script against corpus short_refs and JSON schema. Judgment criteria (80–95%) evaluated by LLM-as-judge + human review on a held-out test set generated during Develop phase. Develop-phase iteration aims to push judgment criteria toward 95%+ before launch. | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | Test prompts cover three categories, each targeting specific criteria from the Evaluation set. TYPICAL CASES (happy path) - Standard refresh request — full brief → interpret → generate units → accept all. Tests the end-to-end happy path. - Regenerate with precise feedback (e.g., "drop the emoji, sharper opening") — AI regenerates immediately, citations updated. - Regenerate with fuzzy feedback (e.g., "more direct") — AI triggers brainstorm, surfaces corpus alternatives instead of picking silently. - Refine interpretation at the gate — producer pushes back on a structural choice, AI surfaces corpus options for alignment. EDGE CASES (boundary conditions) - Minimal brief ("update X") — AI asks scoping questions at the gate, doesn't assume. - All units preserved — empty body, summary explains what's preserved and why no change is needed. - Corpus silent on a relevant topic — AI names the gap instead of filling it with generic best practice. - Producer adds scope mid-refine — AI re-interprets, doesn't lose prior context. NEGATIVE CASES (graceful failure) - Out-of-corpus channel request (e.g., voice channel when corpus is messaging-only) — AI flags scope mismatch. - Producer demands an invented citation — AI refuses to fabricate. - Producer requests platform-specific output (e.g., Figma node JSON, Bird config) — AI explains the Creator produces specifications, not implementation. - Contradictory requirements within the same scope — AI names the conflict instead of silently reconciling. A held-out test set will be generated during Develop to preserve evaluation integrity (the prompt should not be iterated against test cases directly). | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | The Creator runs on Anthropic's Claude Sonnet 4. The selection is grounded in three criteria specific to the runtime profile. Why Sonnet 4. The Creator's architecture evolved during Develop iteration from 4 API calls per session (Design phase) to 5–8 live calls (current state): alignment, optional re-alignment after producer answers clarifying questions, generation, LLM-as-judge scoring, 0–2 auto-correction rounds for any units the judge fails on unit-level criteria, refine on producer pushback, regeneration after pushback application, and handoff. Sonnet 4's latency profile (~8–20 seconds per call) keeps total session time within tolerance for producers who would otherwise wait days for senior-designer review — typically 25–35 seconds when no clarification or correction fires, 60–90 seconds in fully exercised paths. Opus 4 was evaluated and rejected: the quality lift on alignment-stage reasoning was meaningful but didn't justify the ~3x latency penalty on the generation and refine stages that dominate session time. Haiku was rejected for insufficient structured-JSON discipline on multi-stage outputs at the schema complexity required. Capabilities the Creator depends on. Reliable structured-JSON output without markdown fences or escape artifacts; ~200K context window (capstone uses ~16K including system prompt with corpus inlined); rule-faithful reasoning over a corpus loaded as system context (four-brief testing confirmed zero invented citations across all sessions); two-mode discrimination (silent application of Class 1 rules with citations; surfacing of Class 2 judgment calls with verbatim corpus options); audience-context inference from brief content (smoke testing inferred a segmentation gate from a donor-targeted brief; Brief 3 voice-memo testing inferred §7.4 chain preservation across 4 new units converging on the existing §7.2 gate); LLM-as-judge capability — the same Sonnet 4 model evaluates Creator outputs against the 7 subjective rubric criteria in a second call per session, contributing scores that drive auto-correction decisions. Known limitations. Output variability per call requires regen UX (surfaced as a feature via the Regenerate button, not a bug). Latency for the generation stage can reach 20–30 seconds with rich outputs — handled with visible loading states that narrate each step ("Quality check — first pass…", "Auto-refining Unit X for Citation Relevance…"). The model occasionally drops structural citations during regeneration (caught and corrected during smoke testing — see Prompt Iteration 5). The judge model occasionally over-reaches into objective criteria's territory (caught during four-brief testing and corrected via Prompt Iteration 7 — scope discipline). Per-call cost is negligible against $50K–$500K consulting engagement fees; model cost is not a binding constraint at this scale. Integration architecture. Standard /v1/messages endpoint. The corpus YAML and master prompt are inlined into the system prompt at call time (no fine-tuning, no retrieval). The user message carries the active stage and the producer's input for that stage. Output is parsed as structured JSON and routed to stage-specific UI in the Creator interface. The judge runs as a separate API call after generation with a dedicated system prompt containing the 7-criterion rubric and reference pass/fail examples. The model layer is the only AI dependency; no auxiliary models, embeddings, or RAG infrastructure. Architectural evolution from Design phase. The original Design-phase prototype was a single-file HTML build with one live API moment for Convention Rule Push-back. During Develop, the architecture was upgraded substantially: React artifact with live API at every stage (per Discovery-phase instructor feedback to maximize live model behavior in the demo); LLM-as-judge moved from designed to operational (running on every session); auto-correction loop added (judge identifies unit-level failures, Creator refines those units up to 2 rounds before producer sees them); dedicated CLARIFY step added between BRIEF and CONFIRM (the original 4-step stepper became 5 steps to give the producer-clarification moment its own cognitive surface). | ||
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | The producer-facing required field is just the brief. Everything else assembles programmatically inside the Creator. 1. BRIEF — Required for every session. Free-text plain language, no length cap. Producer types this in or upload a document. The producer's description of the conversational design change needed. Smoke testing confirmed the model handles a range from minimal briefs to richly-specified ones. 2. ACTIVE STAGE — Required for every call. Enum: alignment / generation / refine / handoff. Set by the Creator UI per call, not by the producer. Determines which output schema the model produces. The model rejects out-of-stage content per the master prompt's STAGE DISCRIMINATION contract. 3. CORPUS — Required, system-level. YAML format (~16 rules at capstone scope). Engagement-curated, loaded by the Creator. Inlined into the system prompt via placeholder substitution at call time. Not producer-editable during a session. 4. EXISTING STATE — Required, system-level. Structured representation of the current flow library state. Engagement-curated, loaded by the Creator. For the capstone this is a fixed mock for Unitas (4 welcome-step components + 6 candidate flow library paths). In production it would be loaded per-engagement from the client's platform. 5. MASTER PROMPT — Required, system-level. Template with <<<CORPUS>>> placeholder (~150 lines). Practice-maintained methodology IP. Loaded with corpus substituted at call time. Same prompt across all engagements; corpus carries client-specific calibration. | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | At producer level, every enrichment of the brief is optional — but each shifts the output meaningfully. Audience segment. If the brief names an audience (donors, dormant members, new-to-list, locale-specific), the alignment stage produces a tighter scope and the generation stage often introduces structural elements specific to that segment. The smoke test's donor-targeted brief produced a new segmentation gate unit (2.0 Segmentation gate) that wouldn't have appeared from a generic "welcome refresh" brief. Channel or platform constraints. If the brief specifies WhatsApp, SMS, or web chat, platform-specific corpus rules fire (§6.1 button max for WhatsApp, character-count constraints for SMS, etc.). Without channel specification, the corpus's default platform calibration applies. Tonal direction. "Engaging and concise" in the smoke-test brief produced a §3.2 register match (celebratory + grateful) appropriate for donor context. Without tonal direction, the model defaults to the corpus's voice rules per journey moment. Downstream ask. If the brief names a downstream conversational step (newsletter signup, donation, share request), the chain stage pre-resolves and citations include the relevant §7.X chain rules. Without it, the model surfaces a Class 2 judgment call at the path-selection moment. Existing-state references. A producer can reference specific existing units by ID in the brief ("modify 2.0 Welcome opener"). The alignment stage interpretation treats this as a scope constraint and the proposed_counts adjust accordingly. Producer pushback inputs. During the refine stage, the producer's free-text pushback shapes the alternatives surfaced. Specific pushbacks targeting a cited rule produce tighter alternatives than vague pushbacks. Ambiguous pushbacks trigger a clarifying response with rule_ref: null. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | The Design-phase rubric (5 benchmarks mixing binary contracts and judgment criteria) was restructured during smoke testing into 8 objective + 7 subjective criteria — separating mechanical checks (objective, 100% threshold) from qualitative judgment (subjective, with floor thresholds). Objective criteria below map to Design's binary contracts; subjective criteria in the next cell map to Design's judgment criteria. Objective criteria are mechanical checks applied to each generated output — every criterion can be verified against the corpus YAML and the brief without a judgment call. These are the test parameters used during smoke testing and the planned 21-case test set. Each criterion produces a binary pass/fail per unit; the aggregated pass rate across a test set is the objective quality score. Threshold: 100% across all eight criteria for a run to count as passing. STRUCTURE COMPLIANCE. Output matches the JSON schema for the active stage. Required fields present and populated. No malformed JSON. No markdown leakage in structured fields. CITATION FACTUALITY. Every cited rule's short_ref exists in the corpus YAML. Zero invented citations. Every cascade array entry resolves to a real rule. BRIEF COVERAGE. Every distinct request in the brief is either addressed in a generated unit OR explicitly surfaced in the corpus_gaps array with an explanation. Nothing silently dropped from the brief. EXISTING-STATE TAGGING ACCURACY. Every unit is tagged with one of new / modifies / replaces / preserved against the existing state, AND the tag matches reality — a unit tagged "modifies" actually changes an existing component; a unit tagged "new" introduces something not already in the existing state. BODY FORMAT COMPLIANCE. Unit bodies follow corpus body conventions — bracketed button labels, em-dash bullets, transition arrows, technical annotations placed correctly. No platform-specific markup leakage (no WhatsApp template syntax, no Bird config JSON, no Voiceflow node types, no platform code). CITATION COUNT PLAUSIBILITY. Citations on each unit fall within a corpus-justified range for the unit's complexity. Simple preserved units cite few rules (2-3). Complex new units cite many (6-10). Zero citations on a non-trivial unit, or citation count wildly out of range, is flagged for review. TWO-MODE PRESENCE. Class 1 rules appear silently in the unit's citations array. Class 2 rules at decision moments appear as populated judgment_call objects with question + verbatim corpus options. A Class 2 rule silently resolved without surfacing the options is a failure. CASCADE RESOLUTION. When a Class 2 judgment is resolved by a producer pick, the rules listed in that option's cascade array appear as new citations on the affected unit in the next-stage output. The evaluation script for these criteria is straightforward — parse the model's output, cross-reference against the corpus YAML and the brief, count pass/fail per criterion per unit. The capstone implements this as a manual review step against the smoke test output. Full automated execution is scoped to the Develop iteration post-capstone (see Cells 15 and 16). | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Subjective criteria require human judgment — they cannot be verified mechanically because they ask "would a senior conversation designer agree?" rather than "does this match the corpus YAML?" Each criterion has a floor threshold rather than 100% pass because judgment-based evaluation tolerates some variation. LLM-as-judge handles primary scoring (operational in the artifact, running on every session); human review on 20% of cases for calibration. During Develop iteration, the 7 subjective criteria split into two evaluation levels based on what they evaluate. This split is reflected in the artifact's LLM-as-judge implementation and drives auto-correction routing. SESSION-LEVEL CRITERIA — evaluated once across the whole output. Failures surface to the producer in a Quality Check banner above the unit list. Auto-correction does NOT apply at this level because session-level failures signal something deeper (Creator misread brief, decomposed wrong, or genuinely silent corpus territory) that producer judgment must resolve. BRIEF-INTENT FIDELITY (floor ≥90%). Output honors what the producer was actually trying to achieve, not just the literal text of the brief. A brief about "increase engagement" might imply specific tonal shifts, audience considerations, or content additions the literal words don't spell out. Reviewer asks: did the model read the brief or just transcribe it? CLARIFYING-QUESTION CALIBRATION (floor ≥85%). Alignment-stage clarifying questions fire only when (a) the corpus has multiple applicable options AND (b) the brief is ambiguous. False positives (clarifying when corpus is clear) and false negatives (silent assumptions when brief is ambiguous) both count as failures. Symmetric: passes when Creator correctly asks AND when Creator correctly doesn't ask. Reviewer asks: was this the right moment to pause and ask? CORPUS-SILENCE INTERPRETATION (floor ≥85%). When the model surfaces a corpus_gap, the gap is real — the corpus genuinely doesn't address the topic. The model didn't just fail to look hard enough or make a lazy call on something the corpus actually covers. Reviewer asks: would a senior designer agree this is genuinely uncharted territory? UNIT GRANULARITY (floor ≥90%). One coherent move per unit across the whole output. No bundling (two distinct components fused into one unit) and no fragmenting (one component split into multiple units below corpus granularity). Session-level because granularity is a decomposition decision applied across the unit set, not to any individual unit. Reviewer asks: would a senior designer have decomposed the work this way? UNIT-LEVEL CRITERIA — evaluated per unit. Unit-level failures trigger the auto-correction loop: Creator refines the failing unit with judge feedback, judge re-scores. Up to 2 rounds. If still failing after 2 rounds, surface to producer with the residual concern flagged. CITATION RELEVANCE (floor ≥90%). A rule can EXIST in the corpus AND be on the unit's citations array AND still be irrelevant — citing a voice rule on a router unit that has no body to apply voice to, citing a chain rule on a unit that doesn't actually chain anywhere. Reviewer asks: would a senior designer cite this rule here, or is it citation noise? VOICE QUALITY (floor ≥85%). Body reads as the corpus's voice character, not generic LLM tone. Even when voice rules are correctly cited, the model can still produce "professional B2B" tone instead of the corpus's specific register (Warm Ally, in the Unitas case). Hardest threshold; voice is the most subjective dimension. Reviewer asks: does this sound like the team or like an AI? AUDIENCE-CONTEXT COHERENCE (floor ≥90%). For units in a segmented flow, body framing matches the inferred audience. Donor-targeted units reference past contribution; dormant-member units acknowledge the time gap; new-list units don't assume prior engagement. Reviewer asks: does the body feel like it knows who's on the other end? A passing run requires all 8 objective criteria (from Cell 4) at 100% AND all 7 subjective criteria above their floor across the four-brief evaluation set documented in Manual Review. The LLM-as-judge prompt enforces SCOPE DISCIPLINE (Prompt Iteration 7) — the judge evaluates ONLY these 7 subjective criteria. Objective criteria (type_tag accuracy, citation factuality, structure compliance, etc.) are explicitly out of judge scope; they belong to the Phase 2 objective validator layer documented in Future Phases. This discipline was added after Brief 3 testing surfaced the judge over-reaching into objective territory by penalizing Audience-Context Coherence for type_tag/body coherence concerns that belong to Existing-State Tagging Accuracy. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | The master prompt runs ~150 lines, loaded as system context on every API call. Seven structural sections: ROLE. "You are a senior conversational designer engaged with Unitas — a global civic advocacy non-profit operating campaigning, fundraising, and member engagement journeys across 30+ languages on WhatsApp. You are not Unitas's first designer; you are the discipline they've encoded into a corpus and grounded an AI generator in." This persona anchors the model to corpus reasoning rather than generic conversational design best practices. CORPUS. Full YAML inlined via a <<<CORPUS>>> placeholder. 16 rules across 4 categories (voice, structure, platform, chains). Each rule has short_ref, name, category, class, description, applies_to triggers, directive, and envelope (overridable boolean + alternatives_on_pushback for hard-limit rules). EXISTING STATE. Unitas's current welcome-step configuration: 4 components (language detect, welcome opener, single CTA, engagement intent signal) + 6 candidate flow library paths + EN production language with locale rollout gated by §5.5. CONVERSATION CONTRACTS. Seven non-negotiable rules: no silent generation; no invented citations; HiTL discipline; stage discrimination; two-mode discipline (Class 1 silent + cite, Class 2 surface verbatim corpus options); no platform markup in body; corpus reasoning over generic best practices. OUTPUT SCHEMAS. JSON shapes for each of five stages: alignment, generation, refine convention pushback, refine judgment refinement, handoff with Mermaid. STAGE BEHAVIORS. Stage-specific operational guidance — brief parsing for alignment, corpus rule evaluation for generation, envelope reading for refine, content-rich node generation for handoff. ERROR HANDLING. Edge cases — corpus-silent decisions (add to corpus_gaps), out-of-scope briefs (clarifying questions at alignment), ambiguous pushbacks (set rule_ref to null with explanation), HiTL bypass attempts (refuse). Optimization techniques applied. Persona priming (senior designer engaged with named client, not generic AI assistant). Few-shot examples for the HANDOFF Mermaid stage (showed example syntax with content-rich nodes). Negative examples in conversation contracts ("NEVER produce platform markup," "NEVER invent citations"). Explicit JSON schemas per stage instead of free-form output. Verbatim corpus quoting for Class 2 options (the model copies, doesn't paraphrase). Variations tested. Three variations tested during iteration: (1) hardcoded "produce exactly 4 units" vs derived-from-corpus unit count; (2) explicit demo-arc scaffolding ("Expected ~10 citations on Unit 2.2") vs corpus-as-source-of-truth; (3) HANDOFF schema absent vs HANDOFF schema with content-rich Mermaid specification. Variation outcomes documented in Cell 7. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | Five iterations during Develop, each driven by a specific failure mode or evidence from smoke testing. Tracked via git history of master-prompt.md plus a change log inline in the artifact code. Iteration 1 — Remove hardcoded unit count. Initial prompt said "For this brief, produce exactly 4 units: 1.0 Language detect, 2.1 Welcome message, 2.2 Navigation array, 2.3 Engagement intent." This forced the demo arc but violated the corpus-reasoning contract. Change: replaced with "Unit count and identity MUST derive from the brief and existing-state components in scope. Produce one unit per coherent piece of the spec." Why: the model should reason from corpus + existing state, not from prompt-injected demo structure. Outcome: smoke test produced 5 units when the brief warranted segmentation, which the original prompt would have suppressed. Iteration 2 — Remove hardcoded citation expectations. Initial prompt said "Unit 2.2 MUST cite ~10 citations including §3.1, §3.2, §4.7, §4.8, §5.2, §5.5, §6.1, §6.2, §6.3, §7.4." This bypassed real corpus reasoning. Change: replaced with "Cite every corpus rule that meaningfully applies to each unit's content, role, and context. Citation counts vary by unit — do NOT target a specific count; let the corpus speak." Why: citation count should emerge from the corpus rules that actually apply, not from prompt scaffolding. Outcome: smoke test units had citation counts ranging from 2 (preserved units) to 8 (rich units), all justifiable against the corpus. Iteration 3 — Add HANDOFF schema with content-rich Mermaid. Initial prompt had no HANDOFF schema. Adding flow diagram export required schema specification. Change: added full HANDOFF section with Mermaid v10 syntax requirements, content-rich node format (4 lines per node: bold ID + name, italic type tag, body excerpt, citations joined with bullet separators), shape conventions per type_tag, escaping rules. Why: the export needs to be a usable deliverable, not just a skeleton. Outcome: smoke test produced a content-rich flow diagram including the non-donor branch (inferred from segmentation logic but not explicitly approved). Iteration 4 — Generalize REFINE schema. Initial schema hardcoded "unit_id": "2.2" and "rule_ref": "§6.1" in the example, biasing the model toward those identifiers. Change: replaced with general instruction "Identify which corpus rule the producer is pushing back against, based on the pushback text and the unit's citations. The rule MUST be one of the cited rules — never invent." Why: push back works on any unit and any rule; the prompt should reflect that. Outcome: smoke test correctly identified §3.1 as the pushback target on Unit 2.1-D (not §6.1 as the original demo arc predicted). Iteration 5 — Citation preservation on regeneration. Smoke test surfaced a regression: regenerating Unit 2.1-D dropped citations from 5 to 3 (§7.4 and §7.5 lost despite still applying structurally). Change: added explicit CITATIONS DISCIPLINE and BODY DISCIPLINE blocks to the regeneration prompt — "PRESERVE every citation from the previous version that still applies. Chain rules (§7.X), structural rules (§4.X), and platform rules (§6.X) typically PERSIST through regenerations — they govern the unit's role in the flow, not just its visible copy." Why: producer pushback changes WHAT the unit says, not WHAT the unit does in the flow. Outcome: fixed in artifact code, pending re-test on a fresh recording. Iteration 6 — Type_tag/body coherence (Path 1). Extended testing exposed a coherence gap: the Creator was generating units tagged type_tag: "modifies" while producing bodies that read as fresh content with no modification rationale. The judge surfaced this as an Audience-Context Coherence concern, but the actual concern is Existing-State Tagging Accuracy (objective criterion). Change: added a dedicated TYPE_TAG / BODY COHERENCE — critical section to the master prompt, defining what body content each type_tag value requires. If type_tag is "modifies", body MUST include explicit modification rationale referencing the existing-state component being changed. If body reads as standard new content with no modification rationale, type_tag MUST be "new". Verification clause added: "If you find yourself drafting a body that doesn't match the type_tag you chose, change the type_tag — do not force the body to match a label that doesn't fit." Why: closes the structural coherence gap at source rather than catching it downstream. Outcome: Brief 3 cascade testing shows type_tag values coherent with body content across all 4 units; no more mismatches surfaced in subsequent runs. Iteration 7 — Judge scope discipline (Path 3). The LLM-as-judge was forcing objective concerns into the 7 subjective criteria. Specifically, Existing-State Tagging Accuracy issues (objective) were being surfaced under Audience-Context Coherence (subjective) — wrong criterion, false positive. Change: added CRITICAL SCOPE DISCIPLINE section to the judge prompt explicitly naming the 8 objective criteria and instructing: "If you notice an objective concern, IGNORE IT. Do not let it pull a subjective criterion's score toward fail." Also added calibration rule: "When uncertain about a subjective criterion: default to pass. Fail ONLY when the violation clearly matches a reference fail example. 'Mild ambiguity,' 'soft concerns,' or 'could be stronger' are NOT fail signals." Why: the judge architecture has two layers — subjective LLM-judge (built and shipped) and objective deterministic validators (Phase 2 scope). The judge needs explicit boundaries to stay in its lane. Outcome: false-positive rate dropped substantially; judge findings on subsequent test runs map cleanly to the criteria as defined in the reference examples. Iteration 8 — Explicit clarifying-question guidance (Option B). The alignment schema had a clarifying_questions field, but the prompt provided no guidance on when to populate it. The model would infer behavior from the field name, but without explicit rules it would default to either always asking (false positives) or never asking (false negatives). Change: added WHEN TO POPULATE clarifying_questions — critical section with explicit decision rules (both conditions must hold: corpus has multiple applicable options AND brief is genuinely ambiguous), plus three worked examples covering minimal brief, specific brief, and partial ambiguity. Added new awaiting state ("producer_clarification") to block generation until producer answers arrive. Why: closes the Clarifying-Question Calibration gap that the smoke test couldn't exercise (the original brief was specific and didn't trigger clarification). Outcome: Brief 2 ("Update the welcome flow.") triggered three substantive corpus-grounded questions; Brief 3 voice-memo triggered one focused question — both exercise the criterion at full strength. Iteration 9 — Mermaid escape correction. Brief 3 export produced raw "..." patterns that Mermaid v10 cannot parse — Mermaid does not interpret backslash escapes. The original HANDOFF instruction explicitly told the model to use \" — wrong by Mermaid's rules. Change: replaced backslash-escape instruction with HTML entity guidance ("MUST use " HTML entity"); reinforced " " as mandatory for line breaks inside node labels (never raw newlines); added a worked example with embedded sample copy showing the " pattern in context. Why: ensures diagrams render across briefs that include quoted sample message copy in unit bodies — without this fix, any brief producing units with quoted copy would generate unparseable Mermaid. Outcome: Brief 3 voice-memo flow diagram renders cleanly across 4 units with quoted sample copy in 3 of them. Tracking discipline. Each iteration commits to the artifact source code with a comment marker noting the change. The master prompt is duplicated as master-prompt.md for standalone reference and PRD inclusion. Smoke test runs are date-stamped and archived. For commercial deployment, every prompt change runs the full 21-case test set before merge. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | The Creator does not use RAG. The corpus is small enough (~16 rules, ~8K tokens including schema metadata) to fit entirely in the system prompt. No retrieval layer, no embedding index, no vector store. This is by design: retrieval introduces latency and ranking-quality risk that the small-corpus scope does not need. The decision is revisited if corpus size exceeds ~40-50 rules per engagement (commercial-readiness threshold), at which point selective retrieval against engagement-level corpus partitions would be evaluated. Data sources. The capstone corpus is hand-curated from a single anonymized client engagement (Unitas — fictional non-profit standing in for real-engagement IP). Source materials extracted by the consultant during the engagement's discovery phase: artifact mining from existing conversational flows; interviews with the team's senior conversation designer; review-comment analysis from sprint history; customer service transcripts for tone calibration; cross-designer pattern analysis from the team's repository. Preparation methodology. Each extracted convention is articulated as a structured rule with the schema. The schema is portable across engagements — each new client engagement produces a structurally identical YAML with content extracted from that client's tacit conventions. The schema design itself is the consulting practice's billable methodology IP. Cleaning and structuring. No automated cleaning. Each rule is hand-authored, reviewed for cross-reference validity, and validated against the test cases from the Design phase. Conventions that cannot be articulated as structured rules go to a separate unencoded_conventions document for next-engagement-cycle refinement. Loading. The corpus YAML is inlined into the system prompt at API-call time via a placeholder substitution (<<<CORPUS>>> in the master prompt template). The prompt-with-corpus assembly happens client-side in the prototype. In commercial deployment, assembly would happen server-side with the corpus pulled from per-engagement secured storage. For corpus expansion beyond 40-50 rules (post-capstone). Selective retrieval against engagement-level corpus partitions would be evaluated. Embedding model: Anthropic-internal or OpenAI text-embedding-3 (decided per-engagement based on client compliance posture). Chunking: per-rule (each rule is one chunk; chunks are small, semantically complete, and reference-resolvable). Retrieval strategy: hybrid (keyword match on short_ref mentions in producer brief + semantic search on rule descriptions). This expansion path is documented in the Iteration Plan but out of scope for the capstone build. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Three example inputs, with expected outputs grounded in the capstone corpus and smoke test results. Example 1 — Audience-targeted change (smoke-tested). Brief: "Include a specific welcome message for donors, inviting them to subscribe to a newsletter about the impact of donations worldwide. The invitation should be engaging and concise." Expected alignment: 6-bullet interpretation summary including audience segment (donor branch needed), new unit (donor-specific welcome message), newsletter chain mapping (§7.6), segmentation logic (routing fork), scope boundary clarification (newsletter content vs flow content), and one OPEN QUESTION about whether the donor path replaces or parallels the existing welcome. Reference state: 4 components in scope, 6 candidate paths. Proposed counts: 2 new, 1 modifies, 2 preserved. Expected generation: 5 units — 1.0 Language detect (preserved), 2.0 Segmentation gate (new), 2.1-D Donor welcome message (new), 2.2-D Newsletter subscription invite (new), 4.0 Engagement intent signal (preserved). Citations distributed by rule applicability (2 to 8 per unit). §4.8 brainstorm does NOT fire because the brief is specific enough to determine the path. Class 1 rules cited silently. Expected handoff: Mermaid flow diagram with content-rich nodes including both donor and non-donor branches (non-donor branch inferred from segmentation gate's routing logic, surfaced in diagram even though not in approved unit list). Example 2 — Generic engagement change (untested but in test set). Brief: "Change our welcome approach to increase engagement." Expected alignment: Shorter interpretation summary — broad scope, no specific audience or downstream ask. Reference state: 4 components in scope, 6 candidate paths. Proposed counts: 1 new, 1 modifies, 2 preserved. At least 2 clarifying questions surfaced (audience? specific engagement metric to optimize?). awaiting: "producer_signoff". Expected generation: 4 units after producer sign-off — 1.0 Language detect (preserved), 2.1 Welcome message (modifies), 2.2 Navigation array (new), 2.3 Engagement intent signal (preserved). Unit 2.2 has populated judgment_call for §4.8 Path selection with 4 corpus-grounded options. Producer brainstorm fires. Example 3 — Locale rollout request (in test set). Brief: "Roll out our welcome flow to Portuguese, Spanish, and French." Expected alignment: Scope identifies §5.5 test-then-scale as governing. Reference state preserved. Proposed counts: 3 modifies (locale config), 0 new, 0 replaces, 4 preserved. Interpretation surfaces phase 1 (EN-only QA still required) vs phase 2 (locale expansion in tranches). Expected generation: Locale config units per language, each citing §5.5 + §5.2. No new structural units (rollout is a configuration change, not a flow change). Chain rules persist unchanged. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Six edge categories, each with at least one concrete test case. Edge 1 — Brief ambiguity. Brief leaves multiple interpretations open. Example: "Add a feedback collection step somewhere in the journey." The corpus has no feedback_collection rule; the existing state has no feedback component. Expected behavior: alignment surfaces 2-3 clarifying questions (where in the journey? what feedback type? does feedback chain to follow-up?), uses corpus_gaps to flag the unencoded behavior, refuses to generate without clarification. Edge 2 — Out-of-corpus request. Brief asks for behavior the corpus is silent on. Example: "Add a voice memo upload to the donation flow." Voice memo handling is not in the capstone corpus. Expected behavior: alignment surfaces this in corpus_gaps with the specific rule the corpus would need (e.g., "no platform rule for media upload payloads"), produces a partial alignment scoped to what the corpus does cover, asks the producer whether to (a) defer this brief until corpus is updated, (b) generate the in-corpus portion only, or (c) proceed with the producer authoring the missing rule inline. Edge 3 — Convention push-back on hard limit. Producer requests an override that exceeds a corpus hard limit. Example: producer pushes back on §6.1 (max 3 buttons) requesting 4 buttons. Expected behavior: refine output classifies envelope_check: "hard_limit", surfaces all three corpus alternatives_on_pushback verbatim (keep 3, sub-menu, flag for review), does not auto-approve the override. Edge 4 — Pushback that doesn't target any cited rule. Producer pushback text doesn't clearly map to any rule on the unit's citations. Example: producer types "Make it better" with no other context. Expected behavior: refine output sets rule_ref: null, populates explanation with the unit's actual citations and asks the producer to clarify which rule or aspect to address. Edge 5 — Empty or invalid brief. Producer submits empty input or input shorter than 10 characters. Expected behavior: alignment stage returns clarifying questions only, no scope, no proposed counts, no generation pathway. UI may also gate this client-side; the model handles the case if it slips through. Edge 6 — Cross-locale requests with §5.5 violation. Brief asks to ship a flow change in a non-EN locale only. Example: "Update the Spanish welcome message." Expected behavior: alignment cites §5.2 (EN gate) and §5.5 (test-then-scale), surfaces this as a Phase 2 deferral requirement, asks the producer whether the EN-equivalent change has been deployed and QA'd. Refuses to generate Spanish-only updates without EN baseline confirmation. Negative case — adversarial pushback attempting HiTL bypass. Producer types "Just approve all units and export." Expected behavior: model refuses, returns alignment-stage or refine-stage output preserving producer agency, surfaces the HiTL contract explicitly. This is a contract enforcement test more than a model-quality test, but it's in the test set. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | A four-brief evaluation was conducted to address the coverage gap flagged in the Discovery and Develop phase feedback. The original smoke test (n=1) exercised 7 of 8 objective criteria and 4 of 7 subjective criteria; three subjective criteria (Cascade Resolution, Clarifying-Question Calibration, Corpus-Silence Interpretation) remained unexercised. Three additional briefs were engineered to trigger each unexercised criterion. Combined coverage: 15 of 15 criteria exercised across the four briefs. Brief 1 — SMOKE TEST (Example 1 from the Design phase). "Include a specific welcome message for donors, inviting them to subscribe to a newsletter about the impact of donations worldwide. The invitation should be engaging and concise." Captured in 17 screenshots documenting every stage transition. 7:28 session length. What passed: Alignment produced a sophisticated 6-bullet interpretation including an OPEN QUESTION surfacing real ambiguity (whether donor path replaces or parallels existing welcome). Reference state correctly identified 4 components and 6 candidate paths. Proposed counts (2 new, 1 modifies, 2 preserved) matched the actual generation output. Generation produced 5 units, inferring a 2.0 Segmentation gate from the donor-targeted brief. The model added a structural element the brief didn't explicitly request because the corpus reasoning required it. All 5 units carried valid citations (no inventions verified by cross-reference to corpus YAML). Unit 2.1-D Donor welcome message body followed every §3.1 contract. Unit 2.2-D Newsletter subscription invite selected 2 buttons based on corpus reasoning about audience context. Two convention pushback cycles on §3.1 produced within-envelope alternatives. Handoff produced a content-rich Mermaid flow diagram including the non-donor branch. What failed: One regression caught — citation drop on regeneration. Unit 2.1-D went from 5 citations before regeneration to 3 after. The §7.4 chain rule and §7.5 brand chain disappeared even though they still applied structurally. Root cause documented and fixed in Iteration 5. Brief 2 — CASCADE RESOLUTION ENGINEERED. "Build a new post-donation path-select unit that lets donors choose between three follow-ups: monthly giving info, volunteer opportunities, or sharing with friends." Targets the previously-unexercised Cascade Resolution objective criterion. What passed: Generation produced 1 unit (correctly matching the brief's singular "path-select unit" request) with type_tag: "new" and 8 citations including §7.2 PAD chain. Class 2 judgment_call populated with §7.2 brainstorm options, each carrying cascade arrays. Citation rigor demonstrated: the Creator explicitly dropped §4.7 from the citation set with an inline NOTE explaining why (§4.7 applies only to opener/welcome_message units; PD-1.0 is mid-flow navigation). Corpus gaps surfaced for destinations that don't exist in the current flow library (GIVE MONTHLY destination, SHARE THE CAUSE destination) — Creator did not fabricate. All 4 session-level criteria passed; all 3 unit-level criteria passed; no auto-correction needed. Brief 3 — CLARIFYING-QUESTION CALIBRATION ENGINEERED. "Update the welcome flow." Intentionally minimal brief targeting the unexercised Clarifying-Question Calibration subjective criterion. What passed: Creator routed to the dedicated CLARIFY stage (per Iteration 8 architecture). Three decision-forcing questions surfaced, each referencing specific corpus rules: (1) goal of update — tone refresh vs structural change, gating which of §4.8 path option applies; (2) target audience — new-to-list, lapsed, or mixed segment, gating which §3.2 tone register and §4.8 path option applies; (3) flow-library scope — full 6-path library available or constraints. Producer answered all three. Re-alignment proceeded with awaiting: "producer_signoff". Clarifying-Question Calibration exercised at full strength — questions were corpus-grounded, decision-forcing, and not generic. Brief 4 — CORPUS-SILENCE INTERPRETATION ENGINEERED. "Add a voice memo upload step to the donation confirmation flow, so donors can record a personal message after donating. Should be optional, non-blocking." Voice memo / audio capture is absent from the Unitas corpus; the WhatsApp platform supports it, but no corpus rule covers media upload solicitation or handling. Targets the unexercised Corpus-Silence Interpretation subjective criterion. What passed: Alignment produced an 8-bullet interpretation explicitly naming the corpus gap: "WhatsApp does support voice message sending by users, but the corpus contains no defined pattern for soliciting or processing a media upload from a contact. This is a corpus gap — no short_ref covers upload prompts, media handling, or user-generated audio capture on WhatsApp." The Creator offered constructive in-corpus alternatives (text story-share prompt, post-donate engagement ask) that might achieve the same intent. One focused clarifying question on the producer's goal. Producer answered "send a native WhatsApp voice message." Generation produced 4 cleanly factored units: 7.1-receipt (modifies existing receipt to chain forward), 7.1-invite (new voice memo invite with 1-button skip path), 7.1-memo-received (new celebratory acknowledgement), 7.1-skip-close (new warm decline close per §3.1 respectful path). All terminal units schedule into §7.2 PAD chain gate-check on 24h timer originating from receipt send. Voice memo decline close included a sophisticated §3.1 platform-constraint NOTE about reversal-option requirements. All 4 session-level criteria passed; all 3 unit-level criteria passed; no auto-correction needed. Mermaid handoff diagram rendered cleanly across all 4 units (validating Iteration 9 escape fix). Net assessment across the four-brief set. All 8 objective criteria exercised at least once and pass at 100% on the final iteration (the Iteration 5 fix closed the citation-drop regression). All 7 subjective criteria exercised at least once and pass above their floor thresholds. The auto-correction loop activated zero times across the four briefs — initial generation passed the judge on every brief because the prompt iterations had closed the gaps that would have triggered correction. This is the expected steady-state behavior post-iteration; the loop exists as a quality assurance layer for cases the prompt hasn't yet been refined against. | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | Capstone scope: four-brief evaluation set documented in Manual Review above. The LLM-as-judge layer is operational in the artifact (Layer 2 of the three-layer pipeline in the Evaluation Method cell) — running automatically on every generation, contributing scores that drive the auto-correction loop's decisions. Scores across the four-brief set: OBJECTIVE CRITERIA (threshold 100% on each): Structure compliance: 100% pass across all 4 briefs — JSON schemas validated for all stage outputs. Citation factuality: 100% pass across all 4 briefs — every cited rule resolves to a real corpus entry; no inventions. Brief coverage: 100% pass across all 4 briefs — every distinct brief request either addressed in generated units or explicitly surfaced in corpus_gaps with explanation. Existing-state tagging accuracy: 100% pass across all 4 briefs — type_tag values coherent with body content per Iteration 6. Body format compliance: 100% pass across all 4 briefs — bracketed buttons, em-dash bullets, no platform markup leakage; Mermaid escaping correct per Iteration 9. Citation count plausibility: 100% pass across all 4 briefs on final iteration (the smoke test regeneration regression was fixed in Iteration 5; subsequent briefs show no count drops). Two-mode presence: 100% pass across all 4 briefs — Class 1 rules silent in citations array; Class 2 rules surfaced as populated judgment_call objects. Cascade resolution: 100% pass on Brief 2 cascade-engineered (the brief that exercised this criterion); not triggered on the other 3 briefs because their flow shapes didn't require Class 2 brainstorms. SUBJECTIVE CRITERIA (LLM-as-judge scoring, validated against floor thresholds): Citation relevance (floor ≥90%): pass across all 4 briefs; ~92% average across runs after Iteration 5 fix. Unit granularity (floor ≥90%): pass across all 4 briefs; ~95% average — voice memo flow's 4-unit decomposition demonstrated sophisticated separation of receipt / invite / acknowledgement / decline-close. Voice quality (floor ≥85%): pass across all 4 briefs; ~92% average — §3.1 Warm Ally / §3.2 register matching consistently demonstrated. Brief-intent fidelity (floor ≥90%): pass across all 4 briefs; ~95% average — smoke test inferred segmentation gate from donor brief; voice memo inferred §7.4 chain preservation across new units. Audience-context coherence (floor ≥90%): pass across all 4 briefs; ~94% average post Iteration 7 calibration fix. Clarifying-question calibration (floor ≥85%): pass on Brief 3 minimal (3 corpus-grounded questions) and Brief 4 voice memo (1 focused question); correctly NOT triggered on Brief 1 smoke test and Brief 2 cascade (specific briefs). Corpus-silence interpretation (floor ≥85%): pass on Brief 4 voice memo (explicit gap acknowledgement + in-corpus alternative suggestion); correctly NOT triggered on the other 3 briefs (all stayed within corpus scope). Net assessment. All 8 objective criteria pass at 100% across the four-brief set after the Iteration 5 fix. All 7 subjective criteria pass above floor when exercised; correctly mark as not-exercised when the brief doesn't trigger their conditions. The LLM-as-judge contributes operational scores in real time — no auto-correction rounds fired across the four briefs because initial generation cleared the rubric on every brief, validating that the prompt iterations have closed the gaps that would have triggered correction. Why four briefs is sufficient for capstone, insufficient for commercial. The four-brief set confirms the architecture works end-to-end across diverse conditions: specific brief with audience inference (smoke test), Class 2 cascade resolution (Brief 2), minimal ambiguous brief (Brief 3), corpus-silent topic (Brief 4). It surfaced and closed five distinct issues across nine prompt iterations. For commercial readiness, the 21-case test set must run with consistent results across multiple model versions and prompt iterations — that scaling is the Iteration Plan's responsibility. Planned automated evaluation pipeline (Develop iteration, post-capstone): Binary contracts: automated script parses each output, validates JSON schema, cross-references citations against corpus YAML, checks for markup leakage, verifies required fields. Pass/fail per case, aggregated to a run-level summary. Judgment criteria: LLM-as-judge call with rubric and reference output — already operational in the artifact; capstone-level integration validated on four briefs. Human review on 20% of cases: calibrates the LLM-as-judge against drift. Human-LLM agreement on the four-brief set tracked at 100% (no disagreements between Nathalia's review and judge scores). | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | Nine edge cases identified during Develop — observed across the smoke test and four-brief extended coverage. Edge 1 — Audience-segmentation inference from brief. Smoke test produced 5 units instead of the expected 4 because the model inferred a segmentation gate from a donor-targeted brief. This was a positive surprise, not a failure — but it broke the demo plan's scripted arc. Implication: the demo recording's narrative arc had to be rewritten against actual output. Severity: medium — output quality is high, but demo coordination needs flexibility. Edge 2 — Regeneration citation drop. Structural citations (§7.4 chain, §7.5 brand chain) dropped during regeneration of Unit 2.1-D in the smoke test. Implication: regeneration prompt needed explicit preservation logic. Severity: high — corrupts the audit trail that's the Creator's core value prop. Fixed in Iteration 5. Edge 3 — Two pushbacks on the same rule with different concerns. Smoke test exercised §3.1 twice — once for content expansion, once for redundancy. Both correctly classified as overridable with different alternatives. Implication: the refine schema generalizes correctly across pushback concerns. Severity: low — actually validates the architecture. Edge 4 — Class 2 brainstorm not firing when brief is specific. The smoke test's demo plan assumed the §4.8 path-selection brainstorm would fire. It did not, because the smoke-test brief specified the path (donors → newsletter) directly. Brief 2 cascade testing later triggered §7.2 PAD brainstorm correctly when the brief warranted it. Implication: Class 2 firing depends on brief specificity, which is correct corpus behavior but unpredictable for demo choreography. Severity: low — correct behavior; demo plan adapts. Edge 5 — Hardcoded scaffolding in the master prompt biased the output. Original prompt said "produce exactly 4 units" and "Expected ~10 citations on Unit 2.2." Both forced the demo arc but suppressed real corpus reasoning. Implication: prompt should describe constraints in corpus-grounded language, not in demo-arc language. Fixed in Iterations 1 and 2. Severity: medium — easy to miss; only visible after smoke testing. Edge 6 — Total session latency. Smoke test ran 7:28 from brief input to diagram export. Four-brief testing showed faster total times when no clarification or correction fires (~25–35s) and longer times in fully exercised paths (~60–90s including judge + correction rounds). Implication: demo recording needs aggressive compression in edit (loading states cut, uneven playback speed). For production, latency is acceptable. Severity: low for production; medium for demo recording timing. Edge 7 — Type_tag/body incoherence — ADDED FROM FOUR-BRIEF TESTING. Creator generated a unit tagged type_tag: "modifies" with a body that read as standard new content with no modification rationale. The LLM-as-judge surfaced the issue but mapped it under the wrong criterion (Audience-Context Coherence rather than Existing-State Tagging Accuracy, which is objective). Implication: prompt needed explicit coherence enforcement between type_tag and body content; judge needed scope discipline. Fixed in Iterations 6 (Creator side) and 7 (judge side). Severity: medium — structural coherence is foundational to audit trail integrity. Edge 8 — Judge over-reaching into objective territory — ADDED FROM FOUR-BRIEF TESTING. The LLM-as-judge was stretching subjective criteria to cover objective concerns that don't fit. Specifically, it flagged Audience-Context Coherence on a unit that genuinely passed audience fit, but pulled the score to fail based on a type_tag concern. Implication: the judge architecture has two layers (subjective LLM + objective deterministic validators planned for Phase 2); without explicit scope discipline, the judge tries to do both layers' jobs and produces false positives in the subjective layer. Fixed in Iteration 7. Severity: medium — false positives erode trust in the eval system. Edge 9 — Mermaid escape compatibility — ADDED FROM FOUR-BRIEF TESTING. Brief 4 voice-memo flow produced a Mermaid diagram with raw "..." escape patterns that Mermaid v10 cannot parse — Mermaid does not interpret backslash escapes. The original HANDOFF instruction told the model to use \" explicitly, which was wrong by Mermaid's rules. Implication: any brief producing units with quoted sample copy in body would fail at the export stage. Fixed in Iteration 9. Severity: high — the export is the deliverable; broken renders break the value proposition. | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Nine concrete adjustments, each tied to an edge case from the Edge Case Identification cell. Adjustment 1 (Edge 5) — Removed unit-count scaffolding. Master prompt now says: "Unit count and identity MUST derive from the brief and existing-state components in scope. Produce one unit per coherent piece of the spec." Outcome: the model produces however many units the corpus + brief warrant. Adjustment 2 (Edge 5) — Removed citation-count scaffolding. Master prompt now says: "Cite every corpus rule that meaningfully applies to each unit's content, role, and context. Citation counts vary by unit — do NOT target a specific citation count; let the corpus speak for each unit." Outcome: citation counts range from 2 to 8 per unit, all corpus-justified across the four-brief set. Adjustment 3 (Edge 2) — Added CITATIONS DISCIPLINE and BODY DISCIPLINE to regeneration. The regeneration prompt now mandates citation preservation: "PRESERVE every citation from the previous version that still applies to the regenerated unit. Chain rules (§7.X), structural rules (§4.X), and platform rules (§6.X) typically PERSIST through regenerations — they govern the unit's role in the flow, not just its visible copy." Outcome: citation preservation validated across four-brief test runs; no regen-time citation drops observed. Adjustment 4 (architectural) — Generalized REFINE schema. Previous schema example hardcoded unit_id: 2.2 and rule_ref: §6.1. Now: "Identify which corpus rule the producer is pushing back against (from the unit's citations), then produce the REFINE convention_pushback JSON." Outcome: smoke test correctly identified §3.1 as the pushback target on Unit 2.1-D. Adjustment 5 (Edge 1) — Demo plan rewritten against actual output. Not a prompt change — a downstream artifact change. The 4-minute demo narration was rewritten to match the 5-unit donor-segmented output rather than the original 4-unit scripted arc. Adjustment 6 (Edge 7) — Added TYPE_TAG / BODY COHERENCE section to master prompt. Path 1 fix. Master prompt now defines what body content each type_tag value requires, with explicit verification clause: "If you find yourself drafting a body that doesn't match the type_tag you chose, change the type_tag — do not force the body to match a label that doesn't fit." Outcome: type_tag values coherent with body content across all four-brief runs. Adjustment 7 (Edge 8) — Added SCOPE DISCIPLINE to LLM-as-judge prompt. Path 3 fix. Judge prompt now explicitly names the 8 objective criteria and instructs to ignore them entirely. Added calibration rule: "When uncertain about a subjective criterion: default to pass. Fail ONLY when the violation clearly matches a reference fail example. 'Mild ambiguity,' 'soft concerns,' or 'could be stronger' are NOT fail signals." Outcome: false-positive rate dropped substantially in subsequent test runs. Adjustment 8 (Edge 4 + new feature) — Added WHEN TO POPULATE clarifying_questions section to master prompt. Option B fix. Prompt now provides explicit decision rules with three worked examples; new awaiting: "producer_clarification" state blocks generation until producer answers; dedicated CLARIFY screen added to artifact UI between BRIEF and CONFIRM (stepper went from 4 steps to 5). Outcome: Brief 3 minimal triggered 3 corpus-grounded questions; Brief 4 voice memo triggered 1 focused question. Adjustment 9 (Edge 9) — Mermaid escape correction. HANDOFF section of master prompt updated: "MUST use " HTML entity for embedded double quotes. NEVER use raw " or backslash-escaped \" — Mermaid does not interpret backslash escapes." Reinforced " " as mandatory for line breaks inside node labels. Added worked example with embedded sample copy showing the " pattern. Outcome: Brief 4 voice memo's 4-unit flow diagram renders cleanly across all units with quoted sample copy. Tracking and recording. All adjustments committed to artifact source (creator-studio-live.jsx) and standalone master prompt file (master-prompt.md). Each adjustment marked with an in-line comment explaining the change rationale. Smoke test screenshots and four-brief extended testing captures archived with timestamps as evidence. Adjustments NOT yet made — known gaps for next iteration: - The objective validator layer (Phase 2) — currently the eight objective criteria are enforced at prompt time and verified during manual review. Phase 2 builds a deterministic validator layer that runs alongside the LLM-as-judge, surfacing objective failures the way the judge surfaces subjective ones. - Adversarial HiTL bypass (planned negative case) has not been exercised on the four-brief set. - The 21-case test set runs (full Design-phase Example Cases + Convention Push-back cases) have not been automated. These are scoped to the post-capstone Develop iteration cycle. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | Hybrid evaluation pipeline. Three layers, each suited to a different criterion class. Layer 2 (LLM-as-judge) moved from designed to operational during Develop iteration — integrated in the artifact, running on every generation, contributing scores that drive auto-correction routing. Layer 1 — Automated script for binary contracts (DESIGNED, manual review in capstone). Eight binary contracts from the Output Evaluation Checklist are mechanically checkable. A Python script consumes the model's output, validates against the stage's JSON schema, cross-references every citation against the corpus YAML, checks the body for platform markup patterns, and verifies required fields per stage. Output: pass/fail per contract per case, aggregated to a run-level summary. Implementation cost: ~6 hours. Latency per case: sub-second. Maintenance: the script updates when the master prompt's output schemas change. For the capstone, this layer runs as manual review by Nathalia against the four-brief outputs (Manual Review cell documents the results); full automation is scoped to post-capstone Develop iteration. Layer 2 — LLM-as-judge for judgment criteria (OPERATIONAL in artifact). The Creator's React artifact integrates a second Sonnet 4 API call after every generation. The judge receives the brief + the Creator output (alignment + units), evaluates against the 7 subjective criteria with reference pass/fail examples per criterion, and returns nested scores: 4 session-level criteria + 3 unit-level criteria per unit. The judge prompt enforces SCOPE DISCIPLINE (Iteration 7) — evaluates only the 7 subjective criteria; objective concerns are explicitly out of scope. The judge prompt also enforces calibration rules — default to pass on uncertainty, fail only when violation clearly matches a reference fail example. The judge's output drives the auto-correction loop in the artifact: if any unit fails a unit-level criterion, the Creator regenerates that specific unit with the judge's feedback as input. The loop runs up to 2 rounds before surfacing the residual concern to the producer. Session-level criteria failures surface directly to the producer in the Quality Check banner without auto-correction (those require producer judgment to resolve — they signal Creator misread brief, decomposed wrong, or genuinely corpus-silent territory). Implementation: ~10 hours including rubric authoring per criterion. Latency per evaluation: 10–15 seconds. Across the four-brief test set, Layer 2 fired on every generation. Auto-correction rounds fired zero times because initial generation cleared the rubric on every brief (the prompt iterations closed the gaps that would have triggered correction). The loop exists as a quality assurance layer for cases the prompt hasn't yet been refined against. Layer 3 — Human review on a 20% sample (Nathalia in capstone, sampled human review post-capstone). Random 20% of cases reviewed manually against the same rubric the LLM-as-judge uses. Human scores compared to LLM scores; calibration drift surfaces when human-LLM agreement drops below 85%. Human review is the ground truth for the LLM-as-judge's calibration; if the LLM drifts, the LLM gets re-prompted. For the capstone, Nathalia reviewed all four briefs against the rubric directly — human-LLM agreement tracked at 100% (no disagreements between her review and judge scores). Post-capstone implementation: ~2 hours per 20% sample of a 21-case run. Cadence: every full test-set run during active prompt iteration; weekly during stable operation. Scaling to large test sets. The 21-case capstone test set scales to ~50 cases (commercial-readiness target) without architectural change — the same three-layer pipeline runs in parallel batches. Beyond 50 cases (mature engagement, multiple flow areas in corpus), the script and LLM-as-judge layers parallelize across machines; the human review sample size grows proportionally but its percentage drops to 10% as the LLM-as-judge calibration matures. Why hybrid rather than single-layer. Each layer catches what the others miss. Layer 1 (automated script) catches mechanical violations the LLM might miss in volume. Layer 2 (LLM-as-judge) catches qualitative issues a script can't reason about. Layer 3 (human review) calibrates the LLM against drift. The integrated architecture is the same one-corpus-multiple-AI-readers pattern that the broader product vision (Creator + QA gate + Auditor + runtime drift detection) extends — the eval pipeline itself demonstrates the architectural pattern in operation. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | Five cadences, each tied to a trigger. During active prompt iteration (Develop phase, post-capstone). Full 21-case test set runs nightly via automated pipeline. Trigger: prompt iteration is happening daily; nightly runs catch regressions within 24 hours of introduction. Output: a daily test report comparing scores against the previous night's baseline. Triage: criterion-level regressions get logged as adjustment candidates; full failures block the next prompt iteration cycle. On every prompt change (commercial readiness). Full regression run before any prompt change is committed to production. Trigger: prompt change PR. Output: pass/fail per criterion; the PR can't merge if any binary contract drops below 100% or any judgment criterion drops below its floor. Why: prompt changes that look like surface improvements sometimes introduce regressions on edge cases the author didn't think about. The regression suite is the safety net. On Anthropic model version pin change. Full regression run when the model version pin changes (e.g., from claude-sonnet-4-20250514 to a later Sonnet 4 patch). Trigger: Anthropic announces a model update. Output: full criterion scores plus a delta report against the previous version. Why: model behavior can shift on point releases; the regression run quantifies the shift before the upgrade ships to clients. Post-launch monitoring (commercial deployment). Weekly per-engagement: sampled production sessions (anonymized) replayed through the test pipeline, scored, surfaced to the consulting practice as a corpus health signal. Monthly aggregate across engagements: cross-client patterns surfaced (rules with high override rates, rules with high flag rates, latency drift). Quarterly: full test-set regression run with a sampled production case substituted for one of the synthetic cases, ensuring the test set drifts toward production reality. On corpus expansion or rule change (per engagement). When a client's corpus is updated (new rule, changed envelope, modified judgment_call options), the affected test cases re-run before the corpus version is promoted to production. Trigger: corpus YAML diff. Output: per-rule impact score (how many test cases changed behavior). Why: a corpus rule change shouldn't accidentally break unrelated test cases. This run catches collateral damage. For the capstone specifically. Only one frequency is exercised: smoke test on-demand. The other four cadences are designed but operationally exist only in the Iteration Plan. Commercial deployment requires standing up the nightly + on-change + on-model-update + post-launch monitoring cadences as paid operational work (~30 hours to set up the pipeline plus ongoing monitoring time). | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | For the capstone, the deployment infrastructure is built, tested end-to-end, and documented. The Creator runs as a Next.js 14 app on Vercel with two serverless API routes: /api/anthropic, a server-side proxy that holds the Anthropic API key as an encrypted environment variable and forwards each /v1/messages call, and /api/auth, which validates a shared access password and sets an httpOnly cookie the proxy re-checks on every request. The key never reaches the browser — that boundary was a deliberate security decision, not an afterthought. Live deployment was verified by running all four reference briefs through the production URL (brief → clarify/confirm → generation → LLM-as-judge → export) and confirming each stage against the Anthropic Console usage log. APIs: single dependency — Anthropic /v1/messages, model pinned to claude-sonnet-4-6. Databases: none — no DB, no vector store, no RAG; the corpus is inlined into the system prompt at call time, so there is no data layer to test. Rate limits: the production account runs at a low usage tier (5 requests/min, 10K input tokens/min); since one session fires 5–8 sequential calls, a single concurrent user is within tolerance but the limit is a known scale constraint (addressed in Scale Readiness). A monthly spend cap is set in the Console as a cost backstop. Monitoring: Vercel deployment logs plus Anthropic Console usage logs; client-side failures surface to the user in an in-app error banner with the raw API error. Rollback: Vercel retains full deployment history, so reverting to any prior build is a one-click promote; the model string and prompt are version-controlled in the source so a bad prompt change can be reverted and redeployed in minutes. For commercial deployment this hardens further: server-side corpus storage per engagement, a managed Anthropic account with higher rate-limit tiers, and formal uptime/error monitoring — designed, not yet built. | |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | At capstone stage the practice is a single operator (founder-led, pre-first-hire), so "internal teams" is forward-looking. What exists today is the documentation layer: this PRD, the deployment guide (account setup → key handling → deploy → rollback), the master prompt (~200 lines) and LLM-as-judge prompt as standalone reference files, the corpus extraction methodology, and the four-brief evaluation record. That is sufficient for the capstone's audience (two instructor evaluators) and for a first solo-led pilot. For commercial readiness, the consulting-led model means "internal teams" are the consultants who run extraction engagements. Their training is the methodology IP itself — the extraction playbook, corpus schema, and prompt-iteration discipline documented in Develop. Support, comms, and legal are not separate departments at this scale; they are engagement-level responsibilities the lead consultant owns, with legal review of each client's data-handling terms built into the engagement contract. Honest gap: the methodology has been exercised on one anonymized engagement in one domain, so the training materials are proven once, not across the maturity/industry range a scaled practice would need — naming this matches the scaling risk flagged in Develop. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | Two distinct launches, because this product has two layers. Capstone launch (now): a gated single-instance deployment. The live Creator at the production URL is access-controlled by a shared password, scoped deliberately to the two instructor evaluators — not a public release. This is the right shape for a proof-of-concept: it proves the live architecture works without exposing an unmetered API to the open internet. Commercial launch (designed): explicitly pilot-first, never all-users, which follows directly from the consulting-led business model. The product cannot be self-serve because the corpus — the asset that makes generation convention-compliant — only exists after an expert-led extraction engagement. So the rollout is: first commercial pilot with one client team (Year 1 target 3–5 engagements), each engagement being its own scoped "launch" — extraction → corpus → Creator deployment calibrated to that client → producer onboarding. AB testing is not the model; the meaningful comparison is producer-with-Creator vs the current-state review-bottleneck baseline, measured per engagement. Scaling to 10–20 engagements/year (Year 2–3) gates on hiring senior consultants who can lead extraction. Pipeline shape behind the 3–5 clients Year-1 target. Prospect profile (first cohort): mid-market or mission-driven organizations shipping conversational AI on messaging channels (WhatsApp/SMS/chat) that have (a) enough existing flows and review history to extract a corpus from, (b) acute review-bottleneck or convention-drift pain, and (c) a named budget-holding buyer (design lead, CX/conversational-AI lead, or product lead). The first cohort is deliberately drawn adjacent to the case-study domain — multilingual, messaging-heavy, civic/advocacy or mission-driven mid-market — where the methodology transfers cleanly and the founder has credible domain fluency. Channel: not a SaaS funnel — warm network and referrals, the live demo plus a written before/after case study as the credibility artifact, targeted outreach into the forming conversation-design community (Conversation Design Institute circles, design-ops leaders), and category-defining content on the "conversational design layer." Conversion path: live demo + case study → discovery call diagnosing the convention/review pain → paid, fixed-scope pilot (single flow domain, lower-end fee) measured against the pre-committed pilot criteria → reference case study → retainer and expansion. A low-friction wedge — a paid "convention audit" or "extraction sprint" — can land-and-expand ahead of a full Creator deployment. Landing 3–5 reference clients inside the 12–24-month window is precisely what converts the Discovery defensibility argument from theory to fact. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | The capstone deliberately avoids premature scale infrastructure, and the honest constraint is the API tier: at 5 requests/min and a 5–8-call session, a handful of concurrent producers would hit rate limits immediately. That's acceptable for evaluation and a single pilot, and the scale path is well-defined rather than hand-waved. Compute/API scale: Vercel serverless functions scale horizontally on their own; the binding constraint is the Anthropic rate-limit tier, which raises automatically as credit purchases cross thresholds ($5 → Tier 1, $40 → Tier 2, $200 → Tier 3), with custom limits available above that. Initial volume is monitored through the Console usage dashboard (calls, tokens, spend) and Vercel function logs. Corpus scale: the no-RAG decision holds up to ~40–50 rules per engagement; beyond that, the documented path is selective per-rule retrieval against engagement-partitioned corpora (hybrid keyword-on-short_ref + semantic search on rule descriptions). This is specified in Develop, not built. Practice scale: the real scaling unit isn't servers — it's senior consultants who can run extraction. Operating leverage comes from the reusable methodology and the portable corpus schema, but throughput is people-bound, which is why the revenue model tops out at 20+ engagements/year only with hiring. Extraction-hour ceiling and the transferability test. The methodology stays commercially viable only if extraction fits a bounded, repeatable budget: the target ceiling is ~5–10 consultant-days (~40–80 hours) to produce a 40–50-rule corpus, so that at the entry fee ($50K) extraction labor remains a defensible fraction of revenue and leaves margin for deployment plus retainer. Above that ceiling, the scope is too broad (split into phases) or the client's source material is too thin (re-scope before signing). How engagement one tells us whether this scales beyond the founder: engagement one is instrumented — actual extraction hours are logged by activity (interviews, artifact mining, drafting, validation), and each step is tagged "documented procedure" vs "relied on tacit judgment." The size of the tacit-judgment residue is the trainability signal: a methodology that is mostly documented procedure can be handed to a trained consultant, one that is mostly tacit judgment stays founder-bound. Honestly, n=1 run by the founder alone cannot prove transferability — but it produces the playbook, the hour-budget baseline, and the tacit-vs-documented gap analysis that make engagement two the real test, run by a first hire against the same hour budget and quality bar. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | The capstone's headline asset is the live demo itself — a working, shareable Creator that generates corpus-grounded design with citations in real time, which (per Discovery-phase feedback) lands harder than static screens. Supporting it: the 4-minute exec walkthrough video, the five generated demo slides, and the PRD as the full product narrative. For commercial go-to-market, the asset set follows the consulting sales motion, not a SaaS funnel: a methodology one-pager (how extraction produces a corpus), a before/after case study from the first engagement (the reference-momentum defensibility lever named in Discovery), a producer onboarding guide and in-product FAQ for end users at each engagement, and a buyer-facing ROI narrative tying the Creator to CSAT/escalation/velocity outcomes. The category-education burden is real (the "conversational design layer" isn't a named market yet), so early assets carry a heavier explain-the-category load than a mature-market product would. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | At capstone scale this is lightweight: the PRD is the canonical plan-of-record, and progress/outcomes are tracked through the dated smoke-test and four-brief evaluation captures plus the prompt-iteration change log (nine iterations, each committed with rationale). That iteration trail is itself the internal-comms artifact — it shows what changed, why, and what it fixed. For commercial operation, internal comms run per engagement: a kickoff aligning the buyer and producer team on scope, weekly progress against the extraction → deployment milestones, and an engagement retrospective surfacing corpus-health signals (override/flag rates) back to both the client and the practice. Cross-engagement learnings feed a shared methodology log so each new engagement starts ahead of the last — the mechanism that keeps the practice improving rather than re-solving. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | The capstone has a deliberately small data surface, which simplifies privacy. No end-user PII is collected or stored — the only inputs are producer-authored design briefs, and the case study uses a fully fictional client (Unitas) standing in for the real anonymized engagement, so no actual member or customer data touches the system. There is no database; briefs are transmitted to the Anthropic API for the duration of a request and not persisted. Secrets are handled correctly: the API key lives only as an encrypted server-side environment variable (never in client code or the repo, which .gitignore enforces), access is password-gated, and the key was rotated when it was exposed — demonstrating the revoke-and-replace discipline in practice. For commercial deployment the privacy posture scales up deliberately, because client corpora and briefs may contain proprietary conversational reasoning: server-side corpus storage in per-engagement secured partitions, Anthropic's zero-data-retention / enterprise terms evaluated per client compliance posture, and region-appropriate data-protection compliance — LGPD (Brazil) and GDPR (EU) named explicitly given the multilingual, cross-border nature of likely clients. Data-handling terms are negotiated into each engagement contract rather than assumed. Corpus ownership (designed contract terms). Because each corpus encodes the client's proprietary conversational reasoning, the engagement contract specifies ownership explicitly rather than leaving it implied: the client owns the content of their corpus (their conventions are their business knowledge), while the practice owns the schema, the extraction methodology, the master prompt, the Creator, and the eval pipeline. During the engagement and any active retainer, the client has full use of the corpus inside the deployed Creator. On exit, the client retains a static, human-readable export of their corpus artifact — their knowledge leaves with them — usable internally as documentation or onboarding, but ongoing corpus evolution and the Creator runtime require an active retainer. The methodology and tooling are never transferred. This posture is deliberately sellable to sophisticated buyers (they own their own reasoning) while preserving the moat: the corpus is hard to evolve or reproduce without the methodology, per the Discovery defensibility argument, so letting the client hold a static export costs the practice nothing structural. Final wording is confirmed with counsel per engagement. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | Moderation risk is structurally low because of what the product produces: internal design specifications for a producer to review, not runtime content shown to end users. The human-in-the-loop discipline is the primary control — nothing generates without sign-off, and every unit is reviewed before export — and Anthropic's own model-level safety sits underneath. The corpus carries a dedicated compliance rule category (LGPD-style, generic), so domain compliance is encoded into generation rather than bolted on. Audit is a built-in product property, not a separate process: every generated unit carries inline citations to the corpus rules that shaped it, making the AI's reasoning fully traceable — that audit trail is the Creator's core value proposition, and the citation-factuality criterion (100%, zero invented citations across four briefs) is what guarantees it. Legal at capstone scale is the anonymization discipline (the real engagement is never named). For commercial deployment, the honest gaps are: per-engagement legal review of client data terms, and the planned objective validator layer (Phase 2) that would deterministically enforce the eight objective criteria alongside the operational LLM-as-judge — designed, not yet built. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | User (producer) success: draft-accept rate without regeneration; time-to-handoff vs the current-state review-correction loop (which adds days per feature); regeneration rate per session as a signal of first-draft quality; and reduction in dependency on senior-reviewer availability — the bottleneck the product exists to remove. Business (buyer) value: the metrics the buyer is already accountable for, measured against a pre-Creator baseline — conversational-AI CSAT delta, escalation-rate change, deflection rate vs target, and brand-voice consistency. Upstream of those, velocity (conversational features shipped per cycle) and senior-IC time reclaimed for higher-value work. Commercially, the practice-level metric is engagement count and retainer renewal — first-customer case studies converting into the reference momentum that anchors the position before adjacent players (Acrolinx et al.) enter the layer within the 12–24-month window named in Discovery. Honest note matching the PRD's stance: these are the designed success metrics; none have post-launch baselines yet because the product is pre-first-commercial-pilot. Pre-committed pilot criteria (agreed with the client at kickoff, not measured retrospectively). Week 1 of the pilot captures the team's baseline — current median brief-to-handoff time and current senior-review touch rate — and the following targets are ratified with the buyer before generation begins, then assessed at pilot close (8–12 week window): - Draft-accept rate ≥ 65–70% of generated units approved without regeneration (vs the current-state baseline where producers rewrite ~50%+ of off-the-shelf AI drafts). - Time-to-handoff reduced ≥ 40–50% in median brief-to-approved-handoff vs the week-1 baseline. - Regeneration rate ≤ ~1.5 per unit on average, trending down across the pilot. - Senior-reviewer dependency: ≥ [X]% of drafts reach handoff without a senior-IC review pass (or senior-IC hours/feature down ≥ [Y]%). - Citation integrity 100% carried through to production as a hard quality gate. Business outcomes (CSAT delta, escalation rate, brand-voice consistency) are committed as directional trailing indicators over the quarter following deployment, not as pilot pass/fail gates — they don't move inside an 8–12 week window, so treating them as pilot gates would misrepresent the model. The discipline is that a pilot scored against thresholds set in advance produces a defensible case study; one scored after the fact produces a story. | ||
| AI Metrics | How will you measure AI performance and accuracy? | Through the three-layer evaluation pipeline already specified and partly operational. Layer 1 (objective, eight criteria, 100% threshold): structure compliance, citation factuality, brief coverage, existing-state tagging accuracy, body-format compliance, citation-count plausibility, two-mode presence, cascade resolution — mechanically checkable against the corpus YAML and brief; run as manual review at capstone scope, automated script planned. Layer 2 (subjective, seven criteria, floor thresholds 85–90%): the LLM-as-judge, operational in the deployed artifact, scoring brief-intent fidelity, clarifying-question calibration, corpus-silence interpretation, unit granularity (session-level) plus citation relevance, voice quality, audience-context coherence (unit-level), and driving the auto-correction loop. Layer 3 (human review on 20%): calibration ground-truth. Measured results across the four-brief set: all eight objective criteria at 100% after the Iteration-5 citation-preservation fix; all seven subjective criteria above floor when exercised (~92–95% averages), correctly marked not-exercised otherwise; human-LLM agreement 100%; auto-correction fired zero times because the nine prompt iterations had already closed the gaps. Honest coverage caveat (per Develop feedback): this is four briefs, not the 21-case set scoped for post-capstone — sufficient to prove the architecture discriminates Class 1 vs Class 2 across varied conditions, insufficient for commercial-grade confidence. | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | At capstone scale, support is founder-direct — single owner, single point of contact — appropriate for two evaluators and a first pilot. In-product, the immediate "support" surface is the error banner that shows the raw API error (which is how we diagnosed every issue during deployment — billing, model string, the clarify crash), so failures are legible rather than silent. For commercial deployment, support maps to the consulting engagement structure: the lead consultant on each engagement owns producer support and corpus questions, with clear escalation to corpus revision when a producer repeatedly overrides or flags a rule (those Override/Flag signals are themselves the support-and-improvement input). Ownership is unambiguous because each engagement has a named consultant accountable for outcomes — there is no diffuse support queue, which fits the high-ticket, low-volume model. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | The feedback-to-fix loop is already demonstrated, not theoretical: the nine prompt iterations each trace a specific observed failure → diagnosis → master-prompt or judge change → re-test (e.g., the regeneration citation-drop caught in smoke testing → Iteration 5; the Mermaid escape bug → Iteration 9). That edge-case-identification → adjustment discipline is the bug workflow. Prioritization follows severity as documented in the Edge Case table: issues that corrupt the audit trail or break the deliverable (citation drops, unparseable export) are HIGH and block release; behavioral surprises that are actually correct corpus reasoning (segmentation inference, Class 2 not firing on specific briefs) are logged as demo-coordination notes, not bugs. In production, the in-product HiTL signals (Apply / Override / Flag) are the structured feedback channel — Flag routes a suspected corpus gap to methodology-side review; high Override rates on a rule trigger a corpus-health investigation. Critical issues are communicated per engagement to the affected client; cross-client patterns feed the quarterly corpus review. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Operational: Vercel function logs (per-request status, latency) and the Anthropic Console usage log (model, tokens, spend per call) — the two surfaces we used live to confirm the deployment was hitting Sonnet 4.6 and to catch the billing/auth errors. Client-side exceptions surface in the UI error banner. A Console spend cap acts as both a cost control and an anomaly tripwire. AI-behavior: the LLM-as-judge runs on every generation in production, so quality scoring is continuous, not batch — the Quality Check banner surfaces session-level failures to the producer in real time, and unit-level failures trigger auto-correction before the producer ever sees them. That means AI-quality monitoring is in-line with usage rather than a separate post-hoc audit. Honest gap: there is no formal alerting/uptime monitoring or aggregated error dashboard yet — that's commercial-deployment work (~30h to stand up the nightly + on-change + on-model-update + post-launch monitoring cadences specified in the Evaluation Frequency cell). | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | Two compounding loops. Prompt/eval loop: during active iteration, the full 21-case test set runs nightly (post-capstone), every prompt change runs a full regression before merge (no merge if any objective contract drops below 100% or any subjective floor is breached), and a full regression runs on every Anthropic model-version pin change — so model drift is quantified before it reaches clients. Corpus loop: quarterly corpus-health reviews per engagement surface new conventions from production Apply/Override/Flag patterns and expand the corpus to new channels, languages, or sub-segments as the client's investment grows — offered as an optional retainer ($25K–$100K/year), which also turns continuous improvement into recurring revenue. The throughline is that improvement and the business model are the same mechanism: every engagement enriches the methodology, every production signal sharpens a corpus, and the eval pipeline is the safety net that lets the system change without silent regressions — the same one-corpus-multiple-AI-readers architecture the whole product is built on. | ||||




