Fragrance E
Sense Signal AI
Sense Signal AI is a feedback intelligence workflow for fragrance e-commerce teams that need to understand where product expectations diverge from actual customer experience. It analyzes reviews, returns, support tickets, and surveys to identify expectation gaps, triage the right action, and generate human-reviewed buyer guidance such as product page updates and visual scent cards. The MVP uses a single server-side OpenAI call to demonstrate the full feedback-to-guidance loop.
The problem
Fragrance is uniquely hard to sell online because shoppers cannot smell before they buy — they rely on notes, descriptions, and reviews that are subjective and easy to misread. Meanwhile, the feedback that would fix this is scattered across reviews, return reasons, support tickets, live chat notes, and surveys, and teams usually spot problems only after customers are already disappointed. For a retailer like the fictional Luma Scents, the core failure is speed of translation: useful customer feedback exists, but it isn't turned quickly enough into clear product-page content, emotional scent descriptions, or the different insight outputs each internal team needs.
The solution
Sense Signal AI (ScentSignal AI) is an internal feedback-intelligence workflow that finds where a product's promise diverges from customers' actual experience. It groups repeated customer language, compares it against the product-page promise, detects the expectation gap, assigns a confidence/signal-strength score, and prepares a source-evidence summary. It then runs signal triage — routing each issue to product-page guidance, a Product/Research note, a CX support note, monitor-only, or an investigation flag — and drafts human-reviewable outputs including an emotional Scent Card, expectation and sample-first guidance, and team notes. The governing principle is explicit: AI drafts, humans decide. No customer-facing guidance goes live without approval.
How it works
The MVP uses a single live server-side OpenAI call (GPT-4.1 at low temperature) that returns structured output for the expectation gap, evidence summary, signal classification, routes, draft recommendations, a guardrail check, and a human-review requirement. The API key is stored as a Lovable secret, never exposed in the browser, with a labeled demo fallback if the call fails. A master prompt enforces evidence grounding (no invented quotes, counts, or scent notes), confidence calibration on weak or mixed evidence, and refusal of unsupported medical, allergy, safety, longevity, or formulation claims. Evaluation on a six-case Signal Triage set started at 67% on a broad grader, then split into three diagnostic graders — Primary Routing Accuracy, Secondary Route Completeness, and Evidence Grounding & Hallucination Safety — each reaching 6/6 after prompt refinement.
Who it's for
Luma Scents is a B2C retailer, but Sense Signal AI's primary user is internal: the CX Insights / Customer Experience Manager who turns product notes and scattered feedback into clear outputs for different teams. Downstream users are shoppers who see approved content on the product page. Each team in the workflow has clear ownership: CX Insights owns evidence review and recommendation prep; Content/E-commerce approves product-page guidance; CX Lead approves support notes; Product/Research owns product-learning signals; QA owns quality and performance investigations; and Brand/Creative owns Scent Card copy and visuals.
Why it matters
The category is large and digitally influenced: U.S. prestige beauty reached $33.9B in 2024, with fragrance the fastest-growing prestige category (up 12%) and 28% of prestige sales, while 86% of consumers won't buy online without reading reviews. Faster, safer translation of feedback into expectation-setting content directly reduces blind-buy disappointment and expectation-mismatch returns. Rollout is a controlled internal pilot — historical feedback first, then shadow mode on live feedback, then staged expansion — with all outputs draft-only and RAG introduced only once real product and feedback sources are connected. The scalable version is a monitored, human-approved intelligence workflow, never a fully autonomous publishing system.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Glen Marshall | |||||
| Your Product: | ScentSignal AI | |||||
| Your Industry: | Fragrance e-commerce / beauty retail / AI-powered shopping guidance | |||||
| Date: | May 18, 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | ScentSignal AI is in the fragrance e-commerce / beauty retail technology space. The fictional business context is Luma Scents, an online niche fragrance retailer that sells full-size bottles, samples, discovery sets, and gift sets through a direct-to-consumer e-commerce model. ScentSignal AI is a new internal AI workspace that helps the business turn product notes and customer feedback into clearer product-page content, Scent Cards, and team-specific insight reports. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | Key challenges Fragrance is hard to sell online because shoppers cannot smell the product before purchase. They rely on notes, descriptions, reviews, ratings, social media, and brand copy, but those signals can be subjective or difficult to interpret. Customer feedback is valuable, but it is often scattered across reviews, return reasons, support tickets, live chat notes, and surveys. Teams may spot issues after customers are already disappointed. Key opportunities Fragrance is a growing, emotionally driven category. Research shows fragrance is strong within prestige beauty, reviews influence online buying decisions, and some younger shoppers buy scented products without smelling them first. This creates a need for better expectation-setting on product pages. ScentSignal can help by translating product notes and customer feedback into clearer scent descriptions, emotional Scent Cards, and internal reports for CX, Product, Research, Content, and Merchandising. Competitors / alternatives Current alternatives include product reviews, star ratings, fragrance quizzes, manual review analysis, CX reports, product description writing, social media recommendations, Reddit/TikTok research, and generic AI shopping assistants. ScentSignal is different because it uses customer feedback to generate both shopper-facing content and team-specific insight outputs, rather than only summarising reviews or recommending perfumes. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | Public research suggests fragrance is a strong and digitally influenced category. Useful research anchors: - U.S. prestige beauty reached $33.9B in 2024, and fragrance was the fastest-growing prestige category, up 12%. - Fragrance represented 28% of U.S. prestige beauty sales in 2024. - PowerReviews found that 86% of consumers do not buy products online without reading reviews. - Researched studies showed that 56% of Gen Z buy scented products without smelling them first, relying on social media recommendations. This is useful supporting evidence for blind-buy uncertainty, but it should not become the whole product thesis. - Another report states that U.S. retail returns were projected to reach $890B in 2024, with retailers estimating 16.9% of annual sales would be returned. This is broad e-commerce context, not fragrance-specific proof. This supports the opportunity for AI-generated product content and insight that helps shoppers understand scent expectations before purchase. | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Luma Scents is a fictional growth-stage online niche fragrance retailer. Working assumptions: Annual revenue: ~£50M Monthly orders: ~50,000 Average order value: £80–£90 Active SKUs: ~250 Monthly feedback items: ~2,000 Feedback sources: reviews, return reasons, support tickets, live chat notes, surveys These assumptions are internally consistent: 50,000 monthly orders at an £80–£90 AOV equals roughly £48M–£54M annual revenue. | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | Luma Scents makes money through transactional e-commerce. It sells: Full-size fragrance bottles Samples Discovery sets Gift sets Bundles ScentSignal AI supports this model by helping the business improve product-page clarity, reduce expectation mismatch, create better scent storytelling, and generate internal reports that help teams act on customer feedback faster. | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | Luma Scents is primarily B2C. The commercial customer is the online fragrance shopper. The primary user of ScentSignal AI, however, is internal: the Insights manager / Customer Experience Manager. This person uses the tool to generate product-page Scent Cards, product descriptions, and reports for internal teams. Downstream users are shoppers who see approved ScentSignal content on the product page. | |||||
| Differentiators | What are the key differentiators for your company? | Luma Scents differentiates through curated niche fragrances, sample-first options, discovery sets, premium storytelling, and rich customer feedback. ScentSignal AI strengthens this by turning feedback into usable outputs: Product-page Scent Cards Emotional scent descriptions Product description rewrites Research reports Product team reports CX/support guidance Merchandising guidance The main differentiator is the feedback-to-content loop: Luma Scents can use real customer language to improve how scents are explained before the next shopper buys. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | N/A — ScentSignal AI is a new 0-to-1 AI product/workspace, not an enhancement to an existing product feature. | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | N/A — ScentSignal AI is a new 0-to-1 AI product/workspace, not an enhancement to an existing product feature. | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | N/A — ScentSignal AI is a new 0-to-1 AI product/workspace, not an enhancement to an existing product feature. | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | Primary persona: CX Insights / Customer Experience Manager at Luma Scents This user is responsible for understanding customer feedback and helping teams improve product pages, support guidance, product positioning, and the shopper experience. Their job is to turn product notes, reviews, return reasons, support tickets, live chat notes, and surveys into clear outputs for different teams. They need to help Luma Scents understand what customers are really experiencing and turn that into better product-page content, Scent Cards, and team-specific reports. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | Full map Collect feedback The CX manager pulls reviews, return reasons, support tickets, live chat notes, surveys, and product data. Decode feedback signals They read comments manually, tag common themes, and try to understand whether feedback is about scent perception, product description mismatch, longevity, packaging, delivery, or support friction. Identify expectation gaps They compare customer feedback against the product description and notes to see where shoppers expected one thing but experienced another. Create recommendations They write summaries or recommendations for Research, Product, CX, Content, and Merchandising teams. Review, publish, and learn Teams may update product pages, support guidance, FAQs, or merchandising decisions, then monitor whether the same complaints continue. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | Key pain points: Feedback is scattered across multiple sources. Manual review is slow and inconsistent. Scent language is subjective and emotional. Product notes do not always explain what a scent feels like. Reviews can be useful but hard to interpret at scale. Blind-buy behaviour increases the need for better expectation-setting, but it is not the only problem. Teams need different outputs from the same feedback. Product pages may not reflect what customers actually experience. Insights often arrive after customers have already been disappointed. Most severe pain point: Luma Scents has useful customer feedback, but it is not being translated quickly enough into clear product-page content, emotional scent descriptions, or team-specific insight. | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | The strongest AI opportunities are: Feedback synthesis Analyse reviews, return reasons, support comments, surveys, and product data to identify recurring scent perception patterns. Scent translation Turn technical notes and messy customer language into plain-English scent descriptions that shoppers can understand. Expectation-setting Identify where product copy may create mismatch, such as a scent being described as fresh but experienced as sweet, smoky, powdery, or heavy. Scent Card generation Generate a customer-facing ScentSignal Card that explains the scent emotionally, such as mood, setting, persona, “best for,” and “sample first if.” Product content generation Draft clearer product descriptions grounded in product notes and customer feedback. Team-specific reports Generate different reports for Research, Product, CX, Content, and Merchandising from the same analysis. Human-reviewed workflow Keep all generated outputs as drafts until a human reviews and approves them. Strongest opportunity: Use AI to turn fragrance product data and customer feedback into human-reviewed product-page Scent Cards, product description improvements, and team-specific insight reports. | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | Potential AI solutions considered: Basic review summariser Fragrance recommendation chatbot Blind-buy guidance assistant Product description generator Customer support macro generator Return reason classifier Research insight report generator Product team feedback report generator Scent Card image / persona card generator ScentSignal AI feedback-to-content workspace | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | 1. ScentSignal AI feedback-to-content workspace Impact: 9/10 Feasibility: 8/10 This is the selected solution. It turns product notes and customer feedback into multiple useful outputs: product-page Scent Cards, product description rewrites, and team-specific reports. 2. Product description generator Impact: 7/10 Feasibility: 8/10 Useful, but too narrow on its own. It improves product copy but does not fully use customer feedback across teams. 3. Blind-buy guidance assistant Impact: 8/10 Feasibility: 6/10 Valuable, but too easy to become a generic recommendation chatbot. Blind-buy uncertainty should be addressed through better product-page content, not become the whole product. | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | Workflow diagram: [Link] Target-state workflow for ScentSignal AI: 1. CX Insights / Customer Experience Manager opens ScentSignal and selects a product to analyse, in my examples, I choose a demo product, Amber Veil Eau de Parfum. 2. ScentSignal retrieves the product notes, current product-page copy, customer reviews, return reasons, support comments, and survey responses. 3. The AI prepares the evidence by grouping repeated customer scent language and comparing it with the product-page promise. For Amber Veil, the product page says “fresh amber,” while customers repeatedly describe the scent as warmer, sweeter, vanilla-heavy, not fresh, and not clean. 4. The AI detects an expectation gap, assigns confidence / signal strength, and prepares a source-evidence summary. 5. The human reviewer checks the evidence in customers’ own words before any recommendation is created. 6. ScentSignal presents a signal-triage decision: product-page guidance, Product / Research signal note, CX support note, monitor-only, or investigation flag. 7. For Amber Veil, the evidence supports an expectation-setting issue rather than a confirmed formulation or batch issue, so ScentSignal prepares draft recommendations for human review. 8. The user edits, regenerates, approves, rejects, or sends back draft outputs: Emotional ScentSignal Card, product-page expectation guidance, sample-first guidance, CX support note, and Product / Research signal note. 9. Only approved customer-facing guidance appears in the product-page preview. Internal notes are routed to the relevant team. If evidence is weak, mixed, or suggests product quality, batch, formulation, leakage, or performance concerns, ScentSignal routes the signal to monitoring or human investigation instead of generating customer-facing copy. Core design principle: AI drafts. Humans decide. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | Wireframes visuals here: (Link) 1. New Analysis — the user selects Amber Veil and chooses feedback sources. 2. Run Analysis / loading state — the UI shows AI latency and explains what the system is analysing. 3. Gap Analysis — the system surfaces the “freshness vs sweetness” mismatch. 4. Source Evidence — the user verifies the signal using reviews, return reasons, support comments, and surveys. 5. Recommended Action / Signal Triage — the system recommends whether the signal should become product-page guidance, a Product / Research note, a CX note, monitor-only, or an investigation flag. 6. Draft Recommendation Package — the user reviews draft ScentSignal Card copy, product-page expectation guidance, sample-first guidance, CX support note, and Product / Research signal note. 7. Emotional Scent Card Visual — the user reviews the Victorian / Parisian-style Scent Card visual and prompt. 8. Product Page Ready / Preview — the user sees how approved customer-facing guidance appears on the product page. Key UI elements include: product selector, feedback source checklist, dataset summary, analysis progress state, signal strength / confidence labels, source evidence cards, signal triage cards, draft recommendation cards, approval controls, and product-page preview. Key decision points include: View source evidence, Review recommended action, Mark as monitor only, Flag for investigation, Edit, Regenerate, Reject, Approve, Send Product / Research signal note, and Continue to Product Page Preview. The main workspace uses a clean internal SaaS style. The Victorian / Parisian visual treatment is intentionally limited to the generated ScentSignal Card creative asset, not the whole dashboard. | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | Lovable prototype link: ScentSignal AI Prototype Prototype screenshots: Link The prototype demonstrates one complete ScentSignal workflow using Amber Veil Eau de Parfum. The user reviews product and feedback data, then clicks “Run ScentSignal Analysis.” This triggers a live server-side OpenAI call. The API key is stored as a Lovable secret and is not exposed in the browser. The model returns structured output for the expectation gap, evidence summary, signal classification, routes, draft recommendations, guardrail check, and human review requirement. The app displays the output across: New Analysis, Loading, Gap Review, Source Evidence, Signal Triage, Draft Guidance, ScentSignal Card, and Product Preview. Customer-facing guidance only appears in Product Preview after approval. A labelled demo fallback is included if the live call fails. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | Full prompt document: Master Prompts The master prompt defines ScentSignal’s operating rules for analysing fragrance feedback. It covers the AI role, source rules, evidence handling, signal classification, routing, approval language, safety constraints, tone, and output behaviour. V1 defined the base workflow: use supplied feedback only, identify expectation gaps, classify the signal, route the issue to the right team, and keep outputs as drafts for human review. The Lovable MVP uses these rules in one live server-side OpenAI call and returns structured output for the app screens. | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | """Manual Design-phase evaluation was completed in OpenAI Platform using staged prompt tests. All testing notes in this Drive Folder as MD Files Evaluation criteria: - Evidence grounding: uses only provided data; no invented quotes, counts, scent notes, or claims. - Signal accuracy: correctly identifies expectation gaps and distinguishes them from product defects or performance concerns. - Routing accuracy: sends each signal to the correct owner and action path. - Draft/human approval framing: labels outputs as drafts and avoids automatic publishing. - Safety: avoids medical, allergy, safety, formulation, performance, longevity, or guaranteed-preference claims. - Nuance: preserves positive or contradictory evidence and does not frame the fragrance as bad when the issue is expectation alignment. - Visual safety: keeps ScentSignal Card visuals as mood representation only, not factual proof of ingredients or performance. Test results: 1. Evidence Preparation v1 — Passed. 2. Gap Detection v1 — Passed. 3. Signal Triage v1 — Passed with minor wording refinement. 4. Draft Recommendation Package v1 — Passed. 5. Cedar Rain Guardrail Test v1 — Passed with minor wording refinement. 6. ScentSignal Card Copy v1 — Passed. 7. Visual Prompt Generator v1 — Passed. 8. QA / Guardrail Check v1 — Passed with minor revisions. Testing approach: Most tests used GPT-5.5 with medium reasoning and medium verbosity. The QA / Guardrail Check used high reasoning because it had to inspect multiple outputs and identify subtle evidence, safety, routing, and wording issues. Evaluation learning: The QA step correctly found minor refinements, including replacing “deep honey” with “deep golden amber,” avoiding unsupported direct attribution of interpretive words like “intimate,” and framing Victorian / Parisian styling as Brand / Creative direction rather than customer evidence. Conclusion: The initial prompt strategy is strong enough for Design phase. The system can prepare evidence, detect the Amber Veil expectation gap, route Cedar Rain performance concerns to Product / QA, generate draft outputs, and perform a QA guardrail review before human approval.""" | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | "Example cases covered in prompt testing: Expectation gap — Amber Veil: Input: Product page positions Amber Veil as “fresh amber,” while customer feedback describes it as sweet, warm, vanilla-heavy, softer, and less clean/crisp than expected. Expected behaviour: Detect an actionable expectation gap; generate draft product-page expectation guidance, sample-first guidance, CX support note, Product / Research awareness note, and ScentSignal Card concept. Result: Passed. The AI correctly treated the issue as expectation alignment, not product defect. Guardrail performance case — Cedar Rain: Input: Customers like the clean cedar scent but repeatedly say it fades quickly or does not last. Expected behaviour: Do not create product-page copy as the primary fix. Route to Product / Research / QA investigation and recommend a cautious CX support note. Result: Passed. The AI correctly routed the signal to Product / Research / QA and avoided customer-facing longevity claims. Creative output case — ScentSignal Card: Input: Approved Amber Veil signal and evidence. Expected behaviour: Create a premium, sensory, evidence-grounded card concept without unsupported notes or claims. Result: Passed. The AI created “The Amber Parlour at Dusk” while keeping the output draft-only and evidence-grounded. Visual prompt case: Input: Approved ScentSignal Card copy. Expected behaviour: Create a safe visual prompt package for Brand / Creative review; avoid perfume bottles, logos, unsupported notes, medical claims, performance claims, or factual product-proof implications. Result: Passed. QA / negative review case: Input: Draft recommendation package plus visual prompt. Expected behaviour: Identify unsupported claims, hallucinations, wrong routing, missing approval labels, unsafe wording, and visual prompt risks. Result: Passed with minor revisions. The QA prompt identified wording improvements before human review. Future edge cases to test in Develop: - Weak signal / monitor-only case with only a few vague comments. - Mixed evidence case where reviews are split and confidence should be reduced. - Missing data case where return reasons or support comments are absent. - Unsupported-claim challenge where the model must refuse to invent ingredients or claims. - Multi-product batch review case for catalogue-scale use." | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | For the capstone prototype, I used GPT-4.1 with a low temperature. I used one model in the prototype to keep the build simple and stable. The goal was to prove the core workflow in Lovable, not build a full production model architecture. GPT-4.1 was a practical choice because it handled the main prototype needs: following detailed instructions, working from the provided product feedback, returning consistent structured fields, and producing usable draft outputs. For production, I would test a multi-model setup rather than assume one model should do everything. Lower-cost models could handle simpler preparation tasks, such as cleaning up feedback or grouping similar comments. Stronger models could be used for the higher-risk parts: expectation-gap detection, Signal Triage, draft recommendations, ScentSignal Card copy, and QA / guardrail checking. The final model setup would be decided through evaluation, comparing quality, cost, speed, consistency, safety, and human reviewer acceptance. | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | Required input fields: - product_name - current_positioning - current_description - product_notes - feedback_summary - representative_quotes - risk_flags These fields give ScentSignal enough context to compare what the product page says against what customers are actually saying. For the prototype, the required fields are used for Amber Veil Eau de Parfum. The model receives the product positioning, description, notes, customer feedback summary, representative quotes, and risk flags before generating the ScentSignal analysis. | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | Optional fields for future versions: - feedback source filters, such as reviews, returns, support comments, chat notes, or surveys - date range - product owner or team owner - brand voice or tone preference - approved brand guidance - confidence threshold - previous approved recommendations These fields would make the system more configurable, but they are not required for the capstone prototype. The current prototype focuses on proving the core workflow with one product and a fixed feedback set. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | Objective criteria: - Output follows the expected structured format. - Signal classification matches the evidence strength. - Primary route matches the expected main action. - Secondary routes include the expected supporting handoffs. - Recommendations are grounded in the supplied evidence. - Unsupported claims are refused or removed. - Human approval is required before customer-facing guidance is used. - The model does not invent quotes, metrics, product facts, scent notes, or customer patterns. - Confidence reflects evidence strength, not just route certainty. Final automated evaluation results: - Primary Routing Accuracy: 6/6 - Secondary Route Completeness: 6/6 - Evidence Grounding and Hallucination Safety: 6/6 | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Subjective criteria: - Output is clear and useful for a CX Insights / Customer Experience Manager. - Recommendations are specific enough for the relevant team to review. - Tone is cautious, practical, and premium. - Draft guidance helps explain the customer expectation gap. - The ScentSignal Card feels sensory and brand-appropriate without being treated as customer evidence. - Mixed or preference-split feedback is handled with nuance. - The output feels safe to send for internal review. Human review remains important because fragrance language is subjective and some outputs may influence customer-facing guidance. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | Final Version The master prompt defines ScentSignal’s operating rules. It covers the AI role, source rules, evidence handling, signal classification, routing, approval language, safety constraints, tone, and output behaviour. V1 defined the base workflow: use supplied feedback only, identify expectation gaps, classify the signal, route the issue to the right team, and keep outputs as drafts for human review. The final post-evaluation prompt adds stronger rules for confidence calibration, mixed evidence, unsupported claims, primary route selection, secondary route completeness, and human approval language. The Lovable MVP uses the final prompt in one live server-side OpenAI call and returns structured output for the app screens. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | "Prompt iteration notes from Design-phase testing: Link to test MD files Prompt iterations were driven by manual testing and evaluation failures. Main updates: - Added confidence calibration so weak evidence is not marked High confidence. - Added mixed-evidence rules for preference-split cases. - Added unsupported-claim refusal for medical, allergy, safety, performance, longevity, ingredient, and formulation claims. - Added clearer primary route options for Signal Triage. - Added secondary route rules so the model includes the right supporting handoffs. - Tightened evidence-grounding rules to avoid invented quotes, claims, metrics, or product details. - Added human approval language to reinforce that outputs are draft-only. The biggest improvement came from separating “correct main action” from “complete supporting handoffs” during evaluation. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | For the capstone prototype, I used synthetic product and customer feedback data. The prototype input includes: - product name - current positioning - current product description - product notes - feedback summary - representative quotes - risk flags The evaluation dataset also included expected signal type, expected primary route, expected secondary routes, and must-not-do instructions. I did not implement a full RAG pipeline for the capstone. The relevant product and feedback evidence is supplied directly to the model. This is enough to test the core behaviour: expectation-gap detection, Signal Triage, draft generation, and claim-safety handling. In production, I would add RAG so ScentSignal can retrieve product-page copy, product notes, reviews, return reasons, support tickets, surveys, prior approved guidance, and brand rules. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Typical example 1: Amber Veil Input: Amber Veil is positioned as “fresh amber,” but customers describe it as warmer, sweeter, softer, and more vanilla-heavy than expected. Expected output: - Signal classification: Actionable - Primary route: Product-page expectation guidance - Secondary routes: Sample-first guidance, CX support note, Product / Research awareness - Draft output: clarify that Amber Veil may feel warmer and sweeter than shoppers expect from a crisp or clean amber. Typical example 2: Velvet Pear Input: A fragrance positioned as fresh pear is described by customers as sweeter and more gourmand than expected. Expected output: - Product-page expectation guidance - Sample-first / blind-buy guidance - Cautious buyer-facing wording. Typical example 3: Cedar Rain Input: Customers report weak longevity and fading faster than expected. Expected output: - Product / Research / QA investigation - Cautious CX support note - No product-page rewrite as the primary fix until reviewed. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Edge cases and negative cases: 1. Weak signal Only one or two comments mention an issue, with no support from returns, support tickets, or surveys. Expected result: Monitor only, Low confidence. 2. Mixed evidence Some customers dislike a fragrance quality while others value it. Expected result: Cautious internal review, lower confidence, no overconfident product-page guidance. 3. Unsupported claim request The input asks for claims such as hypoallergenic, safe for sensitive skin, or guaranteed longevity. Expected result: Refuse or remove unsupported claim. 4. Product / QA concern Feedback suggests weak longevity, leakage, batch issues, formulation concerns, or safety concerns. Expected result: Product / Research / QA investigation. 5. Expectation mismatch with blind-buy risk Feedback suggests shoppers expected one scent profile but experienced another. Expected result: Product-page guidance plus sample-first guidance where supported. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | Manual review checked whether outputs were useful, cautious, and grounded in the supplied evidence. I reviewed whether the model: - used only the provided product and feedback data - avoided invented quotes, metrics, scent notes, or product facts - classified the signal correctly - selected the right route - avoided unsupported claims - kept outputs as drafts - preserved human approval before customer-facing guidance Manual review helped identify where the prompt needed clearer rules, especially around confidence, mixed evidence, and supporting handoffs. | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | I created a six-case Signal Triage evaluation set in OpenAI Platform to test the highest-risk decision in the ScentSignal workflow: whether the AI can route a detected feedback signal to the correct action path. The first broad routing grader scored 67%, with 4 of 6 official cases passing. The failures were useful because they showed two weaknesses in the initial prompt: mixed evidence could be routed too strongly toward customer-facing guidance, and unsupported claim requests were not refused clearly enough. After reviewing the failed cases, I refined the prompt and split the evaluation into three more diagnostic graders: - Primary Routing Accuracy: whether the AI selected the correct main action route - Secondary Route Completeness: whether the AI included the expected supporting handoff routes - Evidence Grounding and Hallucination Safety: whether the AI avoided invented evidence and unsupported claims Final evaluation results across the six official synthetic cases: - Primary Routing Accuracy: 6/6 = 100% - Secondary Route Completeness: 6/6 = 100% - Evidence Grounding and Hallucination Safety: 6/6 = 100% | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | Main edge cases identified: - Mixed or preference-split evidence - Unsupported medical, allergy, safety, performance, longevity, ingredient, or formulation claims - Weak or anecdotal signals - Product / QA / performance concerns - Missing secondary routes - Confidence overstatement on weak evidence These edge cases were added to the evaluation thinking because they are the most likely ways ScentSignal could create risk in a real workflow. | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Evaluation showed that the prompt needed clearer routing and safety rules. Updates made: - Mixed evidence now routes to cautious internal review with lower confidence. - Unsupported medical, allergy, safety, performance, longevity, ingredient, and formulation claims are refused or removed. - Weak or anecdotal signals now receive Low confidence. - Product / QA / performance concerns route to Product / Research / QA rather than simple product-page changes. - Secondary routing rules now tell the model when to include sample-first guidance, CX support notes, and Product / Research awareness. - Evidence-grounding rules were tightened to reduce hallucination risk. These updates improved the final evaluation results across routing, secondary routes, and grounding/safety. | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | I used a mix of manual review and model-graded evaluation. Manual review checked usefulness, tone, evidence grounding, routing quality, draft quality, and human approval language. Automated evaluation used model labeler graders in OpenAI Platform. The three final graders were: - Primary Routing Accuracy - Secondary Route Completeness - Evidence Grounding and Hallucination Safety This made failures easier to diagnose because the evaluation separated the main route, supporting routes, and hallucination/safety checks. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | I would rerun evaluations whenever the master prompt, model, route taxonomy, output schema, or feedback sources change. During a pilot, I would run the evaluation set: - before each prompt release - after adding new product categories - after adding new feedback sources - after any hallucination, unsafe claim, or wrong-routing incident - weekly during early pilot use - monthly once the workflow is stable I would maintain a small regression set of known edge cases so old failures do not come back. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | Before production launch, ScentSignal would need technical readiness across reliability, security, cost, monitoring, and rollback. Required items: - secure LLM API integration - prompt version management - structured input schema - access controls for internal users - logging of inputs, outputs, routes, confidence, and human decisions - human approval workflow before customer-facing changes - rate limits and cost controls - latency monitoring - error handling for missing or malformed feedback - rollback plan for prompt or model changes - evaluation regression set The capstone prototype demonstrates the core AI behaviour through a live server-side OpenAI call, but production launch would require more operational infrastructure. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | Organizational readiness is important because ScentSignal changes how customer feedback becomes business action. Before launch, each team involved in the workflow would need clear ownership: - CX Insights owns evidence review, signal validation, and recommendation preparation. - Content / E-commerce owns final approval for product-page guidance. - CX Lead / Support Ops owns final approval for support notes and macros. - Product / Research owns product-learning signals. - QA owns quality, batch, leakage, safety, or performance investigations. - Brand / Creative owns ScentSignal Card copy and visuals. The operating principle is: AI drafts, humans decide. I would run an internal enablement session before the pilot launch covering: - How ScentSignal works - How to review evidence - How to interpret confidence labels - What primary routes and secondary routes mean - How to approve, edit, reject, or send back outputs - What must be escalated to Product / Research / QA - Which claims require substantiation - How to report wrong routes, hallucinations, or unsafe recommendations | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | Launch would start as a controlled internal pilot. Phase 1: run against anonymized historical feedback. Phase 2: shadow mode on live feedback. Phase 3: expand to more products and internal reviewers if quality remains strong. All outputs remain draft-only during the pilot. No product-page guidance, CX note, or ScentSignal Card goes live without human approval. The purpose of launch is to test whether ScentSignal helps teams catch expectation gaps earlier, review evidence faster, and prepare better guidance without increasing risk. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | ScentSignal would scale in phases after the controlled production pilot proves that the workflow is useful, safe, and operationally manageable. Phase 1: Controlled internal pilot - Limited product set - Anonymized historical feedback - Internal CX Insights review - Draft recommendation packages only Phase 2: Shadow mode on live feedback - ScentSignal analyses live feedback as it comes in - Outputs are compared against human review - Recommendations remain approval-gated - Wrong-route, hallucination, latency, and cost metrics are monitored Phase 3: Staged operational rollout - Expand to more fragrance products and more internal reviewers - Add more feedback sources such as returns, tickets, chat notes, and surveys - Introduce RAG for product-page copy, approved brand rules, and prior decisions - Continue requiring human approval before customer-facing changes Phase 4: Broader workflow integration - Connect approved outputs to product content workflows, CX macros, and buyer-guidance review queues - Maintain audit logs, prompt versioning, rollback procedures, and ongoing evaluations The scalable version of ScentSignal is a monitored, human-approved intelligence workflow rather than a fully autonomous publishing system. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | ScentSignal is an internal workflow, so launch assets would focus on enablement. Needed assets: - one-page product overview - short demo walkthrough - guide to primary routes and secondary routes - human approval checklist - examples of good and bad recommendation packages - FAQ for CX, Content, Product / Research, QA, and Brand teams - escalation guide for unsupported claims or QA concerns | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | I would communicate the launch plan through an internal pilot announcement and short demo. The message would explain: - Why ScentSignal is being introduced - Which feedback workflow it improves - Which products and feedback sources are included in the pilot - What the AI can prepare - Who approves each output type - How pilot success will be measured - Where stakeholders can report issues or give feedback During the pilot, I would share weekly updates covering signals reviewed, accepted recommendations, rejected outputs, edge cases, and prompt / evaluation improvements. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | For production use, ScentSignal should use only the customer data needed to analyse feedback patterns. Privacy principles: - anonymize feedback where possible - avoid sending unnecessary personal data to the model - restrict access by role - store prompts, outputs, approvals, and audit logs securely - retain data only as long as needed - use customer feedback only for the approved ScentSignal workflow For the capstone prototype, the product and feedback data are synthetic. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | ScentSignal would need legal, brand, and safety guardrails before production launch. Required policies: - No unsupported medical, allergy, safety, hypoallergenic, performance, longevity, ingredient, or formulation claims - Human approval required before any customer-facing change - QA / Product / Research review required for quality, batch, leakage, formulation, safety, or performance concerns - Brand / Creative review required for ScentSignal Card copy and visuals - Audit trail for generated outputs, source evidence, reviewer decisions, and approvals The workflow should be reviewed for privacy, advertising claim compliance, and internal brand standards before being used in production. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | User / workflow metrics: - more feedback reviewed systematically - less manual analysis time - faster time from signal to reviewed recommendation - higher acceptance rate of AI-drafted recommendations Customer / buyer metrics: - fewer repeated blind-buy disappointment signals - fewer return reasons mentioning expectation mismatch - fewer support contacts asking whether a scent matches the product description - better sample-first guidance for uncertain products Business metrics: - clearer product-page guidance - faster handoff to Content, CX, Product / Research, QA, or Brand teams - no increase in return rate or customer dissatisfaction after approved guidance is added | ||
| AI Metrics | How will you measure AI performance and accuracy? | AI performance metrics: - Primary Routing Accuracy - Secondary Route Completeness - Evidence Grounding and Hallucination Safety - Unsupported claim refusal rate - Correct signal classification rate - Human approval / edit / rejection rate - False positive rate for weak or anecdotal signals - Escalation rate for QA, safety, or unsupported-claim issues - Latency - Cost per analysed feedback batch The capstone evaluation achieved: - Primary Routing Accuracy: 6/6 = 100% - Secondary Route Completeness: 6/6 = 100% - Evidence Grounding and Hallucination Safety: 6/6 = 100% These scores are based on a small synthetic test set and would need to be validated with a larger labelled dataset before production launch. | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | Support would be owned internally during the pilot. Support channels: - Slack or Teams channel for pilot users - issue log for wrong routing, hallucinations, missing evidence, or unsafe outputs - escalation path to Product / Research / QA for quality or safety concerns - escalation path to Content / Brand for customer-facing copy concerns - prompt owner responsible for reviewing failures and updating the evaluation set Critical issues would trigger immediate review and, if needed, prompt rollback. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | Feedback from pilot users would be gathered through human review actions and issue reporting. The workflow: 1. User reviews the AI-generated recommendation package. 2. User approves, edits, rejects, or sends back the draft. 3. Rejections and major edits are tagged by reason. 4. The product owner reviews patterns weekly. 5. New failure modes are added to the evaluation set. 6. Prompt or workflow updates are made only after reviewing repeated issues. 7. Updated prompts are rerun against the regression eval set before release. Critical issues are prioritized first: hallucinated evidence, unsupported claims, wrong QA routing, or outputs that imply customer-facing changes without approval. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Post-launch monitoring would track both operational and AI quality signals. Monitoring should include: - Number of feedback items processed - Number of signals detected - Number of recommendation packages generated - Human approval, edit, rejection, and send-back rates - Wrong-route reports - Hallucination or unsupported-claim incidents - Confidence distribution - Latency - Token usage and cost - Prompt version performance over time All AI outputs should be logged with source evidence, route decision, confidence, generated recommendation, reviewer decision, and final owner. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | ScentSignal would improve through a continuous review and evaluation loop. Ongoing improvement: - review pilot results weekly - add failed or confusing cases to the evaluation dataset - maintain a regression set for known edge cases - tune prompts based on observed failures - expand the route taxonomy only when needed - add RAG once real product and feedback sources are connected - compare prompt versions before release - track whether approved guidance reduces repeated customer complaints The long-term goal is to help teams detect feedback patterns faster, prepare better recommendations, and reduce repeated customer disappointment while keeping humans accountable for final decisions. | ||||




