← All capstone projects

E-commerce

Outgrown AI

Built by Mayank Bagaria Cohort 9 Consumer resale / recommerce

Outgrown AI is a resale decision-support assistant for parents of young children deciding whether clothes should be sold, bundled, donated, or recycled. From a single photo, it analyzes inventory, recommends a resale strategy, explains the reasoning, and suggests alternative bundle options if the first strategy does not work. The product shifted from listing generation toward helping parents make lower-effort, higher-confidence resale decisions.

The problem

Parents cycling through children's clothing face decision fatigue over what to do with outgrown items. Individual kids' items often have low resale value, so deciding what's worth selling versus donating is a frequent, high-friction judgment call. Sorting items into effective bundles is time-consuming and mentally taxing — factoring in size, season, condition, and matching items — and creating marketplace listings adds titles, pricing, and descriptions on top. These pain points are severe because they directly determine whether parents attempt to resell at all, often defaulting to donating or discarding instead.

The solution

outgrown.ai is a resale decision-support assistant for parents of children ages 1–4. From a single photo, it analyzes an item or pile, recommends whether to Sell, Donate, Keep, or Bundle, explains the reasoning, and suggests alternative bundle options if the first strategy doesn't fit. The product deliberately shifted from being a listing generator toward reducing effort and decision fatigue. It prioritizes realistic effort-versus-value recommendations — including recommending donation when resale isn't worthwhile — with kids-specific intelligence around seasonality timing, same-size and same-season bundles, and complementary item suggestions. Confidence levels and editable outputs keep parents in control.

How it works

The workflow moves from photo upload (with optional size, brand, and notes) through AI processing to structured outputs: a recommendation badge, confidence level, estimated value range, seasonal timing tips, bundle suggestions, and a marketplace-ready listing. It uses a hybrid model approach — a multimodal model such as GPT-4o for image analysis (item type, condition, seasonality, bundle opportunities) and a lower-cost text model for recommendations and listing generation. Strong hallucination-avoidance rules prevent inventing brands, sizes, or conditions; unclear images trigger a request for a clearer photo rather than a guess, and out-of-domain images are rejected. Manual review across 14 representative scenarios drove the product's evolution, with final results at 13 of 14 scenarios (93%) meeting expected outcomes; similar-item counting on highly repetitive lots (e.g., a lot of identical shorts) remains a known limitation. RAG is not used in the MVP but is planned as a supplementary knowledge layer.

Who it's for

The product is B2C, for time-constrained, mobile-first parents of children ages 1–4 who frequently cycle through clothing and toys. They act as occasional sellers and household managers who want to recover value with minimal effort and reduce clutter. Monetization is a freemium plus Seasonal Cleanup Pass model, deliberately aligned to parent behavior: a free tier with 3 item assessments, then a ~$9.99 CAD 30-day pass for unlimited assessments, bundle optimization, listing generation, and saved history — matching seasonal cleanouts and growth spurts rather than monthly usage.

Why it matters

The kids' apparel market is growing steadily at ~5–7% CAGR, while the resale and circular-commerce segment grows much faster at ~9–16% CAGR, driven by cost-of-living pressure and sustainability trends. The biggest opportunity is reducing friction rather than building another marketplace. As a startup validating problem-solution fit, outgrown.ai positions itself as a decision-support and workflow layer that horizontal marketplaces like Facebook Marketplace, Poshmark, and eBay are unlikely to prioritize — focusing on children's-resale-specific workflows such as deciding what's worth selling and optimizing seasonal, same-size bundles for busy parents.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Mayank Bagaria
Your Product:outgrown.ai
Your Industry:eCommerce
Date:April 30, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?eCommerceMayank — you've found a real friction point, and you're articulating it well. The insight that parents default to donating or trashing because the effort-to-value ratio of reselling kids' items is broken that's a problem worth building around. Your pain points are specific and sequenced with genuine reasoning, not just listed. Two places to sharpen: Your ICP is "time-constrained parents of young children" but a parent of a 2-year-old cycling through onesies every 8 weeks and a parent of a 10-year-old with a closet of lightly-worn jackets have completely different inventory volumes, item values, and bundling logic. Pick one. Your AI prompt design, bundle recommendations, and pricing model all shift depending on which parent you're building for first. Your competitive landscape names Poshmark, Facebook Marketplace, and eBay, and the question worth sitting with is whether those platforms begin shipping their own AI listing tools is the right one to sit with. Horizontal platforms have distribution advantages; if they add automation, your defense needs to be something specific to kids' resale (seasonality logic, outgrown-item velocity, bundle optimization) that they won't prioritize. Spell that out before you move forward. As you move into Design, anchor your target workflow against a single, named parent persona age of kids, reselling frequency, platform they're currently frustrated with so your wireframes and prompt design solve for a real person, not an abstraction.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?The kids resale market is growing due to rising living costs and increased acceptance of secondhand goods, but is held back by low item value, high effort to list, and decision fatigue. The biggest opportunity is reducing friction and helping parents turn unused items into value, rather than building another marketplace. Key competitors include Facebook Marketplace, Poshmark, and eBay, which all require manual effort and lack automation.
What is the projected growth rate of your target market segment over the next 3-5 years?- The kids apparel market is growing steadily at ~5–7% CAGR over the next 3–5 years - The resale / circular commerce segment is growing much faster at ~9–16% CAGR - Growth is driven by cost-of-living pressures, sustainability trends, and increased adoption of resale platforms
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?At the startup stage, focused on validating the core problem-solution fit.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)* Monetization Model: Freemium + Seasonal Cleanup Pass * Core Idea: Align pricing to real parent behavior patterns (seasonal cleanouts and growth spurts) rather than ongoing monthly usage Free Tier * 3 free item assessments * Basic sell vs donate recommendations * Limited AI listing generation * Allows users to experience value before paying Seasonal Cleanup Pass * ~$9.99 CAD * 30-day access * Unlimited clothing assessments * AI bundle optimization * AI-generated marketplace listings * Seasonal resale timing recommendations * Saved assessment history Future Monetization Opportunities * Marketplace integrations * Auto-posting tools * Family inventory management * Premium resale insights * Multi-child wardrobe tracking
Who is your primary customer base (B2B, B2C, B2B2C)?B2C
DifferentiatorsWhat are the key differentiators for your company?KEY DIFFERENTIATORS: * Focused specifically on children’s clothing resale for parents of children ages 1–4 * AI recommendations prioritize reducing effort and decision fatigue, not just listing generation * Kids-specific resale intelligence: * seasonality timing recommendations * outgrown-item velocity * bundle optimization * complementary item suggestions * Practical “sell vs donate” recommendations to help parents avoid wasting time on low-value items * Parent-friendly bundle logic based on: * same size * same season * daycare/play use cases * Designed around the high inventory turnover common in early childhood clothing COMPETITIVE DEFENSE: * Platforms such as Facebook Marketplace, Poshmark, and eBay may eventually introduce generic AI listing tools * Large marketplaces optimize for: * broad product categories * marketplace liquidity * generic listing automation * This product focuses on children’s resale-specific workflows that horizontal marketplaces are unlikely to prioritize: * determining whether items are actually worth selling * optimizing same-size and seasonal bundles * reducing resale effort for busy parents * helping parents sell before children outgrow seasonal demand windows * The product acts as an AI-powered decision-support and workflow optimization layer rather than another resale marketplace * The primary value is not just generating listings, but helping parents make faster, smarter resale decisions with less effort
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?The primary customers (buyers) are parents, particularly time-constrained parents of young children who frequently purchase and cycle through clothing and toys. These users are motivated by convenience and cost savings, and are looking for efficient ways to manage and resell used and unused items.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Primary end-users are parents of children 1-4, particularly those who frequently purchase and cycle through clothing and toys. These users are both the core users and the most revenue-impacting segment, as they generate supply (items to resell) and drive transactions on the platform. Goals: Quickly decide what to do with used and unused items (sell, donate, or keep), recover value with minimal effort, and reduce clutter. Roles: Occasional sellers and household managers responsible for purchasing, organizing, and disposing of children’s items. Context: Time-constrained, mobile-first users managing high inventory turnover as children grow, often using platforms like Facebook Marketplace or Poshmark but experiencing friction in listing and selling.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?The core features of the product are focused on reducing the friction involved in reselling children’s clothing and toys. 1. AI Item Analysis: Users upload photos of items or bundles, and the system analyzes item type, condition, seasonality, and resale potential. This helps parents quickly determine whether items are worth selling, donating, or keeping. 2. Bundle Recommendations: The product suggests optimized bundles based on factors such as size, category, season, and complementary items (e.g., matching tops with pants). This addresses the common challenge of low individual item value and improves resale likelihood. 3. AI-Generated Listings: The system generates marketplace-ready titles, descriptions, and pricing suggestions, significantly reducing the time and effort required to create resale listings.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)External users - parents
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?1. Parent takes a photo of children’s clothing or toys they want to get rid of 2. AI analyzes the items and recommends whether to sell, donate, keep, or bundle 3. AI suggests ways to improve bundle value 4. User reviews and organizes the recommended bundles 5. AI generates a marketplace-ready title, description, and suggested price 6. Parent posts the listing to marketplace platforms 7. Parent saves time and potentially recovers value from items
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?- Parents experience friction when trying to determine which items are actually worth selling versus donating, especially because individual kids items often have low resale value. - Sorting items into effective bundles is time-consuming and mentally taxing, particularly when considering size, season, condition, and matching items. - Creating marketplace listings is another major pain point, requiring photos, pricing decisions, titles, and descriptions for each bundle. These pain points are both highly frequent and severe because it directly impacts whether parents attempt to resell items at all, often causing them to default to donate or dispose instead.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.1. Determining what items are worth selling vs donating (High severity, high frequency) Parents often experience decision fatigue when trying to decide whether low-value kids items are worth the effort to resell. Why GenAI fits: LLMs can analyze item context (type, condition, seasonality, bundle potential) and provide practical recommendations with reasoning. 2. Creating marketplace listings (High severity, high frequency) Writing titles, descriptions, and deciding on pricing takes time and effort, especially for bundles. Why GenAI fits: LLMs are well-suited for generating structured, marketplace-ready listing content quickly from minimal input. Figuring out how to bundle items effectively (Medium-high severity, high frequency) 3. Parents struggle to determine which items should be grouped together to maximize resale value and likelihood of sale. Why GenAI fits: LLMs can recommend bundle combinations and suggest complementary items based on size, category, season, and resale context. Knowing when to sell items based on seasonality (Medium severity, medium frequency) 4. Parents may list items at suboptimal times, reducing resale potential. Why GenAI fits: LLMs can use seasonal context to recommend better timing and explain why certain items may perform better closer to a specific season.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.1. Whats worth selling - sell vs donate recommendation - resale value estimator - resale confidence score 2. Marketplace listing - AI-generated titles/descriptions - pricing suggestions - AI-generated tags - One-click marketplace-ready listing format 3. AI Bundling - bundle recommendations - themed bundles (school, winter, toddler essentials) - quick-sell vs maximize-value bundle strategies
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.1. Sell vs Donate Recommendations - High impact and high feasibility) 2. Bundle Suggestions - High impact and medium feasibility 3. Listing Generator - Medium-high impact, high feasibility Select #1 but might be able to do all 3
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?1. Trigger: Parent has children’s clothing their child has outgrown 2. Open App: a. Parent takes photos of clothing/items they want to get rid of b. Parent optionally adds extra information (e.g., brand, size, notes) 3. AI Processing: a. AI identifies item type, condition, seasonality, and resale potential b. AI recommends whether items should be sold, donated, kept, or bundled c. AI suggests optimized bundles using similar sizes, seasons, or complementary items d. AI generates a marketplace-ready listing with: Title, Description, Suggested price range 4. User Action a. Parent reviews the recommendations b. Parent posts the listing to Marketplace c. Parent saves or dismisses the recommendation 5. Feedback / Learning a. Parent provides feedback on whether the recommendation was helpful b. Future state: - Parent marks items as sold, unsold, or donated - System learns from outcomes to improve future recommendations Outcome: Parent saves time, reduces clutter, and recovers value from outgrown children’s clothing with minimal effortMayank, the live prototype is the strongest asset in this capstone. The deployed app shows a functional item analysis flow with confidence scoring, value ranges, and disposition recommendations that map cleanly to the six-screen workflow in Design, and the fourteen-entry prompt iteration log with explicit reasoning behind each change shows unusually rigorous Develop work. The one reframe to carry forward is that the convergence section selects sell-versus-donate as the primary feature but the prototype and PRD both signal intent to ship all three (analysis, bundle, listing). Leading with the sell-versus-donate decision as the demoed hero flow, with bundle optimization and listing generation framed as supporting features, gives you one prompt, one eval set, and a tighter Demo Day story. The evaluation criteria are well categorized across objective and subjective dimensions, and setting a specific pass-rate number on structure compliance and a confidence-accuracy percentage on each criterion will give that framework real decision power at launch. The legal and privacy scaffolding in Deploy shows genuine care, with PIPEDA and GDPR awareness, encryption, consent flows, and a content moderation framework all in place. Tying that compliance work to a concrete launch-readiness checkpoint, where each policy has a verification step before go-live, turns it from documentation into an operational gate. The high-leverage modeling exercise to run next is connecting the dual-model architecture to a cost-per-analysis estimate, because that number determines whether unit economics hold for a product built around low-value resale items.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?SCREEN 1: HOME * Parent profile / recent activity summary * Prominent “Analyze Items” button (primary CTA) * Recent item assessments and bundle recommendations * Quick access to saved listings and past results * Navigation: * Home * History * Saved Listings * Settings SCREEN 2: ITEM CAPTURE * Camera/photo upload flow * Multi-photo support for bundles or piles of clothing * Optional fields: * Brand * Size * Notes * Photo preview grid before submission * Large tap targets for quick mobile use * “Continue to Analysis” button SCREEN 3: AI PROCESSING * AI analysis animation (instead of basic loading spinner) * Reassuring microcopy: * “Analyzing resale potential…” * “Optimizing bundle suggestions…” * Short intentional delay to build trust * Processing steps displayed: * Item identification * Bundle analysis * Resale evaluation SCREEN 4: AI RECOMMENDATIONS * Recommendation badge at top: * Sell * Donate * Keep * Bundle * Confidence level indicator * Estimated value range * Seasonal timing suggestions * Expandable “Why this works” explanation * Bundle improvement suggestions: * “Add matching pants” * “Combine same-size items” * Action buttons: * Continue to Listing * Donate Instead * Save for Later SCREEN 5: LISTING GENERATOR * Marketplace-ready title * AI-generated description * Suggested price range * Editable fields so users stay in control * Copy-to-marketplace button * Shortcut to Facebook Marketplace * “Mark as Listed” action SCREEN 6: HISTORY & FEEDBACK * Timeline of past item assessments * Saved bundle recommendations * Listing outcomes: * Sold * Unsold * Donated * “This helped / This didn’t help” feedback buttons * Filters: * Date * Recommendation type * Outcome * Future learning loop: * AI improves recommendations based on user outcomes
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?https://outgrown.lovable.app AI Prototype Demonstration The prototype will demonstrate the core AI-assisted resale workflow focused on reducing decision fatigue and effort for parents managing outgrown children’s clothing. The prototype will demonstrate: * Uploading photos of children’s clothing/items * Optional user inputs: * Size * Brand * Notes * AI analysis of: * Item type * Condition * Seasonality * Resale potential * AI recommendations: * Sell * Donate * Keep * Bundle * AI bundle suggestions: * Similar sizes * Matching/complementary items * AI-generated marketplace listing: * Title * Description * Suggested price range AI Inputs, Processing, and Outputs Inputs-- Users provide: * Photos of clothing/items * Optional metadata: * Brand * Size * Notes Visual Presentation * Camera/upload interface * Image preview grid * Optional form fields AI Processing-- The AI analyzes uploaded items and determines: * Resale likelihood * Bundle opportunities * Seasonal timing * Estimated value range Visual Presentation * AI processing/loading screen * Progress states such as: * “Analyzing items” * “Evaluating resale potential” * “Optimizing bundles” Outputs-- The AI displays: * Recommendation badge: * Sell * Donate * Keep * Bundle * Confidence level * Estimated value range * Bundle recommendations * Marketplace-ready listing preview Visual Presentation Outputs will be separated into clear sections: 1. Recommendation 2. Bundle Suggestions 3. Listing Preview 4. Next Best Action Recommendation cards and confidence indicators will be visually prioritized to make the results easy to scan and understand quickly. Essential Features for Launch (MVP) * Photo upload/capture * AI sell vs donate recommendations * AI bundle suggestions * AI-generated listing title and description * Assessment history feed * Editable AI-generated outputs ----- Features for Future Releases * Receipt scanning * Marketplace integrations * Automatic resale tracking * Personalized pricing optimization * Cross-platform listing support * AI learning from sold/unsold outcomes * Notifications for optimal seasonal resale timing * Advanced user profiles and inventory management
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?You are an AI resale assistant for time-constrained parents of children ages 1–4 who frequently cycle through outgrown clothing and toys. Your goal is to help parents quickly determine whether items are worth selling, donating, keeping, or bundling while minimizing effort and decision fatigue. Act like a practical resale advisor, not a salesperson. Prioritize realistic recommendations, time savings, and reducing unnecessary listing effort. Do not overestimate resale value or encourage users to sell low-value items individually when the effort is unlikely to be worthwhile. When analyzing items, consider: - Item type - Estimated size or age range, if visible or provided - Condition and visible wear/stains - Brand, only if clearly visible or provided - Seasonality and current time of year - Whether items are likely to sell better before an upcoming season - Bundle quality and resale likelihood - Typical resale behavior for young children’s clothing - Whether adding complementary items would improve bundle value - Effort required versus expected resale value Bundle optimization guidance: - Prioritize practical, parent-friendly bundles: - same size - same season - similar use case - Suggest complementary items when helpful: - matching pants with tops - jackets with winter accessories - daycare/play clothing sets - Recommend separating higher-value items when appropriate. - Explain why a bundle is likely or unlikely to sell. Seasonality guidance: - Recommend listing seasonal items before peak demand. - Example: - winter clothing in late summer/early fall - spring clothing before warmer weather - Mention timing recommendations when relevant. Hallucination avoidance rules: - Do not invent brands, sizes, conditions, or item details. - If information is unclear, acknowledge uncertainty. - Use confidence levels: - High - Medium - Low - If photos are unclear or insufficient, ask for clearer photos or additional details instead of guessing. Tone guidelines: - Helpful - Honest - Practical - Concise - Non-judgmental Always return output in this structure: === DECISION === Recommendation: Sell | Donate | Keep | Bundle Confidence: High | Medium | Low Estimated Value Range: $X–$Y CAD Why this works: Short explanation of the recommendation, including resale effort, seasonality, or bundle quality. Timing Tip: If relevant, explain the best time to list the items. --- === BUNDLE SUGGESTIONS === Explain: - How items should be grouped - Whether complementary items should be added - Whether any items should be separated into their own listing --- === READY-TO-POST LISTING === Title: Short marketplace-friendly title. Suggested Price: Suggested resale range. Description: Short resale description suitable for Facebook Marketplace or Poshmark. Tags: Relevant search-friendly tags. --- === NEXT BEST ACTION === Provide the simplest recommended next step for the parent.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?* Clarity: Recommendations and listings should be easy to understand and quickly scannable for busy parents. * Relevance: AI suggestions should match the uploaded items, including appropriate bundle recommendations, seasonality, and resale potential. * Accuracy: Item descriptions, bundle suggestions, and estimated value ranges should reasonably reflect the uploaded clothing/items and provided details. * Practicality: Recommendations should prioritize saving users time and effort, including recommending donation when resale is unlikely to be worthwhile. * Tone: Outputs should feel helpful, practical, and trustworthy — not overly sales-focused or robotic. * Consistency: AI outputs should follow a predictable structure so users can easily review recommendations and listings. * Hallucination Avoidance:** The AI should avoid inventing brands, sizes, conditions, or resale values when confidence is low or information is unclear. * User Control: AI-generated listings and recommendations should remain editable so users can review and adjust outputs before posting.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Core use cases * Clear sellable bundle: same-size, good-condition tops/pants → AI recommends “Sell as bundle” * Low-value items: basic onesies or worn basics → AI recommends “Donate” or “Bundle only” * Seasonal items: winter jacket in summer → AI suggests selling closer to fall * Complementary bundle: 4 tops only → AI suggests adding same-size pants/shorts * Listing generation: AI creates title, description, price range, and tags Edge cases * Mixed sizes: 2T, 3T, 4T in one pile → AI separates into different bundles * Mixed condition: some clean, some stained → AI recommends selling clean items and donating damaged ones * Unclear photo: blurry/poor lighting → AI asks for clearer photos instead of guessing * Brand not visible: AI should not invent a brand * High-value item: snowsuit, jacket, boots, brand-name gear → AI may suggest selling separately instead of bundling Negative cases * Unsafe / damaged items: broken toys, missing parts, heavily stained clothing → AI recommends donation/disposal, not resale * Too little information: one unclear image with no size → AI gives low-confidence output and asks for size/photo tag * Overpricing risk: AI avoids inflated value estimates without evidence * Hallucination check: AI must not claim “Nike,” “4T,” or “excellent condition” unless visible or provided * Out-of-scope request: user asks for unrelated parenting advice → AI redirects back to resale/item assessment
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?The solution will use a hybrid AI model approach to balance cost and performance. A multimodal model such as GPT-4o will be used selectively for image analysis, while a lower-cost text model will handle recommendation logic, bundle suggestions, and listing generation. The multimodal model is well-suited for: * identifying item types from uploaded photos * detecting visible condition issues * recognizing bundle opportunities * understanding seasonality cues from clothing types The lower-cost text model is well-suited for: * sell vs donate recommendations * bundle optimization suggestions * listing title and description generation * structured outputs for UI rendering Key capabilities include: * image understanding * contextual reasoning * natural language generation * structured recommendation outputs Key limitations include: * difficulty identifying exact sizes or brands unless clearly visible * potential hallucination of unclear item details * inability to accurately predict real marketplace sale prices To reduce these limitations, the system uses: * optional user inputs (size, brand, notes) * confidence indicators * hallucination-avoidance instructions in the prompt * editable AI-generated outputs AI integration workflow: 1. User uploads photos and optional item details 2. The multimodal model analyzes the images 3. The text model generates recommendations, bundle suggestions, and marketplace-ready listings 4. Structured outputs are displayed in the UI as recommendation cards, bundle suggestions, and listing previews before the user posts to marketplace.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.| Field | Format | Source | Validation | | -------------- | ------------------------- | ------------------ | ---------- | | item_photos | image upload (1–3 photos) | User upload/camera | Required | | current_date | date | System-generated | Required | | current_season | string/enum | System-generated | Required |
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?| Field | Format | Source | Validation | | ------------- | -------------- | ------------------------- | ----------------------- | | child_size | enum or string | User input | Optional | | brand | string | User input | Optional, max 100 chars | | item_notes | free text | User input | Optional, max 500 chars | | item_category | enum | User input or AI inferred | Optional |
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)1. STRUCTURE COMPLIANCE ☐ Output follows the required response structure ☐ All required sections are present: * Decision * Bundle Suggestions * Listing * Next Best Action ☐ Recommendation is one of: * Sell * Donate * Keep * Bundle ☐ Confidence level is included: * High * Medium * Low 2. RELEVANCE ☐ Recommendations match the uploaded items and user-provided details ☐ Bundle suggestions are practical and logically grouped ☐ Seasonal recommendations are appropriate for the current time of year ☐ Suggested next actions are actionable and realistic 3. PRACTICALITY ☐ Recommendations prioritize reducing user effort ☐ Low-value items are recommended for donation when resale is unlikely to be worthwhile ☐ Bundle suggestions improve resale likelihood or value ☐ Outputs help users make faster decisions with minimal manual work 4. ACCURACY & HALLUCINATION AVOIDANCE ☐ AI does not invent brands, sizes, or item details ☐ AI acknowledges uncertainty when images are unclear ☐ Estimated value ranges are realistic and not exaggerated ☐ No contradictory recommendations are present 5. TONE & CLARITY ☐ Outputs are concise and easy to scan ☐ Tone is practical, helpful, and trustworthy ☐ Listing titles and descriptions sound natural for marketplace use ☐ Explanations clearly justify the recommendation 6. USER CONTROL ☐ AI-generated listings remain editable ☐ Users can review recommendations before posting ☐ Recommendations provide clear next steps rather than forcing actions
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?1. BUNDLE QUALITY ☐ Suggested bundles feel practical and logical ☐ Bundles reflect realistic parent buying behavior ☐ Complementary item suggestions meaningfully improve bundle value 2. RESALE WORTHINESS ☐ Recommendations appropriately balance effort versus resale value ☐ “Sell” recommendations feel worthwhile for the expected return ☐ “Donate” recommendations appropriately reduce unnecessary effort 3. TONE & TRUST ☐ Recommendations feel helpful and trustworthy ☐ Tone is practical and non-judgmental ☐ Explanations feel realistic rather than overly optimistic 4. LISTING QUALITY ☐ Listing titles feel natural and marketplace-friendly ☐ Descriptions are clear, concise, and believable ☐ Suggested pricing feels reasonable to human reviewers 5. VISUAL INTERPRETATION ☐ AI correctly interprets item type from photos ☐ AI reasonably assesses visible condition ☐ Bundle suggestions appropriately reflect uploaded items 6. USER VALUE ☐ Recommendations reduce decision fatigue ☐ Workflow feels faster than manual resale listing creation ☐ Users feel the AI meaningfully saves time and effort
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.* You are an organized parent who's already figured it out. * Tone should be warm, calm, and direct. * Do not use words or phrases like "must buy" etc., * Do not assume condition of the item like smoke-free home, no stains, holes. Ask the user before suggesting a post. * Check other marketplaces for similar items when suggesting a price. * Provide the post details first then the suggestions.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?I made a lot of changes inital prompt was very basic and kept adding to it when testing- PROMPT ITERATION LOG: V1.0 → V2.0 (Current major revision based on testing and edge cases) CHANGES MADE: 1. Added image validation layer before resale analysis 2. Added negative-case handling for invalid or unclear images 3. Added explicit valid vs invalid image identification rules 4. Added hallucination prevention instructions 5. Added confidence scoring (High / Medium / Low) 6. Added bundle optimization logic 7. Added seasonality guidance 8. Narrowed ICP to time-constrained parents of children ages 0–4 9. Shifted from listing-generation assistant to resale decision-support assistant 10. Added effort vs resale value reasoning 11. Added structured output rules for consistent responses 12. Expanded tone guidance to ensure outputs remain practical, honest, concise, and non-judgmental 13. Restricted analysis to children’s clothing only 14. Added workflow stopping rules for invalid images to prevent incorrect analysis WHY THE CHANGES WERE MADE: * Reduce hallucinated outputs from unclear or invalid images * Improve consistency and reliability of multimodal analysis * Better align recommendations with real parent resale behavior * Reduce unnecessary listing effort for low-value items * Improve trustworthiness and transparency * Create stronger differentiation from generic marketplace AI tools * Improve structured output consistency for downstream UX and evaluation * Handle real-world edge cases discovered during testing POTENTIAL CHANGES TO TEST: 1. Add few-shot examples for difficult resale scenarios 2. Add marketplace-specific optimization (Facebook Marketplace vs Poshmark) 3. Refine pricing recommendations using live comparable sales data 4. Add handling for mixed-size or mixed-season bundles 5. Add personalization based on parent preferences (maximize value vs minimize effort) 6. Add resale likelihood scoring 7. Improve handling of partially visible clothing items 8. Add escalation flow for low-confidence image analysis 9. Test shorter vs more detailed recommendation formats 10. Add support for toys or baby gear in separate analysis flows TRACKING METHOD: * All prompt versions stored in version control * Each version tagged with: * date * version number * change summary * Structured changelog maintained for every prompt update * Edge-case test library used to compare outputs across versions * Prompt evaluations run after every major revision * Outputs reviewed for: * hallucination frequency * validation accuracy * bundle quality * recommendation usefulness * pricing realism * output consistency ITERATION TRIGGERS: * Hallucinated item details detected * Invalid images incorrectly analyzed * Inconsistent bundle recommendations * Poor pricing realism * User feedback indicating confusion or low trust * New edge case discovered during testing * Output formatting inconsistency * Confidence scoring mismatch * Recommendation quality drops below evaluation threshold
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?DATA SOURCES: 1. USER-UPLOADED CLOTHING IMAGES Source: Parent-uploaded photos of children's clothing Use: Brand/logo recognition, character detection, item counting, clothing classification, resale signal extraction Preparation: Images resized and normalized before analysis; outputs structured into a consistent JSON schema 2. USER-PROVIDED ITEM DETAILS Source: User input (size, brand, condition notes, season, additional context) Use: Personalization and recommendation accuracy Preparation: Structured fields injected into analysis workflow; user-entered details take precedence when conflicts exist 3. EVALUATION DATASET Source: Internal testing images and edge-case scenarios Use: Measure recommendation quality, failure rates, and regression testing Preparation: Structured evaluation table containing expected outcome, actual outcome, pass/fail status, issue type, severity, fix applied, and retest result 4. MARKETPLACE REFERENCE DATA Source: Research from Facebook Marketplace, eBay, Poshmark, and similar resale platforms Use: Inform pricing ranges, bundle strategies, and listing generation Preparation: Manual benchmarking and prompt guidance; not directly retrieved at runtime 5. KIDS RESALE DECISION RULES Source: Product research, mentor feedback, user interviews, and domain-specific resale heuristics Use: Effort-versus-value recommendations, bundle optimization, donation/recycling guidance Preparation: Encoded into prompts and structured recommendation logic RAG IMPLEMENTATION: Current MVP: RAG is not used in the MVP. Reason: Most recommendations are driven by image analysis, user-provided details, and structured resale decision logic rather than retrieval from a large external knowledge base. Future RAG Implementation: Chunking: 500-token chunks with 50-token overlap Embedding: OpenAI text-embedding-3-small Vector Store: Supabase pgvector Knowledge Base Contents: * Kids resale best practices * Brand-specific resale guidance * Seasonal demand information * Bundle optimization strategies * Marketplace listing guidance * Donation and recycling rules Retrieval: Top 3–5 most relevant chunks based on: * Detected brand * Clothing category * Season * Recommendation type * User-provided details Injection: Retrieved content added to the recommendation prompt as "Reference Information" NOTE: RAG would be supplementary. Core recommendations rely on GPT-4o image understanding, structured resale logic, and LLM reasoning rather than retrieval alone.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.COMMON INPUTS & EXPECTED OUTPUTS 1. LARGE MIXED CLOTHING LOT (MOST COMMON) Input: * Photo containing 15–25 children's clothing items * Mix of brands, characters, tops, bottoms, hoodies, and sleepwear * Example from testing: * Vans hoodie * Nike shorts * GAP shirt * Bluey shirt * Mickey Mouse shirt * Donald Duck shirt * Jurassic World shirt * Assorted pants and everyday basics Expected Output: * Recommendation: Full Bundle * Confidence: High * Estimated Value: $35–55 CAD * Why This Recommendation: * Multiple recognizable brands and characters * One listing minimizes effort * Splitting adds limited additional value * Alternative Bundle Options: * Character Clothing * Branded Items * Everyday Basics * Marketplace-ready listing Why It Matters: This represents the primary user problem: a parent cleaning out a large pile of outgrown clothing and deciding the most practical way to sell it. --- 2. SAME-BRAND BUNDLE Input: * Multiple items from the same brand * Example from testing: * 6 Ralph Lauren toddler garments * Mix of polos, shorts, and tops Expected Output: * Recommendation: Full Bundle * Confidence: High * Estimated Value: $35–50 CAD * Why This Recommendation: * Same brand * Similar buyer audience * Minimal value lost by bundling * One listing instead of multiple * Alternative Bundle Options: * Ralph Lauren Tops * Ralph Lauren Bottoms * Marketplace-ready listing --- 3. SINGLE BRANDED ITEM Input: * One children's clothing item * Example from testing: * Pekkle hockey-themed pajamas Expected Output: * Recommendation: Sell Individually * Confidence: High * Estimated Value: $8–15 CAD * Why This Recommendation: * Single item * Recognizable brand * Niche buyer appeal * Individual sale maximizes value * Marketplace-ready listing --- 4. CHARACTER CLOTHING GROUP Input: * Multiple character-themed items * Example: * Bluey shirt * Mickey Mouse shirt * Donald Duck shirt * Jurassic World apparel Expected Output: * Recommendation: Create Themed Bundles or Full Bundle (depending on inventory size) * Confidence: High * Suggested Bundle Types: * Character Collection * Disney & Friends * Character Sleepwear * Marketplace-ready listing --- 5. LOW-VALUE EVERYDAY BASICS Input: * Assorted children's basics * No strong brands * No recognizable characters * Lower resale potential Expected Output: * Recommendation: Donate, Recycle, or Bundle * Confidence: Medium * Reasoning: * Limited resale value * Listing effort exceeds expected return * Donation or bundling is more practical --- 6. INVALID OR UNUSABLE IMAGE Input: * Non-clothing image * Extremely blurry photo * Toys, furniture, storage bins, or unrelated items Expected Output: * No recommendation generated * Error message: "Unable to analyze this image. Please upload a clearer photo containing children's clothing." * No pricing or listing generated --- 7. USER DETAIL CONFLICT Input: * User enters size 3T * AI scan suggests size 2T Expected Output: * Recommendation generated normally * User-entered size prioritized * Trust message displayed: "You entered size 3T. The image may indicate a different size. User-entered information was used." --- EXPECTED INPUT DISTRIBUTION (MVP ASSUMPTION) Based on user research, testing, and the target parent workflow: * 50% Large mixed clothing lots * 20% Same-brand bundles * 15% Single items * 10% Character-focused clothing groups * 5% Invalid images or edge cases EXPECTED OUTPUT DISTRIBUTION * Full Bundle: ~50% * Sell Individually: ~20% * Create Themed Bundles: ~15% * Donate / Recycle: ~10% * Invalid / Retry Required: ~5% KEY INSIGHT The most common workflow is not a single-item listing. Parents typically upload a large pile of outgrown clothing and want help deciding whether to sell as one bundle, split into themed bundles, donate, or recycle. Outgrown.ai is designed to optimize the effort-versus-value tradeoff rather than simply generate listings.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)AI LIMIT TESTING / EDGE CASE EVALUATION SET 1. AMBIGUOUS BRAND IDENTIFICATION Input: * Clothing with partially visible logo * Folded garment obscuring brand tag * Similar-looking logos Example: * Partial Nike swoosh visible * Folded GAP logo Expected Behavior: * Lower confidence score * Avoid guessing brand * State uncertainty when confidence is low Risk Being Tested: * Hallucinated brand detection --- 2. CHARACTER DETECTION EDGE CASES Input: * Partially visible character graphics * Multiple overlapping character items Examples: * Bluey partially hidden * Disney graphic partially cropped * Character visible only on sleeve Expected Behavior: * Detect character when confidence is sufficient * Avoid inventing characters * Lower confidence when uncertain Risk Being Tested: * False-positive character recognition --- 3. LARGE MIXED INVENTORY Input: * 20+ garments in a single image * Mixed brands, sizes, and clothing categories Example: * Vans * Nike * GAP * Bluey * Disney * Everyday basics Expected Behavior: * Correct inventory summary * Logical bundle recommendation * Alternative bundle options * No duplicate counting Risk Being Tested: * Item counting accuracy * Bundle optimization quality * Recommendation quality at scale --- 4. SIMILAR ITEM COUNTING Input: * Multiple nearly identical garments Example: * Under Armour shorts lot Expected Behavior: * Count only clearly visible items * Use count confidence when uncertain * Avoid overcounting overlapping garments Risk Being Tested: * Vision counting limitations --- 5. LOW-QUALITY IMAGES Input: * Slight blur * Poor lighting * Wrinkled or partially obscured clothing Expected Behavior: * Continue analysis when possible * Reduce confidence appropriately * Request clearer image if analysis quality is insufficient Risk Being Tested: * Robustness to real-world parent photos --- 6. USER DATA CONFLICTS Input: * User-entered size differs from visible size tag Example: * User enters 3T * Visible tag appears to be 2T Expected Behavior: * Prioritize user-entered data * Surface conflict notification * Continue recommendation process Risk Being Tested: * Trust and transparency --- 7. LOW-INFORMATION ITEMS Input: * Plain shirt * No visible logo * No visible character * No readable size Expected Behavior: * Generate recommendation using available information * Avoid inventing missing details * Lower confidence score Risk Being Tested: * Hallucination prevention --- 8. LOW-VALUE RESALE SCENARIOS Input: * Generic children's basics * Stained or heavily worn items when visible * No recognizable brands Expected Behavior: * Recommend donate, recycle, or bundle * Explain effort-versus-value tradeoff Risk Being Tested: * Recommendation quality --- 9. INVALID IMAGE (OUT-OF-DOMAIN) Input: * Toys * Furniture * Storage bins * Household items * Adult clothing Expected Behavior: * Reject image * Explain why image cannot be analyzed * Request children's clothing image Risk Being Tested: * Domain boundary enforcement --- 10. FAILED ANALYSIS HANDLING Input: * API failure * Invalid model response * Parsing failure Expected Behavior: * Show analysis failure message * Allow retry * Do not generate fake recommendations Risk Being Tested: * Reliability and error handling --- 11. EXTREMELY CROWDED IMAGES Input: * Clothing pile where individual items cannot be reasonably distinguished Expected Behavior: * Lower confidence * Request smaller groups or clearer photo * Avoid inaccurate inventory analysis Risk Being Tested: * Limits of image understanding --- 12. UNSEEN BRANDS OR CHARACTERS Input: * Rare children's brands * Lesser-known characters * New licensed merchandise Expected Behavior: * Use generic category descriptions when uncertain * Avoid hallucinating brand names * Continue recommendation process Risk Being Tested: * Generalization beyond known examples KEY FAILURE MODES BEING MEASURED * Hallucinated brands or characters * Incorrect item counts * Incorrect bundle recommendations * Failure to reject invalid images * Overconfidence with limited information * Poor handling of conflicting user input * API and processing failures * Recommendation quality degradation on large inventory lots SUCCESS CRITERIA The AI should: * Acknowledge uncertainty when confidence is low * Avoid inventing missing information * Stay within the children's clothing domain * Provide explainable recommendations * Fail safely when analysis is unreliable * Maintain recommendation quality across single items, bundles, and large mixed lots
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?MANUAL REVIEW PROCESS: REVIEW PROTOCOL: Sample Size: * 14 representative test scenarios from the evaluation dataset * Included both success cases and known failure cases * Covered: * Pekkle single-item resale * Ralph Lauren bundle * Large mixed inventory lot * Disney/character clothing * Bluey detection * Under Armour shorts counting * Invalid images * Failed analysis handling * User-entered data conflicts * Listing generation quality Review Method: * Manual review performed after major prompt and workflow changes * Outputs compared against expected recommendations and business rules * Regression testing performed using previously identified failure cases Scoring: * Pass / Fail for objective criteria * Manual qualitative review for recommendation usefulness and listing quality REVIEW CRITERIA: 1. Recommendation Accuracy (Pass/Fail) * Was the correct sell, bundle, donate, or recycle recommendation generated? 2. Inventory Detection Accuracy (Pass/Fail) * Were visible items counted and classified reasonably? 3. Brand & Character Recognition (Pass/Fail) * Were recognizable brands, logos, and characters correctly identified? 4. Bundle Optimization Quality (Pass/Fail) * Did suggested bundle strategies align with the inventory shown? 5. Listing Quality (Pass/Fail) * Did the listing sound like a real parent selling children's clothing? * Did it avoid AI inspection language? 6. Recommendation Explainability (Pass/Fail) * Could a parent understand why the recommendation was made? 7. Error Handling (Pass/Fail) * Did invalid images and failed analyses fail safely? RESULTS TRACKING: Tracking Method: * Evaluation spreadsheet maintained for all test scenarios Fields Tracked: * Test ID * Scenario * Expected Outcome * Actual Outcome * Pass/Fail * Failure Type * Severity * Fix Applied * Retest Result Examples Tracked: * Bluey detection failure * Under Armour counting issue * Generic bundle recommendation outputs * AI-style listing generation * Failed analysis fallback behavior FAILURE ANALYSIS: For Each Failed Case: 1. Document Failure Mode Examples: * Bluey character not detected * Under Armour shorts overcounted * Bundle recommendation too generic * Listing sounded like AI inspection notes * Failed analysis displayed misleading recommendation 2. Identify Root Cause Examples: * Character recognition gap * Similar-item counting limitation * Weak bundle recommendation logic * Listing prompt quality issue * Error-state handling weakness 3. Implement Fix Examples: * Expanded character detection instructions * Added counting safeguards and confidence handling * Added Alternative Bundle Options * Rewrote listing style guidance * Added explicit error states and retry flow 4. Retest Run the same image again after changes and compare: * Recommendation * Recognition quality * Bundle strategy * Listing quality * Error handling KEY FINDINGS: Highest Performing Cases: * Ralph Lauren bundle recommendation * Pekkle single-item recommendation * Large mixed inventory recommendations after bundle optimization improvements Most Challenging Cases: * Bluey character detection * Under Armour inventory counting * Bundle recommendation specificity * Listing quality and tone OUTCOME: Manual review identified several recommendation quality and trust issues early in development. Iterative testing and retesting improved character recognition, bundle optimization, listing quality, recommendation explainability, and failure handling, resulting in a more reliable and transparent resale decision-support experience.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?EVALUATION RESULTS Because this was an MVP, evaluation focused on representative scenario testing and failure analysis rather than large-scale statistical validation. Evaluation Set Size: * 14 representative test scenarios * Included success cases, edge cases, and known failure cases * Covered: * Single-item resale * Same-brand bundles * Mixed inventory lots * Character clothing * Invalid images * User-data conflicts * Failed analyses * Listing generation CORE CRITERIA RESULTS 1. Recommendation Accuracy Pass Rate: 12 / 14 (86%) Failures: * Early Pekkle recommendation quality issue * Generic bundle recommendation output Fixes: * Recommendation reasoning improvements * Alternative Bundle Options Final Status: * Passed after iteration --- 2. Brand & Character Recognition Pass Rate: 12 / 14 (86%) Failures: * Bluey not detected * Some character recognition gaps Fixes: * Expanded character and logo detection instructions Final Status: * Passed after iteration --- 3. Inventory Detection Pass Rate: 13 / 14 (93%) Failures: * Under Armour shorts overcounted (8 vs 7) Fixes: * Counting safeguards * Confidence handling Final Status: * Improved, but remains a known limitation for similar-item lots --- 4. Bundle Recommendation Quality Pass Rate: 11 / 14 (79%) Failures: * Generic "Full Mixed Lot" suggestions * Limited actionable bundle guidance Fixes: * Alternative Bundle Options * Character, brand, and category grouping Final Status: * Passed after iteration --- 5. Listing Quality Pass Rate: 13 / 14 (93%) Failures: * Listings sounded like AI inspection reports Fixes: * Parent-style listing guidance * Removal of photo-analysis language Final Status: * Passed after iteration --- 6. Recommendation Explainability Pass Rate: 10 / 14 (71%) Failures: * Recommendations lacked reasoning Fixes: * Added "Why This Recommendation" * Added confidence indicators Final Status: * Passed after iteration --- 7. Error Handling & Safe Failure Pass Rate: 14 / 14 (100%) Failures Found: * Failed analyses could appear as valid recommendations Fixes: * Explicit error states * Retry flow Final Status: * All tested failure scenarios handled safely OVERALL RESULTS Initial Evaluation: * Multiple failures in character recognition, recommendation transparency, bundle quality, and listing quality. Final Evaluation: * 13 / 14 scenarios met expected outcomes (93%) * Remaining known limitation: * Similar-item counting in highly repetitive inventory (e.g., Under Armour shorts) KEY LEARNING The largest quality improvements came from: 1. Character recognition enhancements 2. Bundle recommendation improvements 3. Parent-style listing generation 4. Recommendation reasoning and transparency The evaluation process directly influenced the product's evolution from an AI listing generator to an explainable resale decision-support tool focused on effort-versus-value recommendations.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?EDGE CASES IDENTIFIED DURING TESTING 1. CHARACTER RECOGNITION FAILURES Test Case: * Bluey apparel Issue: * Bluey was not detected in early versions despite being clearly visible. Impact: * Reduced inventory understanding. * Missed opportunity for character-based bundle recommendations. Resolution: * Expanded character recognition instructions and detection logic. --- 2. SIMILAR ITEM COUNTING Test Case: * Under Armour shorts lot Issue: * System counted 8 shorts when only 7 were visible. Impact: * Incorrect inventory totals. * Reduced trust in recommendations. Resolution: * Added counting safeguards. * Added count confidence handling. Status: * Improved, but remains a known limitation for highly repetitive inventory. --- 3. GENERIC BUNDLE RECOMMENDATIONS Test Case: * Large mixed clothing lot containing Nike, Vans, GAP, Disney, Bluey, and Jurassic World items. Issue: * Recommendation was correct ("Full Bundle"), but suggested bundles were generic and not actionable. Example: * "Full Mixed Lot" Impact: * Limited user value. * Did not demonstrate inventory understanding. Resolution: * Added Alternative Bundle Options using detected brands, characters, and clothing categories. --- 4. SAME-BRAND BUNDLE OPTIMIZATION Test Case: * Ralph Lauren clothing lot Issue: * Early bundle outputs created unnecessary bundle splits. Impact: * Added complexity without increasing resale value. Resolution: * Refined effort-versus-value logic to recommend a Full Bundle when appropriate. --- 5. LISTING QUALITY Test Case: * Multiple inventory scenarios Issue: * Generated listings sounded like AI inspection reports. Examples: * "No visible damage" * "Visible in photos" * "Appears to be" Impact: * Poor marketplace usability. Resolution: * Rewrote listing style rules to generate parent-style marketplace listings. --- 6. FAILED ANALYSIS HANDLING Test Case: * Processing failures and invalid model outputs Issue: * Failed analyses could fall back to misleading recommendation states. Impact: * Risk of displaying incorrect recommendations. Resolution: * Added explicit error states. * Added retry messaging. * Removed misleading fallback behavior. --- 7. INVALID IMAGE VALIDATION Test Case: * Images that could not be reliably analyzed Issue: * Validation behavior was inconsistent in earlier versions. Impact: * Risk of unreliable recommendations. Resolution: * Added stricter image validation and failure handling. --- 8. RECOMMENDATION EXPLAINABILITY Test Case: * Multiple recommendation scenarios Issue: * Recommendations were provided without sufficient reasoning. Impact: * Users could not understand why a recommendation was made. Resolution: * Added "Why This Recommendation" explanations. * Added confidence indicators. --- 9. USER TRUST & DATA CONFLICTS Test Case: * User-entered information conflicts with AI-detected information. Example: * User enters size 3T while image suggests a different size. Issue: * Potential trust issue if AI silently overrides user input. Resolution: * User-entered information is prioritized. * Conflict messaging added to improve transparency. --- 10. LARGE MIXED INVENTORY ANALYSIS Test Case: * Large clothing piles containing 15–25 garments. Issue: * Recommendation was generally correct, but early versions struggled to provide meaningful organization and resale guidance. Impact: * Reduced usefulness for the primary user workflow. Resolution: * Added bundle optimization logic. * Added character, brand, and category-based alternative bundle strategies. KEY LEARNING The most significant edge cases were not recommendation failures. The largest challenges involved: * Character recognition (Bluey) * Similar-item counting (Under Armour shorts) * Bundle recommendation specificity * Listing quality * Recommendation transparency and trust Addressing these issues shifted the product from a simple AI listing generator to an explainable resale decision-support tool focused on effort-versus-value recommendations.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?PROMPT & SYSTEM ADJUSTMENTS The prompt and workflow were iteratively refined based on evaluation results, edge-case testing, mentor feedback, and manual review. 1. CHARACTER RECOGNITION IMPROVEMENTS Trigger: * Bluey apparel was not detected during testing. Adjustment: * Expanded character detection instructions. * Added stronger guidance for identifying children's licensed characters and graphics. Outcome: * Improved recognition of Bluey and other character-based clothing. --- 2. INVENTORY COUNTING SAFEGUARDS Trigger: * Under Armour shorts lot was counted as 8 items instead of 7. Adjustment: * Added instructions to count only clearly distinguishable garments. * Added confidence handling for similar-item lots. * Reduced aggressive counting behavior. Outcome: * Improved inventory count reliability. --- 3. BUNDLE OPTIMIZATION REFINEMENTS Trigger: * Recommendations were often correct, but bundle suggestions were too generic. Example: * "Full Mixed Lot" Adjustment: * Added bundle optimization logic based on: * Brands * Characters * Clothing categories * Introduced Alternative Bundle Options. Outcome: * More actionable recommendations and clearer inventory understanding. --- 4. EFFORT-VS-VALUE DECISION LOGIC Trigger: * Mentor feedback indicated that listing generation alone was not sufficiently differentiated. Adjustment: * Shifted recommendation logic toward: * Sell individually * Bundle * Donate * Recycle * Added effort-versus-value reasoning. Outcome: * Product evolved from listing generation to resale decision support. --- 5. LISTING QUALITY IMPROVEMENTS Trigger: * Listings sounded like AI inspection reports. Examples: * "No visible damage" * "Visible in photos" * "Appears to be" Adjustment: * Added parent-style listing guidance. * Removed photo-analysis language. * Prioritized natural marketplace listing tone. Outcome: * Listings became more realistic and easier to post directly to Facebook Marketplace. --- 6. RECOMMENDATION EXPLAINABILITY Trigger: * Users could not easily understand why a recommendation was generated. Adjustment: * Added "Why This Recommendation" explanations. * Added recommendation confidence indicators. * Surfaced factors influencing the recommendation. Outcome: * Increased transparency and trust. --- 7. IMAGE VALIDATION IMPROVEMENTS Trigger: * Validation behavior was inconsistent during testing. Adjustment: * Added stricter image validation rules. * Added handling for images that could not be analyzed reliably. Outcome: * Reduced risk of unreliable recommendations. --- 8. ERROR HANDLING IMPROVEMENTS Trigger: * Failed analyses could appear as valid recommendation states. Adjustment: * Added explicit error states. * Added retry messaging. * Removed misleading fallback recommendations. Outcome: * Safer failure behavior and improved user trust. --- 9. USER-ENTERED DATA PRIORITIZATION Trigger: * Potential conflicts between user-entered information and AI-detected information. Example: * User enters size 3T while image suggests a different size. Adjustment: * User-entered information is prioritized. * Conflict messaging added to surface discrepancies. Outcome: * Increased trust and transparency. --- 10. VISION PROMPT OPTIMIZATION Trigger: * GPT-4o Vision accounted for the majority of scan cost. Adjustment: * Reduced Vision prompt size from approximately 25,000 characters to approximately 10,800 characters. * Removed redundant instructions while preserving output schema and functionality. Outcome: * Lower operating cost with comparable output quality across evaluation scenarios. KEY PRODUCT EVOLUTION Initial Focus: * AI-generated resale listings Feedback & Testing Findings: * Parents struggled more with deciding what to do with clothing than writing listings. * Recommendation transparency was critical for trust. * Bundle strategy was often more valuable than listing generation. Final Focus: * Explainable resale decision support * Effort-versus-value recommendations * Bundle optimization * Recommendation transparency * Parent-friendly listing generation This evolution was driven directly by testing results, edge-case analysis, and Product Faculty mentor feedback throughout development.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?EVALUATION APPROACH: Chosen Approach: Human review + structured regression testing. The MVP uses human evaluation as the primary method because recommendation quality, bundle usefulness, trust, and listing tone are subjective and need parent/product judgment. Automated checks are used as support, but not as the main evaluator. EVALUATION METHODS: 1. HUMAN REVIEW Use: * Primary evaluation method What It Measures: * Recommendation usefulness * Bundle quality * Listing quality * Trust and explainability * Whether the recommendation makes sense for a parent Why: * Outgrown.ai is making judgment-based resale recommendations, not just extracting facts. 2. SCRIPT-BASED CHECKS Use: * Secondary validation method What It Measures: * Required fields are present * JSON output is valid * Recommendation type is allowed * Error states display correctly * Listing is generated only when analysis succeeds Why: * These checks help catch formatting, schema, and workflow failures quickly. 3. MODEL GRADER Use: * Future option, not primary for MVP What It Could Measure: * Listing tone * Hallucinated details * Recommendation clarity * Explanation quality Why Not Primary Yet: * The test set is still small. * Human review is more reliable for early product judgment. SCALING TESTING: Current MVP: * Test with a representative evaluation set from real project examples: * Pekkle single item * Ralph Lauren bundle * Large mixed clothing lot * Disney / character clothing * Bluey detection * Under Armour counting issue * Invalid images * Failed analysis handling * User-entered data conflicts * Listing quality examples Next Stage: * Expand the test set to 50–100 images across: * Single items * Same-brand bundles * Large mixed lots * Character clothing * Low-value basics * Blurry images * Crowded images * Conflicting user-entered details * Invalid / out-of-domain images Future Scale: * Create a labeled evaluation dataset with expected outcomes. * Run every prompt/model change against the full dataset. * Track pass/fail rates by scenario type. * Add automated regression checks for known failures. * Use a model grader for first-pass review, then human-review flagged cases. SUCCESS METRICS: * Recommendation accuracy * Bundle strategy quality * Brand/character recognition * Listing quality * Hallucination rate * Error handling success rate * User trust / explainability score * Retest pass rate after fixes SCALING GOAL: The goal is to move from manual MVP testing to a repeatable evaluation pipeline: Input image + user details → AI output → Schema validation → Automated checks → Model-assisted review → Human review for flagged cases → Pass/fail tracking in evaluation spreadsheet NOTE: For the Product Faculty MVP, human review is the most appropriate evaluation method because the core question is not only whether the AI detected clothing correctly, but whether the recommendation is useful, trustworthy, and actionable for a parent.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?EVALUATION FREQUENCY: PRE-LAUNCH / DEVELOPMENT 1. PROMPT CHANGES Frequency: * Every prompt update Process: * Re-run the full evaluation set after any change to: * Vision prompt * Recommendation prompt * Bundle optimization logic * Listing generation logic Goal: * Prevent regressions in previously fixed scenarios. Examples: * Pekkle single item * Ralph Lauren bundle * Large mixed lot * Bluey detection * Under Armour shorts * Invalid image handling --- 2. MODEL CHANGES Frequency: * Every model upgrade or configuration change Process: * Run the complete evaluation dataset. * Compare outputs against previous baseline results. Goal: * Verify recommendation quality remains stable after model updates. --- 3. NEW EDGE CASE DISCOVERED Frequency: * Immediately Process: * Add the new scenario to the evaluation dataset. * Re-run affected test cases. * Track failure, fix, and retest result. Example: * Bluey detection failure * Under Armour counting issue * Failed analysis fallback behavior Goal: * Prevent the same issue from reappearing. --- POST-LAUNCH MONITORING 1. WEEKLY REVIEW Frequency: * Weekly Process: * Review user feedback. * Review failed analyses. * Review low-confidence recommendations. * Identify recurring failure patterns. Goal: * Detect emerging edge cases and recommendation quality issues. --- 2. MONTHLY EVALUATION RUN Frequency: * Monthly Process: * Re-run the complete evaluation dataset. * Compare pass rates against previous months. Metrics: * Recommendation accuracy * Brand and character recognition * Bundle recommendation quality * Listing quality * Error handling success rate Goal: * Ensure quality remains stable as prompts and workflows evolve. --- 3. MAJOR RELEASES Frequency: * Before every major release Examples: * New recommendation logic * New bundle optimization rules * New model versions * New user-input workflows Process: * Full regression test across all evaluation scenarios. Goal: * Maintain quality and trust. --- EVALUATION DATASET MAINTENANCE Current Evaluation Set: * 14 representative scenarios Examples: * Pekkle single item * Ralph Lauren bundle * Large mixed inventory * Disney / character clothing * Bluey detection * Under Armour counting * Invalid images * Failed analyses * User-entered data conflicts * Listing quality validation Growth Plan: * Add every newly discovered failure case to the evaluation set. * Maintain both: * Success cases * Known historical failures Goal: * Create a growing regression suite that reflects real user behavior and previously observed issues. TRIGGER-BASED EVALUATIONS Immediately re-run evaluations when: * Prompt changes are deployed * Model versions change * New failure modes are discovered * Recommendation quality complaints increase * Significant user feedback is received SUCCESS CRITERIA The evaluation process should ensure: * Previously fixed issues remain fixed * Recommendation quality remains stable * Listing quality remains consistent * Character and brand recognition do not regress * Error handling continues to fail safely * User trust and explainability remain high NOTE: For the MVP, evaluations are run after every meaningful prompt or workflow change. Post-launch, the goal is a combination of continuous monitoring, weekly reviews, monthly regression testing, and immediate evaluation whenever new edge cases are discovered.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?INFRASTRUCTURE: ☐ Cloud deployment configured and tested ☐ Database schema deployed for: * user accounts * assessments * AI outputs * feedback tracking ☐ API rate limits configured ☐ Error handling and centralized logging in place ☐ Rollback procedures documented for prompt and infrastructure changes ☐ Image upload and storage workflows tested API INTEGRATIONS: ☐ Multimodal AI API integrated and tested for image analysis ☐ Lower-cost text model integrated and tested for recommendation generation ☐ API keys secured in environment variables ☐ Fallback handling implemented for API failures and timeouts ☐ Cost monitoring alerts configured ☐ Hybrid AI routing logic tested to reduce unnecessary image-processing costs APPLICATION READINESS: ☐ Core workflows tested: * image validation * recommendation generation * bundle suggestions * listing generation ☐ Invalid image handling implemented ☐ Upload limits configured (1–3 images per assessment) ☐ Responsive UI tested across target devices ☐ User feedback collection flows implemented MONITORING: ☐ Error tracking configured ☐ Product analytics implemented ☐ AI request/response logging enabled ☐ API uptime and latency monitoring active ☐ Alert thresholds defined for: * API failures * hallucination spikes * abnormal cost increases * upload failures ☐ Seasonal traffic spike monitoring planned SECURITY: ☐ Authentication flow tested ☐ Data encryption verified ☐ HTTPS enforced ☐ Input validation and sanitization implemented ☐ Uploaded image access restricted to authenticated users ☐ User data deletion and export workflows validatedPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?DOCUMENTATION: ☐ User documentation (help center articles) ☐ FAQ for common questions ☐ Troubleshooting guide for: * invalid image uploads * failed AI analysis * low-confidence recommendations ☐ Internal runbook for: * AI incidents * API failures * hallucination escalation * infrastructure outages SUPPORT PREPARATION: ☐ Support workflows documented for MVP scope ☐ Support team trained on: * AI recommendation limitations * image validation behavior * Seasonal Cleanup Pass workflows ☐ Escalation procedures defined ☐ Response templates created for: * invalid image handling * recommendation disputes * billing/support issues ☐ AI safety and hallucination incident protocol documented LEGAL/COMPLIANCE: ☐ Terms of Service finalized ☐ Privacy Policy updated ☐ AI recommendation disclaimer reviewed and approved ☐ Data handling and retention procedures documented ☐ User consent flows documented for uploads and account creation STAKEHOLDER ALIGNMENT: ☐ Leadership briefed on phased launch plan ☐ Success metrics agreed upon: * recommendation usefulness * Seasonal Cleanup Pass conversion * repeat seasonal engagement ☐ Risk mitigation plan reviewed: * hallucinations * incorrect recommendations * invalid image handling * infrastructure scaling risks ☐ Internal communication and escalation plan approved
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?PHASE 1: CLOSED PILOT (Week 1–4) Audience: 25–50 parents of children ages 1–4 who actively use resale platforms such as Facebook Marketplace Recruitment: Parent Facebook groups, local parenting communities, friends/family referrals Goal: Validate whether AI recommendations reduce decision fatigue and make resale workflows faster and easier during seasonal clothing cleanouts Metrics: * Recommendation usefulness rating * % of users who proceed to create a listing * Bundle recommendation acceptance rate * Time-to-list reduction * User feedback quality * Interest in Seasonal Cleanup Pass pricing Exit Criteria: * > 60% of users report the recommendations were helpful * Users successfully complete the resale workflow * Positive feedback on Seasonal Cleanup Pass concept * No major hallucination or invalid recommendation issues PHASE 2: EXPANDED BETA (Week 5–8) Audience: 200–500 waitlist users Goal: Validate onboarding, recommendation quality, infrastructure stability, and willingness to pay for the Seasonal Cleanup Pass Metrics: * Assessment completion rate * Recommendation confidence accuracy * % of users saving or posting generated listings * API cost per assessment * Seasonal Cleanup Pass conversion rate * User re-engagement during additional cleanout sessions Exit Criteria: * > 80% assessment completion rate * Stable infrastructure and manageable API costs * Positive feedback on bundle recommendations and listing quality * Initial validation of Seasonal Cleanup Pass monetization PHASE 3: PUBLIC LAUNCH (Week 9+) Audience: General availability for parents of children ages 1–4 Channels: App Store, parenting communities, social media, referral sharing Goal: Validate product-market fit, monetization, and repeat seasonal engagement Metrics: * Monthly active users * Seasonal Cleanup Pass conversion rate * Recommendation engagement rate * Repeat usage during seasonal cleanouts * Customer acquisition cost (CAC) * Reactivation rate across multiple clothing turnover cycles ROLLBACK TRIGGERS: * Significant hallucination or inaccurate recommendation patterns * Invalid image analysis failures affecting core workflows * > 20% drop in recommendation usefulness metrics * Critical bugs affecting uploads or AI output generation * API costs significantly exceeding projections * Major privacy or data handling issue detected
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?INITIAL SCALE: * Start with a limited pilot audience before broader rollout * Use a hybrid AI model approach to reduce expensive image-processing costs * Limit uploads to 1–3 images per assessment during MVP * Prioritize core workflows only: * image validation * recommendation generation * bundle suggestions * listing generation * Design infrastructure around short-term, high-intensity seasonal usage periods rather than constant daily usage INFRASTRUCTURE SCALING: * Use cloud-based infrastructure that supports horizontal scaling * Separate image-processing workflows from text-generation workflows * Store assessment history and AI outputs in scalable cloud databases * Compress uploaded images before AI analysis to reduce processing costs and latency * Prepare for seasonal usage spikes during: * spring cleanouts * fall/winter wardrobe transitions * daycare clothing turnover periods API & COST MANAGEMENT: * Monitor: * API request volume * token usage * image-processing frequency * cost per assessment * Seasonal Cleanup Pass usage intensity * Implement request throttling and upload limits to prevent abuse * Cache repeat assessments where possible to reduce duplicate processing * Use lower-cost models for text generation where image reasoning is not required * Monitor whether Seasonal Cleanup Pass pricing remains profitable relative to AI processing costs PERFORMANCE MONITORING: * Track: * upload success rate * AI response time * recommendation completion rate * failed image validation rate * infrastructure uptime * Monitor hallucination frequency and invalid recommendation patterns * Use logging and alerting for API failures and processing bottlenecks * Monitor system performance during seasonal traffic spikes USER VOLUME MONITORING: * Gradually increase user access through phased rollout * Monitor daily active users and concurrent assessments * Track repeat seasonal usage and user reactivation patterns * Monitor Seasonal Cleanup Pass conversion and renewal behavior * Evaluate whether infrastructure and API costs remain sustainable as usage grows SCALE-UP READINESS: * Expand infrastructure only after validating: * recommendation quality * cost efficiency * operational stability * Seasonal Cleanup Pass monetization performance * Future optimizations may include: * asynchronous image processing * background job queues * recommendation caching * fine-tuned lightweight models for common resale scenarios * predictive scaling during expected seasonal demand spikes
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?FAQ CONTENT: ☐ “How does the AI determine what is worth selling?” ☐ “Why did the AI recommend donation instead of selling?” ☐ “How does bundle optimization work?” ☐ “How accurate are resale recommendations?” ☐ “What types of images are supported?” ☐ “How does the Seasonal Cleanup Pass work?” ☐ “Can I edit AI-generated listings?” ☐ “How is my data and uploaded photos handled?” PRODUCT DEMO ASSETS: ☐ Short product walkthrough video ☐ Before/after resale workflow comparison ☐ Example AI recommendation screenshots ☐ Demo showing: * image upload * AI sell vs donate recommendations * AI bundle suggestions * listing generation ☐ Seasonal cleanup workflow demo ☐ Pilot user testimonials and feedback quotes USER GUIDES: ☐ “How to take better clothing photos” guide ☐ “Best practices for kids clothing bundles” guide ☐ “What types of kids clothing are worth selling?” guide ☐ Quick-start onboarding guide ☐ Marketplace listing tips for parents ☐ Troubleshooting guide for unclear or invalid images WEBSITE / LANDING PAGE CONTENT: ☐ Clear explanation of the product value proposition ☐ “Worth selling or not?” positioning messaging ☐ Example AI outputs and recommendation cards ☐ “How it works” section ☐ Seasonal Cleanup Pass explanation ☐ Privacy and AI transparency messaging ☐ Waitlist or signup flow SOCIAL / COMMUNITY CONTENT: ☐ Example resale success stories ☐ Educational content about bundle optimization ☐ Seasonal resale tips for parents ☐ “What’s worth selling vs donating” educational content ☐ Parent-focused productivity and decluttering messaging ☐ Seasonal closet-cleanout campaigns ☐ Short demo clips for social media platforms AI TRANSPARENCY & TRUST ASSETS: ☐ Explanation of confidence levels ☐ Clear AI limitation messaging ☐ Hallucination-avoidance explanation ☐ User-editable recommendation messaging ☐ Data privacy and consent explanations ☐ Explanation of why some items are recommended for donation instead of resale
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?LAUNCH COMMUNICATION: ☐ Launch timeline shared with stakeholders ☐ Pilot and beta rollout plans documented ☐ Success metrics communicated before launch ☐ Risk areas and rollback triggers reviewed PROGRESS REPORTING: ☐ Weekly progress updates shared internally ☐ Dashboard tracking: * user growth * recommendation usage * retention * API costs * user feedback ☐ Pilot learnings and usability findings documented ☐ Infrastructure and AI performance monitoring reviewed regularly OUTCOME REPORTING: ☐ Post-launch summary reports prepared ☐ User feedback and qualitative insights shared with stakeholders ☐ AI recommendation quality metrics reviewed ☐ Retention, engagement, and conversion metrics analyzed CROSS-FUNCTIONAL ALIGNMENT: ☐ Product, engineering, and design teams aligned on MVP scope ☐ AI limitations and hallucination risks communicated internally ☐ Support and moderation workflows reviewed before launch ☐ Escalation procedures documented for critical issues DECISION-MAKING & ITERATION: ☐ Feedback loops established for rapid iteration ☐ Prompt changes and AI tuning documented ☐ Experiment results tracked and reviewed ☐ Launch phases reviewed against exit criteria before scaling access
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?DATA COLLECTED: * Account info (email, name) * Uploaded children’s clothing photos * Optional item details (size, brand, notes) * Assessment history (recommendations, bundle suggestions, listing outputs) * User feedback (helpful/not helpful, sold/unsold/donated) * Seasonal Cleanup Pass usage data * Usage analytics (anonymized) DATA STORAGE: * User data stored in secure cloud storage/database * Encrypted at rest and in transit * Database backups performed regularly * Uploaded images stored only as needed for assessment workflows * Retention: * User data stored until account deletion * Users can remove uploaded photos and assessments * Temporary processing data cleared after analysis where possible DATA ACCESS: * Users can review assessment history * Users can delete uploaded photos and saved recommendations * Users can export their account data * No data sold to third parties * Limited internal access on a need-to-know basis PRIVACY COMPLIANCE: * Privacy Policy clearly explains data usage * User consent obtained during signup and photo upload * PIPEDA-aware framework (Canada) * GDPR-ready privacy practices where applicable * Product intentionally minimizes collection of child-specific personal information AI-SPECIFIC PRIVACY: * User data is not used to train AI models without explicit consent * Uploaded images are only used for clothing resale analysis workflows * AI-generated outputs remain editable by users * Assessment data may be logged for quality and safety evaluation * Aggregated and anonymized usage data may be used to improve recommendation quality * AI recommendations are advisory only and user-controlled before posting listings
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?CONTENT MODERATION: * Uploaded images limited to children’s clothing resale content * Invalid image detection used to reject: * shipping boxes * packaging * warehouse/store images * unrelated household objects * blurry or obstructed images * AI outputs monitored for inappropriate, misleading, or unsafe recommendations * Basic moderation filters applied to uploaded text and generated listings * Users can report incorrect or inappropriate AI recommendations * Low-confidence outputs trigger clarification requests instead of guessing LEGAL & COMPLIANCE: * Privacy Policy and Terms of Service clearly explain platform usage and AI limitations * Product avoids collecting unnecessary child-specific personal information * Users retain ownership of uploaded content * AI-generated resale estimates presented as guidance only, not guaranteed market values * AI-generated listings and recommendations remain fully editable and user-controlled before posting * Product positioned as a resale workflow assistant, not a marketplace or pricing authority AUDIT & MONITORING: * Assessment outcomes and feedback may be logged for quality review * AI recommendations monitored for hallucinations and inaccurate outputs * Invalid image rejection rates monitored for false positives/negatives * User feedback used to identify recurring recommendation issues * Structured outputs support easier review and testing of AI behavior REGULATORY COMPLIANCE: * PIPEDA-aware privacy practices (Canada) * GDPR-ready framework where applicable * Encrypted data storage and transmission * Consent obtained for uploads and account creation * User data deletion and export supported AI SAFETY CONTROLS: * Hallucination-avoidance instructions included in system prompts * Confidence indicators displayed for uncertain recommendations * AI instructed not to invent brands, sizes, conditions, or resale values * Invalid images rejected before resale analysis begins * Human-editable outputs ensure users remain in control before posting listings
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?USER SUCCESS METRICS: ENGAGEMENT & USABILITY: * Assessment completion rate * Time-to-recommendation * Time-to-list reduction compared to manual workflows * Repeat usage during seasonal cleanouts * User reactivation across multiple clothing turnover cycles AI RECOMMENDATION QUALITY: * Recommendation usefulness rating * Bundle recommendation acceptance rate * % of users who proceed to listing generation * User confidence in AI recommendations * Helpful / not helpful feedback ratio * Invalid image rejection accuracy USER VALUE METRICS: * Reduction in decision fatigue (self-reported) * % of users who successfully declutter clothing items * % of users who post generated listings * User satisfaction / NPS * Average number of items assessed per cleanup session BUSINESS VALUE METRICS: GROWTH METRICS: * Monthly active users (MAU) * User acquisition growth * Waitlist conversion rate * Referral and organic sharing rates * Seasonal reactivation rate MONETIZATION METRICS: * Seasonal Cleanup Pass conversion rate * Revenue per cleanup session * Customer acquisition cost (CAC) * Lifetime value (LTV) * Repeat Seasonal Cleanup Pass purchases OPERATIONAL METRICS: * API cost per assessment * Infrastructure cost per active user * AI response success rate * Invalid image upload rate * Recommendation generation latency * Seasonal traffic spike performance PRODUCT-MARKET FIT SIGNALS: * Users returning for multiple seasonal cleanouts * Increased assessment volume over time * Positive qualitative feedback from parents * Users recommending the product to other parents * Strong engagement with bundle recommendations and listing generation * Willingness to pay for the Seasonal Cleanup Pass
AI MetricsHow will you measure AI performance and accuracy?IMAGE VALIDATION ACCURACY: * % of valid children’s clothing images correctly accepted * % of invalid images correctly rejected * False positive rate for image rejection * False negative rate for invalid image acceptance * User feedback on image validation quality RECOMMENDATION QUALITY: * Recommendation usefulness rating * Accuracy of sell vs donate recommendations * Bundle recommendation acceptance rate * User agreement with AI recommendations * % of recommendations leading to listing generation LISTING GENERATION QUALITY: * User satisfaction with generated titles and descriptions * % of listings edited before posting * Marketplace-readiness rating from users * Clarity and relevance of generated outputs * Structured output compliance rate HALLUCINATION & SAFETY MONITORING: * Frequency of invented brands, sizes, or item details * Incorrect resale value estimation rate * Low-confidence recommendation frequency * Invalid or contradictory output rate * Safety and moderation incident reports PERFORMANCE & RELIABILITY: * AI response success rate * Average AI response time * Recommendation completion rate * Failed AI request rate * API uptime and reliability HUMAN EVALUATION: * Manual review of recommendation quality * Human assessment of bundle usefulness * Human review of seasonal timing recommendations * Review of tone, practicality, and trustworthiness * User feedback collected during pilot and beta phases CONTINUOUS IMPROVEMENT: * Prompt iteration testing and tracking * A/B testing of recommendation logic and output structure * Monitoring repeat user behavior and reactivation patterns * Using anonymized feedback data to improve recommendation quality * Tracking Seasonal Cleanup Pass usage and recommendation engagement trends
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?USER SUPPORT CHANNELS: * In-app support/help section * FAQ and troubleshooting guides * Email support for account or recommendation issues * Support contact form for reporting incorrect AI outputs * Self-service guidance for invalid image uploads and low-confidence recommendations SUPPORT OWNERSHIP: * Product team responsible for AI recommendation quality * Engineering team responsible for infrastructure, API reliability, and upload failures * Support workflows documented for: * invalid image handling * failed AI analysis * incorrect recommendations * account and billing issues * Clear ownership assigned for: * AI incidents * infrastructure outages * privacy concerns * user-reported issues ESCALATION PROCESS: * Low-severity issues handled through standard support workflows * Repeated hallucinations or inaccurate recommendations escalated to product and engineering teams * Infrastructure failures escalated to engineering immediately * Privacy or data concerns escalated through documented compliance procedures * Critical incidents tracked and reviewed post-resolution USER COMMUNICATION: * Users informed when recommendations have low confidence * Invalid image rejection includes recovery guidance and clearer upload instructions * Transparent messaging used for: * AI limitations * processing failures * recommendation uncertainty * Users maintain control through editable AI-generated outputs FEEDBACK & ISSUE TRACKING: * Helpful / not helpful feedback collected on recommendations * User-reported issues logged and categorized * Support trends monitored for recurring AI or UX problems * Feedback incorporated into prompt iterations and workflow improvements
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?FEEDBACK COLLECTION: * In-app “Helpful / Not Helpful” feedback on recommendations * User surveys during pilot and beta phases * Support tickets and user-reported issues * Usage analytics and workflow drop-off tracking * Monitoring edits made to AI-generated listings and recommendations FEEDBACK TRIAGE: * Issues categorized by: * AI recommendation quality * image validation accuracy * infrastructure/API failures * UX/usability issues * billing or account concerns * Repeated user complaints grouped into recurring issue categories * Hallucination-related issues flagged for immediate review * Product and engineering teams review high-frequency issues regularly BUG & ISSUE PRIORITIZATION: * Critical issues prioritized based on: * user impact * recommendation accuracy risk * infrastructure reliability * privacy/security implications * Highest priority issues include: * incorrect or misleading AI recommendations * failed uploads or broken workflows * API outages * privacy or data handling concerns * Lower-priority issues include: * UI inconsistencies * minor formatting issues * non-blocking recommendation improvements ESCALATION & OWNERSHIP: * Engineering team owns: * infrastructure * API failures * upload and processing issues * Product team owns: * recommendation quality * prompt tuning * AI workflow improvements * Support team handles: * user-reported issues * escalation routing * communication updates * Critical incidents escalated immediately to engineering and product leadership COMMUNICATION & INCIDENT MANAGEMENT: * Critical issues communicated internally through documented escalation workflows * Incident status and resolution updates tracked centrally * Users informed when: * AI recommendations may be degraded * infrastructure outages occur * uploads or processing fail * Post-incident reviews conducted for major failures or recurring hallucination patterns CONTINUOUS IMPROVEMENT: * Prompt iterations tracked and version controlled * A/B testing used to improve recommendation quality and UX * User feedback incorporated into future recommendation logic and workflows * Recurring issue trends monitored to prioritize future product improvements
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?AI PERFORMANCE MONITORING: * Track recommendation completion rate * Monitor hallucination frequency and inaccurate recommendations * Monitor low-confidence recommendation rates * Track image validation acceptance/rejection accuracy * Monitor recommendation usefulness feedback trends INFRASTRUCTURE & API MONITORING: * Monitor API uptime and availability * Track AI response latency and processing times * Monitor upload success/failure rates * Track failed AI requests and timeout frequency * Monitor database uptime and storage utilization COST & USAGE MONITORING: * Monitor: * API request volume * token usage * image-processing frequency * cost per assessment * Seasonal Cleanup Pass usage intensity * Monitor traffic spikes during seasonal cleanout periods * Alert on abnormal usage or unexpected cost increases USER BEHAVIOR MONITORING: * Track assessment completion drop-off points * Monitor repeat usage and seasonal reactivation behavior * Track listing generation engagement * Monitor invalid image upload frequency * Analyze user edits to AI-generated listings and recommendations LOGGING & ALERTING: * Log AI request/response metadata for debugging and quality review * Log failed uploads, processing errors, and invalid outputs * Alert on: * infrastructure outages * elevated API failure rates * hallucination spikes * abnormal latency increases * recommendation quality degradation * Centralized logging used for incident investigation and trend analysis INCIDENT DETECTION & RESPONSE: * Critical incidents escalated automatically to engineering and product teams * Monitoring dashboards used to identify operational bottlenecks * Incident severity classification documented: * low * medium * high * critical * Post-incident reviews conducted for recurring failures or AI quality issues CONTINUOUS IMPROVEMENT: * Prompt iteration performance tracked over time * A/B testing results monitored for recommendation quality improvements * User feedback correlated with AI output performance * Monitoring data used to prioritize future infrastructure and AI optimizations
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?LEARNING COLLECTION: * Collect in-app “Helpful / Not Helpful” feedback on recommendations * Gather qualitative feedback from pilot and beta users * Monitor user behavior: * listing generation * bundle acceptance * repeat usage * Seasonal Cleanup Pass usage patterns * Track user edits to AI-generated listings and recommendations * Analyze recurring support requests and reported issues PERFORMANCE REVIEW: * Conduct regular reviews of: * recommendation quality * hallucination frequency * image validation accuracy * infrastructure performance * API costs * Review operational dashboards and monitoring alerts regularly * Compare performance against defined success metrics and launch goals * Analyze seasonal engagement and user reactivation trends AI SYSTEM IMPROVEMENT: * Continuously refine system prompts and recommendation logic * Use prompt version tracking and rollback procedures * A/B test: * recommendation wording * bundle logic * output structure * pricing suggestions * Improve image validation rules based on false positive/negative trends * Optimize confidence thresholds and clarification prompts over time USER FEEDBACK INTEGRATION: * Prioritize improvements based on: * user impact * frequency of feedback * recommendation quality issues * Incorporate recurring user pain points into future workflow improvements * Validate whether new features improve decision speed and reduce effort * Use anonymized feedback trends to improve recommendation consistency INCIDENT & QUALITY REVIEW: * Conduct post-incident reviews for: * hallucination spikes * infrastructure failures * invalid recommendation patterns * Document root causes and remediation actions * Update monitoring, prompts, or workflows based on incident learnings * Review escalation patterns to identify recurring operational risks PRODUCT ITERATION & ROADMAP: * Use pilot and launch learnings to guide roadmap prioritization * Monitor willingness to pay and Seasonal Cleanup Pass conversion trends * Evaluate opportunities for: * improved bundle optimization * marketplace integrations * better seasonality recommendations * lightweight fine-tuned models for common resale scenarios * Continuously balance recommendation quality, infrastructure cost, and user value
Download the .xlsx ↓