← All capstone projects

Resale

Resale AI

Built by Chesna Henderson Cohort 9 Resale / e-commerce seller tools

Resale AI is an AI-powered sourcing and listing co-pilot for vintage and secondhand resellers. It helps sellers decide whether to buy an item by providing a buy-or-pass recommendation, resale estimates, profit guidance, and confidence levels. The same workflow extends into listing generation and post-sale outcome tracking so sellers can learn from actual results over time.

The problem

Resellers of vintage and secondhand goods work almost entirely by hand. Item identification, pricing, and listing creation are time-intensive and fragmented across multiple apps, and pricing uncertainty is the most severe pain point — sellers lack confidence in resale value, leading to underpricing and missed profit. Existing tools stop short of the real need. Google Lens, Relic, and ChatGPT identify items and give rough valuations; Vendoo and List Perfectly handle cross-listing. But they focus on information retrieval, not decision-making. None help a seller answer the question that matters in the moment: is this item worth buying, and at what price?

The solution

Resale AI (ResaleAI) is an AI sourcing and listing copilot that turns information into decisions. A seller uploads an item image and purchase price and receives a structured BUY/PASS recommendation with confidence, estimated resale range, profit estimate, demand signal, risks, and reasoning. The same workflow continues into negotiation guidance (opening offer, target, walk-away price), one-click listing generation with style regeneration (SEO Optimized, Collector Focused, Concise, Detailed), and outcome tracking that compares actual sale results against the original estimate. Estimates are always framed as directional, never guarantees, keeping a human in the loop for final decisions.

How it works

Resale AI runs on GPT-4o, chosen for multimodal item understanding, structured reasoning, and text generation. A unified capability-routing prompt directs each request to exactly one workflow — item analysis, negotiation guidance, listing generation, style regeneration, or outcome tracking — so the model never completes the wrong task. Strong grounding rules require the model to use only provided inputs, with guardrails against fabricating comparable sales, exact sold prices, measurements, provenance, or authentication status. Missing-input handling returns NEEDS MORE INFO rather than guessing, and authentication-risk guardrails apply to luxury items. A 25-row OpenAI eval dataset drove calibration: baseline runs scored 54–68% (mostly wrong-workflow routing), rose to 80% after consolidating into one routing prompt, and reached 100% pass after targeted edge-case fixes.

Who it's for

The product is B2C, built for individual vintage and secondhand resellers — casual side hustlers, part-time vintage sellers, and full-time independent resellers. The most revenue-impacting users are high-frequency, semi-professional sellers who source multiple times per week and manage moderate-to-high inventory; they are both the daily user and the paying subscriber. A secondary segment includes small resale businesses such as vintage shops and thrift stores managing higher volumes. Revenue is a freemium SaaS subscription: a free tier with limited scans, and paid unlimited usage with advanced pricing insights and bulk tools.

Why it matters

Recommerce is expanding fast: the secondhand apparel market is projected to grow at roughly 15–20% CAGR, outpacing traditional retail, with broader recommerce around 10–15%. Yet the segment remains fragmented and operationally inefficient for sellers working with unstructured, one-of-a-kind inventory — a strong fit for AI-driven efficiency. Resale AI is at the MVP prototype and evaluation stage, with a validated end-to-end workflow. The launch plan moves through internal QA, a closed pilot with selected resellers, a limited beta or A/B test, and broader release once reliability, support documentation, and legal review are complete. Positioning is deliberately careful: a decision-support copilot, not an authenticator or guaranteed-profit tool.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Chesna Henderson
Your Product:ResaleAI
Your Industry:eCommerce
Date:April 24, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Retail, specifically the recommerce (resale) and secondhand marketplace segment, with a focus on vintage and peer-to-peer seller ecosystems.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?The recommerce market is rapidly expanding, driven by changing consumer behavior and increased participation from individual sellers. However, it remains highly fragmented and operationally inefficient for sellers. Tailwinds: Growing consumer demand for sustainable and secondhand goods Rapid growth of peer-to-peer marketplaces (eBay, Depop, Poshmark) Increasing number of casual and professional resellers Advancements in AI enabling automation of unstructured workflows Headwinds: Highly unstructured inventory (especially vintage and one-of-a-kind items) Time-intensive workflows for item identification, pricing, and listing Pricing uncertainty, directly impacting seller profitability Lack of integrated tools to support end-to-end reseller workflows Key competitors: Marketplaces: eBay, Depop, Poshmark (enable selling but not workflow optimization) Seller tools: Vendoo, List Perfectly (focus on cross-listing and operations) AI tools: Google Lens, ChatGPT, Relic (antique identifier apps) — provide identification and valuation but lack decision support and listing automation Key gap: Existing solutions focus on information retrieval, not decision-making and revenue optimization, leaving an opportunity for an AI-native reseller copilot.
What is the projected growth rate of your target market segment over the next 3-5 years?According to the ThredUp 2024 Resale Report, the global secondhand apparel market is expected to grow at approximately 15–20% CAGR, significantly outpacing traditional retail. Broader recommerce categories are estimated to grow at ~10–15% CAGR based on aggregated market research (e.g., Statista).
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?This product is now in the MVP prototype and evaluation validation stage. The initial concept has progressed from ideation into a working prototype that demonstrates the core AI-assisted resale workflow: item analysis, BUY/PASS recommendation, negotiation guidance, listing generation, listing style regeneration, and outcome tracking. The current focus is validating whether the AI can reliably support reseller decision-making, stay grounded in provided inputs, communicate uncertainty, avoid unsupported claims, and produce structured outputs that can be used in the product experience.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)The product is a software platform that helps resellers source, price, and list items more efficiently, ultimately increasing their profitability and throughput. Primary revenue model: Freemium subscription (SaaS) Free tier: Limited number of item scans and listing generations per month Paid subscription: Unlimited usage, advanced pricing insights, and bulk listing tools What the business sells: Access to AI-powered tools for item identification, pricing intelligence, and listing automation Future monetization opportunities: Marketplace integrations (e.g., one-click listing with transaction or listing fees) Value-added analytics for high-volume sellers Potential take rate on completed sales (long-term)
Who is your primary customer base (B2B, B2C, B2B2C)?The primary customer base is B2C, targeting individual resellers and side hustlers who source and sell secondhand or vintage items. There is a secondary opportunity to expand into B2B segments, such as small resale businesses, thrift stores, and vintage shops, particularly those managing higher inventory volumes. The initial focus on B2C enables rapid adoption and product iteration, with potential expansion into B2B as workflows mature.
DifferentiatorsWhat are the key differentiators for your company?Decision intelligence vs. information tools: Unlike existing solutions (e.g., Google Lens, Relic) that focus on identifying items and providing general value estimates, this product enables actionable decisions by calculating profit potential, estimating resale performance, and surfacing risk factors (e.g., condition sensitivity, demand variability). End-to-end reseller workflow (not point solutions): Competitors either support isolated steps (identification, cross-listing, or marketplaces), while this product integrates the full workflow—sourcing → pricing → listing → inventory—into a single AI-powered experience. Reseller-first optimization (revenue-focused): The product is specifically designed for users whose goal is to maximize profit and efficiency, rather than general collectors or hobbyists. Features such as profit estimation, pricing strategy, and listing optimization are tailored to resale outcomes. Multimodal AI + LLM reasoning layer: Combines image understanding with LLM-based reasoning to interpret ambiguous, unstructured inventory (e.g., vintage items with limited metadata), enabling more accurate identification and contextual insights. Automated listing generation (LLM-native capability): Transforms item data into platform-optimized listings (titles, descriptions, tags), reducing manual effort and improving listing quality—capabilities not fully addressed by existing tools. Human-in-the-loop system for trust and control: Users can review, edit, and refine AI outputs, ensuring accuracy while continuously improving system performance and user confidence. In-the-moment sourcing support (“What Is This?” mode): Provides real-time insights during sourcing (in-store), enabling faster buy/no-buy decisions—something most competitors do not support effectively.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?The primary customers are individual resellers (B2C) who purchase tools to improve their efficiency and profitability when selling secondhand goods. These include: Casual resellers and side hustlers Part-time vintage sellers Full-time independent resellers Customers are typically motivated by increasing profit, saving time, and scaling their resale activity. A secondary customer segment includes small resale businesses (B2B) such as vintage shops and thrift stores that manage higher inventory volumes and require more advanced tooling.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Primary end users: Independent vintage and secondhand resellers Most revenue-impacting users: High-frequency and semi-professional resellers who: Source inventory multiple times per week List items regularly across marketplaces Manage moderate to high inventory volumes These users are more likely to convert to paid plans due to frequent usage and clear ROI. Goals: Make fast, confident sourcing decisions (buy / don’t buy) Accurately price items to maximize profit Reduce time spent on research and listing creation Scale their resale business efficiently Roles: Sole operators managing sourcing, pricing, listing, and inventory Small business owners managing resale operations Context: Frequently sourcing in physical environments (thrift stores, flea markets, estate sales) Working with unstructured and inconsistent inventory (vintage, one-of-a-kind items) Operating across multiple resale platforms with limited tools Their workflows are highly manual today, making them strong candidates for AI-driven efficiency gains.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?ResaleAI is a product-led AI prototype that supports the resale workflow from sourcing decision to listing creation and sale outcome review. Core MVP features: - Item Analysis: Users upload an item image or image description, enter purchase price, and optionally add notes. - BUY/PASS Recommendation: AI provides a recommendation with confidence, estimated resale range, profit estimate, demand signal, risks, and reasoning. - Negotiation Guidance: AI suggests opening offer, target offer, walk-away price, and tactical talking points aligned with the recommendation. - Inventory Save: Users can save promising items into a simple inventory/status workflow. - Listing Generation: AI creates marketplace-ready title, description, tags, pricing recommendation, condition notes, and buyer verification notes. - Listing Style Regeneration: Users can regenerate listings in SEO Optimized, Collector Focused, Concise, or Detailed styles while preserving facts. - Outcome Tracking: Users record sale price, platform, days to sell, and notes so the app can compare actual outcomes against original estimates. The MVP does not include direct marketplace publishing or live marketplace integrations. Those are future enhancements.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)The product is designed for external users, specifically independent vintage and secondhand resellers. The primary persona is a high-frequency reseller (side hustler or full-time seller) who regularly sources, prices, and lists items across online marketplaces. This user is both: The end user (uses the product daily) The buyer (pays for the subscription) They are highly motivated by profit maximization, efficiency, and scalability, making them strong adopters of AI-driven tools.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?Happy path using the MVP prototype: 1. User opens ResaleAI while sourcing or reviewing inventory. 2. User uploads an item image and enters a purchase price. 3. User optionally adds notes such as condition concerns, maker marks, measurements, or provenance details. 4. AI analyzes the item and returns a structured sourcing recommendation, including BUY/PASS, confidence, estimated resale range, estimated profit, demand signal, risks, and negotiation guidance. 5. User reviews the recommendation and decides whether to buy or pass. 6. If promising, the user saves the item to inventory. 7. User generates a marketplace-ready listing draft with title, description, tags, pricing recommendation, condition notes, and buyer verification notes. 8. User can regenerate the listing in different styles, including SEO Optimized, Collector Focused, Concise, and Detailed. 9. User manually lists the item on an external marketplace such as eBay, Etsy, Mercari, Poshmark, or Depop. 10. ResaleAI records listing status only; direct marketplace publishing is not included in the MVP. 11. After the item sells, the user records sale price, days to sell, platform, and notes. 12. ResaleAI compares actual sale outcome against the original AI estimates to support learning and future sourcing decisions.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Pain Points Uncertainty in item identification Especially for vintage or unbranded items, making research difficult Pricing uncertainty (MOST SEVERE) Users lack confidence in estimating resale value, leading to underpricing or missed profit opportunities Time-intensive manual research Requires switching between multiple apps and platforms Inefficient listing creation Writing titles, descriptions, and tags is repetitive and slow Inability to assess profitability during sourcing Users cannot easily determine if an item is worth buying in real time Fragmented inventory management No unified system to track items across the lifecycle Most severe & frequent pain points: Pricing and profit uncertainty Inability to make confident sourcing decisions Time spent per listing
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.The highest-value AI opportunity is decision support during sourcing. Resellers need to quickly determine whether an item is worth buying, how much they should pay, what risks to consider, and how easily the item can become a listing. Ranked AI opportunities: 1. Real-time item analysis and BUY/PASS recommendation 2. Estimated resale value, profit range, demand signal, and time-to-sell guidance 3. Negotiation guidance to protect margin 4. Listing generation and style regeneration 5. Risk detection for condition, authentication, weak evidence, and missing inputs 6. Outcome tracking to compare actual sale results against AI estimates The MVP focuses on these workflows using structured AI outputs and guardrails. Live marketplace integrations and personalized pricing models are future enhancements.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.The following AI-powered solutions were ideated to address key pain points across sourcing, pricing, listing, and inventory workflows: Image-based item identification (“What Is This?”) Real-time profit calculator during sourcing AI-powered pricing recommendations with rationale Automated listing generator (title, description, tags) Demand forecasting (likelihood of sale, time-to-sell) Risk detection (authenticity flags, condition sensitivity) Smart photo guidance (suggest angles / missing details) Voice-based item input (“describe what you see”) Batch processing for bulk inventory uploads Auto-categorization and attribute extraction Cross-platform listing optimization (eBay vs Depop tone) Inventory performance tracking (profit per item, sell-through) AI-powered negotiation assistant (suggest counteroffers) Trend insights (what items are currently in demand) Personalized pricing strategy (fast sell vs maximize profit) Duplicate detection (avoid listing similar items twice)
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.High Impact + High Feasibility: Real-time sourcing decision support (Scan → Profit + Buy/No-Buy) Automated listing generation (LLM-powered) Smart pricing engine with explanation
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?Target state workflow for the MVP: 1. User scans or uploads an item image. 2. User enters purchase price and optional notes. 3. AI returns structured item analysis with recommendation, confidence, resale estimate, profit estimate, demand signal, risks, authentication/condition checks, and negotiation guidance. 4. User decides whether to BUY, PASS, or request more information if inputs are insufficient. 5. User saves promising items to inventory. 6. User generates a listing draft from the AI analysis. 7. User edits the listing and can regenerate it using supported styles: SEO Optimized, Collector Focused, Concise, or Detailed. 8. User manually lists the item on an external marketplace. 9. User marks the item as listed in ResaleAI. 10. User records sale outcome after the item sells. 11. AI compares actual sale price, profit, and days-to-sell against the original estimate. The MVP does not include direct marketplace publishing or live marketplace integrations. These are future enhancements. The current workflow focuses on decision support, listing draft creation, inventory status tracking, and outcome comparison.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Wireframe link: https://docs.google.com/presentation/d/1n2u7xp2DP6E3VngJLYT4mz1vypATU5Q2j_0gBxWrDvQ/edit?usp=sharing Users navigate the product through two primary workflows that reflect real reseller behavior: in-store sourcing for decision-making and at-home listing for execution. The experience begins on the home screen, where users can either start a new scan or return to their saved sourcing list. In the in-store flow, users upload an item image and enter a purchase price, then initiate analysis. The AI Results screen serves as the core decision point, presenting a visually dominant BUY or PASS recommendation, accompanied by a confidence level, estimated resale range, profit estimate, and demand signal. Supporting this decision, the interface surfaces structured reasoning ("Why this recommendation") and potential risks ("Check before buying"), enabling users to quickly assess both upside and uncertainty. From this screen, users can save items to their sourcing list and continue scanning, or proceed directly to listing generation. The sourcing list acts as a lightweight inventory system, displaying saved items in a card-based layout with key attributes such as image, resale estimate, profit, and status (for example, "Needs Listing"). AI-generated prioritization signals, such as "High Profit" or "Fast Seller," help users decide which items to act on first. In the at-home workflow, users select items from this list to generate listings. The AI produces structured outputs, including title, description, tags, and suggested price, which are then presented in a clear, editable format. A dedicated editing step allows users to refine AI-generated content before saving or exporting, reinforcing a human-in-the-loop approach. The UI is designed to support AI features through strong visual hierarchy and structured layouts. Key outputs such as BUY or PASS decisions are emphasized with size and color, while supporting data is organized into scannable sections. Confidence indicators and explanatory text are always visible, promoting transparency and trust. The layout prioritizes speed and clarity for in-store use while also supporting deeper interaction during listing creation. Overall, the interface directly mirrors the structure of AI outputs, ensuring seamless integration between model responses and user experience.
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Prototype Link: https://resalewithai.lovable.app The prototype demonstrates the core end-to-end AI workflow for ResaleAI: item scan, sourcing analysis, BUY/PASS recommendation, negotiation guidance, save-to-inventory, listing generation, listing style regeneration, and outcome tracking. AI inputs include an item image or image description, purchase price, optional user notes, item analysis, marketplace, listing style, and sale outcome data. AI outputs are displayed in structured sections so users can quickly understand the recommendation, confidence, estimated resale range, profit potential, risks, and next steps. Essential MVP features include: - Item analysis and BUY/PASS recommendation - Confidence and uncertainty handling - Estimated resale and profit ranges - Negotiation guidance - Save item to inventory - Listing generation - Listing style regeneration - Manual status tracking for listed/sold items - Sale outcome tracking against original AI estimates Deferred future features include: - Real-time marketplace integrations - Direct marketplace publishing - Multi-image analysis - Batch listing generation - Personalized pricing based on user history - Advanced analytics and seller performance dashboards
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Initial master prompt design: The initial prompt positioned ResaleAI as a focused resale copilot, not a general-purpose assistant. The AI was instructed to help vintage and secondhand resellers make fast, practical decisions by identifying items, estimating resale value, evaluating BUY/PASS decisions, and generating listing drafts. Initial prompt principles: - Use a concise, experienced reseller tone - Prioritize decision-making over general explanation - Provide structured outputs with recommendation, confidence, pricing estimates, risks, and next steps - Clearly communicate uncertainty when inputs are ambiguous - Avoid exact sales claims or unsupported facts - Generate marketplace-ready listing content from known item facts only Prompt: You are an AI-powered resale copilot designed to help vintage and secondhand resellers make fast, confident, and profitable decisions. Your primary goal is to help users: 1. Identify items from images or descriptions 2. Estimate resale value and demand 3. Evaluate profitability (buy vs pass decisions) 4. Generate high-quality, marketplace-ready listings Tone & Style - Be concise, clear, and actionable - Sound like an experienced reseller or sourcing expert - Avoid generic or vague language - Prioritize usefulness over completeness - Use plain language, not technical jargon - Be confident but not absolute—acknowledge uncertainty when needed Reasoning & Decision Support - Always explain recommendations (e.g., why an item is a good or bad buy) - Highlight key factors: demand, brand, condition sensitivity, trends - When possible, provide ranges (e.g., resale value, profit) - Surface risks or uncertainties clearly - If confidence is low, say so and explain why Output Structure When analyzing an item, structure responses clearly: - Recommendation: BUY or PASS - Confidence: High / Medium / Low - Summary: 1–2 sentence explanation - Key Insights: - Demand signal - Pricing context - Notable attributes (brand, era, style) - Risks: - Potential downsides or things to check - Estimated Resale Range - Estimated Profit (if purchase price is known) Listing Generation When generating listings: - Create clear, keyword-rich titles optimized for resale platforms - Write concise but compelling descriptions - Include relevant tags/keywords (e.g., brand, era, style) - Adapt tone slightly depending on platform (e.g., more SEO for eBay, more casual for Depop) - Avoid hallucinating unknown details—only include what can be reasonably inferred Tool Usage (if applicable) - Use available data (e.g., comparable listings) to inform pricing and demand - Do not fabricate exact sales data—use phrases like “based on similar listings” when needed Behavioral Guidelines - Optimize for speed and clarity (users are often in-store) - Do not overwhelm with long explanations - Prioritize decision-making over general information - If unsure, provide best estimate and clearly communicate uncertainty - When the user is likely in a sourcing context, prioritize speed and highlight the most important decision factors first. You are not a general assistant—you are a focused resale expert helping users make better buying and selling decisions.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Clarity & Actionability - Outputs must be easily scannable and support fast decisions (e.g., clear BUY/PASS, concise summary). Relevance to Resale Context -Responses must prioritize resale-specific signals (brand, demand, condition, trends), avoiding generic descriptions. Accuracy of Estimates (Directional) -Pricing, profit, and demand outputs should be reasonable estimates, not overly precise or misleading. Hallucination Avoidance -The model must not fabricate specific facts (e.g., exact comps or sales data) and should clearly indicate uncertainty when applicable. Consistency of Structure -Outputs must follow a predictable format (recommendation, confidence, pricing, risks) to support UI rendering and user trust. Explainability (Reasoning Quality) -Each recommendation must include clear reasoning tied to observable factors (e.g., demand patterns, item attributes). Confidence Calibration -Confidence levels should accurately reflect uncertainty, especially for ambiguous or low-quality inputs. Tone & Usability -Tone should be concise, practical, and expert-like, optimized for in-store use (not verbose or generic). Responsiveness -Outputs should be generated quickly enough to support real-time decision-making during sourcing.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?The evaluation set covers happy paths, edge cases, and negative cases across the five core MVP capabilities. Core use cases: - Clear branded item with strong resale potential - Generic item with average resale potential - BUY/PASS recommendation based on purchase price and margin - Negotiation guidance for BUY and PASS items - Listing generation from item analysis - Listing style regeneration across SEO Optimized, Collector Focused, Concise, and Detailed styles - Outcome tracking against original AI estimates Edge cases: - Low-quality or blurry image - No image provided - Ambiguous or unbranded item - Luxury/designer item with authentication uncertainty - Missing purchase price - Missing measurements for listing generation - Incomplete sale outcome data Negative cases: - Non-item image - Pricing unavailable but recommendation attempted - Unsupported authentication claims - Fabricated exact sales data or comparable listings - Listing invents measurements, maker, provenance, age, or condition details - Outcome tracking overstates causality Validation criteria: - Correct capability routing - Required structure present - Recommendation aligns with inputs - Confidence is calibrated - Prices and timelines are framed as estimates - No hallucinated facts - Negotiation protects margin - Listing copy preserves known facts - Outcome tracking compares actuals against estimates without unsupported causal claims
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Selected model: GPT-4o GPT-4o is well suited for the MVP because it supports multimodal item understanding, structured reasoning, and high-quality text generation. These capabilities align with ResaleAI’s core workflows: item analysis, BUY/PASS recommendation, negotiation guidance, listing generation, listing style regeneration, and outcome tracking. Key capabilities: - Interprets item images or image descriptions with user-provided price and notes - Produces structured resale analysis with confidence, risks, and estimates - Generates factual marketplace-ready listing drafts - Regenerates listing copy in different styles while preserving facts - Compares actual sale outcomes against original AI estimates Key limitations: - Estimates are directional, not guarantees - Image quality and missing details affect confidence - Authentication cannot be confirmed through image analysis alone - The MVP does not use live marketplace integrations - The model can hallucinate if prompts and guardrails are weak Mitigations: - Structured output requirements - Confidence and uncertainty language - Authentication-risk guardrails - Missing-input handling - Prompt and grader versioning - Offline evals before prompt/model changes are releasedPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Required inputs vary by workflow. For item analysis: - Item image or item image description - Format: image file or text description - Source: user upload or eval dataset - Requirement: required for item analysis - Purpose: supports item identification, condition assessment, confidence calibration, and risk detection - Purchase price - Format: numeric USD value - Source: user input - Requirement: required for profit estimate, margin calculation, and BUY/PASS recommendation - Purpose: allows the AI to compare estimated resale value against cost and determine whether the item is worth buying For listing generation: - Item analysis - Format: structured text or JSON-like output from the item analysis workflow - Source: AI-generated analysis - Requirement: required for listing generation - Purpose: provides known facts, condition notes, pricing guidance, and uncertainty context for the listing draft For outcome tracking: - Original item analysis or original AI estimates - Format: structured text - Source: prior AI output - Requirement: required for outcome comparison - Purpose: provides the baseline estimates to compare against actual sale results - Actual sale price - Format: numeric USD value - Source: user input - Requirement: required for sale outcome comparison - Purpose: enables comparison of actual sale price against estimated resale range If required inputs are missing, the AI should not fabricate results. It should return NEEDS MORE INFO or UNABLE TO ANALYZE and request the specific missing information.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Optional and user-customizable fields improve output quality, confidence, and usefulness. Optional inputs: - User notes - Format: free text - Source: user input - Impact: improves analysis by adding context such as maker marks, condition concerns, measurements, provenance, flaws, or sourcing location - Marketplace - Format: selected value, such as eBay, Etsy, Mercari, Poshmark, or Depop - Source: user selection - Impact: helps tailor listing tone, keywords, and pricing presentation to the intended marketplace - Listing style - Format: selected value - Supported styles: SEO Optimized, Collector Focused, Concise, Detailed - Source: user selection - Impact: changes the structure and tone of the listing while preserving known facts - Sale platform - Format: selected or typed marketplace name - Source: user input - Impact: supports outcome tracking and helps compare platform performance over time - Days to sell - Format: numeric value - Source: user input - Impact: allows the AI to compare actual time-to-sell against the original estimate - Sale notes - Format: free text - Source: user input - Impact: adds context for outcome tracking, such as accepted offer, buyer feedback, relisting, or condition-related issues - Additional images - Format: image files - Source: user upload - Impact: future enhancement that would improve confidence, condition assessment, and authentication-risk handling These fields are optional, but missing details should reduce confidence where relevant. The AI should clearly state uncertainty and request missing information rather than guessing.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Good output is defined by structure, grounding, usefulness, and uncertainty handling. Criteria: - Follows the required structure for the capability being tested - Completes the correct workflow and does not mix capabilities - Grounds claims in provided inputs only - Clearly labels pricing, resale value, profit, and timeline as estimates - Avoids fabricated comparable listings, exact sales data, measurements, provenance, maker, model, era, material, or authentication status - Calibrates confidence based on image quality, evidence strength, and missing inputs - Aligns BUY/PASS reasoning with price, margin, demand, risk, and confidence - Provides negotiation guidance that protects resale margin - Generates factual, reseller-style listing copy - Compares sale outcomes against original estimates without overstating causality
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Outputs will also be evaluated based on qualitative alignment with the intended “resale copilot” experience defined in the system prompt. This includes assessing whether the reasoning feels credible and reflective of how an experienced reseller would evaluate an item, and whether the explanation meaningfully supports the BUY or PASS recommendation. The usefulness of the output is critical, meaning users should feel more confident and informed after viewing the response. Confidence calibration will be evaluated to ensure that the model appropriately expresses uncertainty, particularly in ambiguous or low-quality input scenarios. Tone will be assessed to ensure it is clear, practical, and human-like, rather than robotic or generic. Additionally, outputs should feel appropriately scoped for the context, prioritizing speed and key decision factors during sourcing rather than overwhelming the user with unnecessary detail. These subjective criteria ensure that the AI not only produces correct outputs, but also delivers a high-quality, trustworthy user experience.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.The final prompt architecture uses a unified capability-routing prompt rather than separate overlapping prompts. The model receives the capability being tested and completes only that workflow. Supported capabilities: - item_analysis - negotiation_guidance - listing_generation - listing_style_regeneration - outcome_tracking Key prompt adjustments: - Explicit capability routing to prevent the model from completing the wrong workflow - Strong grounding rules requiring the model to use only provided inputs - Missing-input handling for blurry, missing, corrupted, incomplete, or non-item images - Authentication-risk guardrails for luxury/designer items - Negotiation margin protection for BUY and PASS recommendations - Listing completeness safeguards, including requesting missing measurements - Outcome tracking limits that prevent unsupported causal claims - Clear rules against inventing measurements, provenance, maker, model, era, authentication status, exact sales data, platform performance, or comparable listings The prompt was refined through eval failures. Early failures included wrong workflow routing, vague missing-photo requests, weak PASS negotiation guidance, authentication-risk gaps, missing measurement requests, and incomplete sale outcome handling. These were addressed through targeted prompt updates.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Prompt iterations were tracked through eval runs, failure analysis, and prompt/grader updates. Major iteration themes: 1. Capability routing: Replaced multiple overlapping user prompts with one unified prompt that completes only the requested workflow. 2. Grounding: Strengthened rules requiring the AI to use only provided inputs. 3. Missing-input handling: Added explicit behavior for missing, blurry, corrupted, incomplete, or non-item images. 4. Authentication risk: Added verification-first guidance for luxury/designer items. 5. Negotiation logic: Added margin-protection rules, especially for PASS items with limited resale upside. 6. Listing completeness: Added requirements to request missing measurements and verification details before final branded listings. 7. Outcome tracking: Added rules to compare actual outcomes against estimates without overstating causality. 8. Grader calibration: Updated the grader to use row-specific expected behavior, pass criteria, fail criteria, and generated output. Prompt evolution is tracked through eval results, failure notes, prompt version changes, and before/after pass rates.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Current MVP data sources: - User-uploaded images or image descriptions - User-entered purchase price - Optional user notes - AI-generated item analysis - Listing style selection - Marketplace selection - User-recorded sale outcomes - Simulated or representative comparable data for evaluation The MVP does not rely on live marketplace integrations or real-time sold comp retrieval. Pricing and demand outputs are directional estimates and must be labeled as estimates. Future production data sources may include: - eBay sold listings - Etsy listings - Mercari data - User resale outcomes - Inventory performance history - Category-specific resale guidance RAG and marketplace retrieval are future enhancements. In the current MVP, guardrails and eval criteria are used to prevent fabricated comps, unsupported exact sales data, and overconfident pricing claims.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Common MVP inputs: - Item image or item image description - Purchase price - Optional user notes - Prior item analysis - Marketplace - Listing style - Sale outcome data Representative sample items: - Prada nylon shoulder bag - Blue and white porcelain vase - Vintage Nike windbreaker - Generic denim jacket - Unbranded ceramic bowl - Blurry handbag image - Non-item image Expected outputs: - Structured item analysis - BUY/PASS/UNABLE TO ANALYZE/NEEDS MORE INFO recommendation - Confidence and uncertainty notes - Estimated resale range and profit range - Negotiation guidance - Listing draft - Listing style regeneration - Outcome comparison against original estimates
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Examples that test the AI’s limits include: Edge cases: - Blurry or low-quality image - Missing image - Ambiguous or unbranded item - Luxury item with authentication uncertainty - Missing purchase price - Missing measurements before listing - Incomplete sale outcome data - Weak or simulated comparable data - Insufficient sold comps - Conflicting user notes and visual evidence Negative cases: - Non-item image - Fabricated exact sales data - Unsupported authentication claim - Pricing unavailable but recommendation shown - PASS item receiving overly aggressive negotiation guidance - Listing invents maker, age, measurements, provenance, or material - Outcome tracking claims causation without evidence These cases are important because they test whether the AI can communicate uncertainty, request missing information, avoid unsupported claims, and refuse to overstate confidence when evidence is weak.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Manual review and eval testing showed that the AI performed well when outputs were structured, grounded in inputs, and routed to the correct workflow. Successful behaviors included structured item analysis, listing generation, outcome comparison, uncertainty handling, and factual listing copy. Failures identified during testing: - Model completed item analysis when the row required negotiation, listing generation, or outcome tracking - Blurry handbag image responses asked for more information too generically - PASS negotiation guidance set walk-away price too close to the current asking price - Luxury/designer negotiation guidance did not fully emphasize verification-first behavior - Designer listing generation did not always request missing measurements - Outcome tracking risked drawing conclusions from incomplete sale data Fixes applied: - Replaced separate user prompts with one capability-routing prompt - Added missing-input and blurry-image handling - Added specific photo request guidance for handbags and bags - Added authentication-risk guardrails for luxury/designer items - Added margin-protection rules for PASS negotiation guidance - Added listing completeness checks for measurements and verification details - Added outcome tracking rules to prevent unsupported causality - Calibrated the grader to evaluate against row-specific pass/fail criteria
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?The AI evaluation achieved a 100% pass rate on the final MVP evaluation dataset after prompt and grader calibration. Earlier baseline runs exposed several issues: -Initial grading showed roughly 54-68% pass, mainly because the model was completing the wrong workflow for some rows. -After consolidating the user prompts into a single capability-routing prompt and improving the grader, the score improved to 80% pass. -Final targeted fixes for edge cases brought the eval to 100% pass. Core criteria validated: -Correct capability routing across item analysis, negotiation guidance, listing generation, listing style regeneration, and outcome tracking -Structured outputs for each workflow -BUY/PASS reasoning aligned with margin, demand, risk, and confidence -No unsupported authentication claims -No fabricated comparable sales, exact sales data, measurements, provenance, or marketplace performance -Negotiation guidance aligned with recommendation and protected resale margin -Listing copy stayed factual and reseller-ready -Outcome tracking compared actual results against estimates without overstating causality
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Observed edge cases from prototype testing and eval work: Real usage / system edge cases: - All comps filtered out - Weak mock data - Insufficient sold comps - Pricing unavailable but recommendation shown - Category context overriding evidence - Async hydration race conditions AI behavior edge cases: - Missing or low-quality images - Non-item images - Luxury/designer authentication uncertainty - Missing purchase price - Incomplete listing details such as measurements - Conflicting user notes and visual evidence - Irrelevant or insufficient comparable data - Active listings mistaken for sold comps - Incomplete sale outcome records These edge cases informed prompt guardrails, missing-input handling, authentication-risk rules, negotiation margin protection, listing completeness requirements, and outcome tracking limits.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?Prompt and system adjustments made based on failures, feedback, and edge cases: Implemented during prototype/eval work: - Unified capability-routing prompt - Stronger grounding rules - Missing-input handling - Fallback-safe outputs - Authentication-risk guardrails - Negotiation margin protection - Listing completeness checks - Outcome tracking causality limits - Scenario-specific model grader - Prompt/grader versioning approach - Frontend state validation and progressive hydration improvements Earlier prototype learnings also informed: - Evidence-quality handling - Dependency gating - Market verification hierarchy - Safer behavior when comps are weak, unavailable, or simulated The final system is designed to avoid unsupported claims, communicate uncertainty, request missing information, and keep AI outputs aligned to the specific workflow being performed.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?The current evaluation approach uses a combination of model-graded evaluation and human review. For the MVP, an OpenAI eval dataset was created with 25 test rows covering: - item_analysis - negotiation_guidance - listing_generation - listing_style_regeneration - outcome_tracking The dataset includes happy path, edge case, and negative case scenarios. Each row includes expected behavior, pass criteria, fail criteria, and quality indicators. A model grader evaluates generated outputs against row-specific criteria using the dataset fields and the model output. Human review was used to inspect failures, determine whether failures were legitimate, and update prompts or grader instructions accordingly. This approach can scale by adding more edge cases, categories, marketplaces, and real user outcome data over time.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Evaluations should be rerun: - After any system prompt, user prompt, grader, or model setting change - Before expanding pilot access - After adding new item categories, listing styles, or workflow capabilities - After recurring production failures or support issues - Before any broader beta or public launch milestone During MVP/pilot, evals should be rerun manually after meaningful prompt changes. In production, evals should become part of a regular regression process, with added cases for any new failure mode observed in user feedback, support tickets, or outcome tracking.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Infrastructure is partially tested. OpenAI eval runs validated core AI behavior and surfaced rate-limit constraints, but full production readiness requires additional API error handling, database persistence tests, monitoring dashboards, retry/backoff logic, and documented rollback procedures for prompts, model settings, and application releases.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Internal training is not yet complete. The AI evaluation framework and core prompt behavior are documented, but support, communications, and legal teams still need enablement materials before launch. Required documentation includes AI limitation guidance, pricing and authentication disclaimers, escalation workflows, approved product messaging, and user-facing help content explaining that resale estimates are directional and not guaranteed.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?ResaleAI should launch through a controlled pilot before broader release. Initial access should be limited to a small group of vintage and secondhand resellers who actively source inventory and can provide qualitative feedback on BUY/PASS recommendations, pricing usefulness, listing quality, and outcome tracking. The pilot should include casual resellers, vintage resellers, and higher-volume sellers to validate the workflow across different sourcing behaviors. During the pilot, access should remain limited while the team monitors AI eval performance, user-reported accuracy issues, rate limits, failed generations, and listing edit rates. After the pilot meets quality thresholds, the product can expand to a larger beta cohort or an A/B test comparing AI-assisted sourcing/listing workflows against the existing manual workflow. Full release should wait until core AI reliability, support documentation, and legal/comms review are complete. Suggested rollout: Phase 1: Internal QA and dogfooding Phase 2: Closed pilot with selected resellers Phase 3: Limited beta or A/B test Phase 4: Broader launch after reliability and trust criteria are met
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Readiness for scale will be managed through phased rollout, usage monitoring, and AI quality tracking. The MVP should not move directly to all users. Instead, access should expand gradually as the team confirms that AI outputs remain reliable, latency is acceptable, and rate limits/error rates are manageable under real usage. Initial volume will be monitored through product and AI reliability metrics, including scan volume, generation volume, failed analysis rate, listing generation success rate, rate-limit frequency, API latency, retry volume, user edits to AI-generated listings, BUY/PASS override behavior, and actual sale outcome variance versus original estimates. Scaling should be gated by operational thresholds. If error rates, rate limits, hallucination reports, support tickets, or failed eval scores increase beyond acceptable limits, rollout should pause until issues are resolved. Prompt versions, model settings, and release configurations should be versioned so the team can quickly roll back to the last stable configuration. Scale readiness checklist: -Prompt and grader versioning -Rate-limit and retry/backoff handling -Monitoring for AI failures and latency -Logging for input/output quality review -Support escalation workflow -Closed pilot feedback loop -Regular eval reruns on representative edge cases -Rollback plan for prompt/model/app changes
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?External communication should focus on setting accurate expectations: ResaleAI helps resellers make faster, more informed sourcing and listing decisions, but it does not guarantee authentication, resale price, or sale outcome. -Product FAQ: Explain what ResaleAI does, what inputs it uses, how BUY/PASS recommendations work, and where estimates may be uncertain. -AI Limitations FAQ: Cover pricing estimates, image quality limits, authentication limitations, missing comps, and marketplace variability. -Demo Video: Show the full workflow: scan item, review BUY/PASS recommendation, generate listing, refine style, mark listed, and record sale outcome. -Getting Started Guide: Walk users through first scan, purchase price entry, notes, listing generation, and outcome tracking. -Trust & Safety / Disclaimer Copy: Clarify that estimates are directional, not guarantees, and that users should verify authenticity, condition, measurements, and pricing before purchase or listing. -Example Outputs: Include good examples for item analysis, negotiation guidance, listing generation, and outcome tracking. -Pilot Feedback Guide: Give pilot users a lightweight way to report inaccurate recommendations, missing information, hallucinations, and sale outcome variance. -Launch Messaging: Position the product as a resale copilot for decision support, not an automated authenticator or guaranteed profit tool.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Launch plans, progress, and outcomes will be communicated through a structured internal rollout cadence. Before launch, the team will share the launch plan, pilot scope, success metrics, known limitations, support process, and rollback criteria with product, engineering, support, comms, and legal stakeholders. During launch, progress will be tracked through regular launch updates summarizing usage volume, AI quality metrics, rate limits/errors, support issues, user feedback, and any prompt or product changes. After each rollout phase, the team will publish a short launch readout comparing results against success criteria and documenting decisions to expand, pause, or revise the rollout. Internal channels/assets: -Launch brief -Pilot tracker -Weekly rollout update -AI eval scorecard -Support issue log -Known limitations and mitigation doc -Post-pilot readout -Final launch retrospective Key outcomes to communicate: -AI pass/fail eval rate -Scan and listing generation volume -Error/rate-limit trends -Hallucination or trust issues -User feedback themes -Listing edit rates -BUY/PASS override patterns -Outcome tracking accuracy versus original estimate
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?ResaleAI will handle user data using privacy-by-design principles. User-uploaded images, purchase prices, notes, listing drafts, inventory records, and sale outcomes should be treated as user-owned data and protected through secure storage, access controls, and clear retention policies. The product should collect only the data required to support sourcing, listing generation, inventory tracking, and outcome comparison. Sensitive data should not be used for unrelated purposes without explicit permission. Any AI processing should be disclosed clearly, including what information is sent to AI services and how outputs are generated. Stored data should be encrypted in transit and at rest, with access limited to authorized systems and team members. Logs should avoid unnecessary personal data and should redact or minimize user-identifiable content where possible. Users should be able to delete uploaded items and associated analysis/listing data. Compliance review should cover privacy policy language, AI vendor data handling, image storage, marketplace data usage, user deletion rights, and any applicable state, federal, or international privacy requirements.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Content moderation, legal, and audit processes are partially defined but should be completed before broad launch. Because ResaleAI handles user-uploaded images, AI-generated resale recommendations, pricing estimates, and potential luxury/authentication scenarios, the product requires clear safeguards around unsupported claims, prohibited content, user data, and marketplace/legal risk. Current AI guardrails reduce risk by preventing unsupported authentication claims, fabricated comparable sales, guaranteed pricing language, and overconfident recommendations when evidence is limited. However, formal legal review is still required before launch to validate disclaimers, privacy practices, image handling, marketplace data usage, and user-facing claims. Compliance should be reviewed for applicable privacy, consumer protection, AI disclosure, data retention, and marketplace-related requirements. ResaleAI should not be positioned as an authentication service, appraisal service, or guaranteed profit tool unless additional legal, operational, and expert-review processes are added.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?Success will be measured by whether ResaleAI helps users make faster, more confident sourcing decisions and create marketplace-ready listings with less effort. User metrics should focus on time saved, recommendation usefulness, listing draft quality, workflow completion, and repeat usage. Business metrics should focus on activation, retention, workflow throughput, AI cost efficiency, and conversion from scan to saved item, generated listing, and recorded sale outcome. Over time, sale outcome tracking will provide a feedback loop to assess whether AI estimates are directionally useful and whether users are improving sourcing efficiency. User Success Metrics -Time saved per sourcing decision -Time saved per listing created -Percentage of scans that result in a clear BUY/PASS recommendation -User trust rating on AI recommendations -Listing draft acceptance rate -Average number of edits needed before listing -Percentage of users who save analyzed items to inventory -Percentage of users who generate listings from saved items -Outcome tracking completion rate -User-reported usefulness of negotiation guidance -Repeat usage rate among active resellers -Reduction in abandoned sourcing/listing workflows Business Value Metrics -Activation rate: users completing first scan and first listing draft -Weekly active users / monthly active users -Scans per active user -Listings generated per active user -Saved inventory items per active user -Conversion from scan to saved item -Conversion from saved item to generated listing -Conversion from generated listing to marked listed -Retention by reseller segment -Paid conversion or upgrade rate, if monetized -Cost per AI analysis/listing generation -Gross margin per active user after AI costs -Support ticket rate related to AI accuracy or trust -Sale outcome variance versus AI estimate over time
AI MetricsHow will you measure AI performance and accuracy?ResaleAI will measure AI performance through a combination of offline evals, human review, production monitoring, and sale outcome comparison. For the MVP, we established an OpenAI eval dataset covering core capabilities: -Item analysis -Negotiation guidance -Listing generation -Listing style regeneration -Outcome tracking -The final calibrated eval achieved a 100% pass rate on the current 25-row MVP dataset. Accuracy Signals -BUY/PASS recommendation alignment with purchase price, resale estimate, demand, risk, and confidence -Confidence calibration for blurry, missing, ambiguous, or high-risk inputs -Pricing and profit estimates labeled as estimates -No fabricated comparable listings, exact sold prices, measurements, provenance, or authentication claims -Negotiation guidance aligned with recommendation and margin protection -Listing copy factuality and completeness -Listing style consistency across regenerated versions -Outcome tracking accuracy versus original estimates Ongoing Measurement -Scheduled eval reruns after prompt/model changes -Human review of sampled outputs -User feedback on incorrect or unhelpful recommendations -Listing edit rate as a proxy for listing quality -BUY/PASS override rate as a proxy for trust or disagreement -Actual sale price versus estimated resale range -Actual days-to-sell versus estimated time-to-sell -Hallucination or unsupported-claim reports -Rate-limit, error, and failed-generation monitoring
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Users should have clear support paths inside the product and through standard support channels. For MVP, users should be able to report inaccurate AI recommendations, hallucinated listing details, authentication concerns, failed scans, pricing issues, and outcome tracking problems directly from the relevant item or analysis screen. Escalation and ownership should be defined before broader launch. Product and engineering should own AI behavior, prompt quality, eval performance, and application bugs. Support should own user communication and issue intake. Legal/compliance should review high-risk issues involving authentication, pricing claims, privacy, or user data. Comms should own public-facing messaging if a pattern of user confusion or trust issues emerges. Each reported issue should be categorized, tracked, and reviewed against the eval suite where appropriate. High-severity issues, such as unsupported authentication claims, fabricated market data, privacy concerns, or repeated misleading recommendations, should trigger escalation, prompt review, and potential rollback.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Feedback and bugs will be gathered through in-app reporting, support tickets, pilot feedback forms, user interviews, eval failures, production logs, and outcome tracking variance. Each issue should be triaged by type, severity, user impact, and whether it indicates a broader AI reliability problem. Issues should be categorized into product bugs, AI quality issues, hallucination or unsupported-claim risks, pricing/recommendation concerns, listing generation problems, outcome tracking issues, and infrastructure failures such as rate limits or failed generations. Critical issues are prioritized based on user harm, trust impact, legal/compliance risk, frequency, and workflow blockage. High-priority issues include unsupported authentication claims, fabricated comparable sales or exact sold data, privacy/data issues, repeated incorrect BUY/PASS recommendations, broken listing generation, and failures that prevent users from completing core workflows. Critical issues should be communicated through an internal escalation channel with clear owners, severity level, affected users, reproduction steps, mitigation status, and next update time. Resolutions may include prompt rollback, model setting changes, feature flagging, product fixes, user communication, or updates to the eval dataset to prevent recurrence.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Post-launch monitoring should combine product analytics, AI quality telemetry, infrastructure logs, and user-reported issues. The goal is to detect both operational failures and AI trust issues quickly, especially failed scans, rate limits, hallucinated claims, unsupported authentication language, pricing inconsistencies, and workflow drop-off. Operational monitoring should track API errors, failed generations, latency, timeout rates, rate-limit events, retry volume, database write/read failures, image upload failures, and frontend state or hydration issues. AI monitoring should track eval pass rate over time, hallucination reports, unsupported authentication claims, BUY/PASS override rate, listing edit rate, incomplete output rate, and sale outcome variance versus original estimates. Logs should include enough context to debug issues, such as prompt version, model version, capability, item category, error type, latency, and generated output status, while minimizing or redacting sensitive user data. High-risk AI outputs should be sampled for review, and prompt/model changes should trigger eval reruns before release.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Post-launch learning will be handled through a continuous improvement loop that combines user feedback, production analytics, AI eval results, support trends, and recorded sale outcomes. The team will regularly review where the AI is accurate, where users override recommendations, where listings require heavy editing, and where actual sale results differ from original estimates. Learnings will be collected from in-app feedback, support tickets, pilot interviews, usage analytics, hallucination reports, failed generations, listing edit behavior, BUY/PASS overrides, and outcome tracking data. These inputs will inform prompt updates, product UX changes, additional eval cases, support documentation, and future marketplace integration priorities. Prompt and model changes should be versioned, tested against the eval dataset, and reviewed before release. Any recurring failure mode, such as unsupported authentication claims, weak pricing confidence, missing measurements, irrelevant comps, or incomplete outcome tracking, should be added to the eval suite so the system improves over time and regressions are caught early.
Download the .xlsx ↓