Media
Clue Watch
Clue Watch is an AI recommendation tool for people who subscribe to multiple streaming services and waste time bouncing between siloed catalogs. Users select their streaming providers, rate movies they know, and receive six grounded recommendations that are filtered to titles actually available on their services. The system uses one LLM to enrich movie metadata into taste-based tags and a second LLM to generate personalized recommendations with explanations.
The problem
People who subscribe to multiple streaming services waste enormous time deciding what to watch — an estimated 110 hours a year — bouncing between siloed catalogs. Recommendations inside each app are trapped in that app's library, existing tools are only about 30% accurate, and 49% of subscribers say they'd cancel over poor discovery. Competitors sit at the extremes. JustWatch overwhelms without real personalization; MovieWiser under-delivers; TasteRay demands 25+ questions just to start. Nothing reliably matches a viewer's mood and taste across all their services at once, quickly.
The solution
Clue Watch (CueWatch) is an AI recommendation layer that sits on top of fragmented streaming ecosystems. Users pick their streaming providers, rate a set of familiar movies with one click, and receive six personalized recommendations — each filtered to titles actually available on their services and explained in a short, friendly match reason. The onboarding is designed to deliver value in 1–3 minutes before asking for a login. Recommendations improve with each rating cycle, and returning users can pick from taste/mood clusters the system has learned over time. The revenue model is hybrid freemium: a free first month, then premium at $3.99/month, unlimited and ad-free.
How it works
Clue Watch uses two LLMs with distinct jobs. An enrichment model (GPT-4.1) reads each movie's title, year, and overview from the TMDB API and extracts grounded tags — emotional tones, pacing, themes, storytelling style — stored in an enrichment table of ~400 movies, using only evidence from the metadata to avoid hallucination. A recommendation model (gpt-5-mini via the Responses API) then selects the final six from a candidate list. Retrieval is two-stage: the user's affinity-weighted taste profile filters the catalog to 20 candidates, which are passed as structured JSON into the prompt. The recommendation prompt applies combination-pattern reasoning over loved and disliked tags and a strict REASON SPEC for conversational, two-sentence explanations. Enrichment scores 6.54/7 relevance with a 7% hallucination rate; switching from gpt-4o-mini to gpt-5-mini eliminated the recommendation hallucinations seen across all earlier test cases.
Who it's for
Clue Watch is B2C, aimed at US viewers aged 25–50 who subscribe to 2–5 streaming services and lead busy lives with work and family. These users have high decision fatigue and value time-saving tools — the segment where nearly half would cancel over poor discovery. The experience is built for the recurring "I want to watch something tonight" moment, spanning trigger, discovery, decision, watching, and reaction. It supports both first-time users who rate movies to build a profile and returning users who select from learned mood clusters.
Why it matters
The U.S. streaming market is projected to grow at a 19.8% CAGR from 2025 to 2030, and streaming fragmentation is a durable, growing pain as households add services. A fast, grounded discovery layer that respects what people already pay for addresses a problem the platforms themselves have little incentive to solve. The MVP is a working end-to-end pipeline on Supabase Edge Functions with deployed database tables and tested flows. Launch runs from a capstone demo with 5–10 real users to an invite-only magic-link beta. Known gaps are documented honestly: rate limiting and rollback aren't yet configured, the ~400-movie catalog must expand past 500+ titles before public launch, and TMDB commercial licensing must be resolved.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Louise Lund Nielsen (Shared with assignments@productfaculty.com) | |||||
| Your Product: | CueWatch | |||||
| Your Industry: | Streaming entertainment | |||||
| Date: | May 8th 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Streaming entertainment | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | **Challanges** Streaming fragmentation across 5+ platforms causes decision fatigue—users spend 110 hours/year deciding what to watch. Existing recommendations only 30% accurate. App fatigue from managing multiple interfaces. *Opportunities:** AI-driven personalization improving. **Competitors:** JustWatch, Reelgood, TasteRay, MovieWiser | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | The U.S. market is anticipated to grow at a CAGR of 19.8% from 2025 to 2030 | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | startup | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | Hybrid freemium model. Freemium first month, the user generate data to improve personalization/algorithm, Goal: Hook users on value. Then premium: $3.99/month, unlimited, no ads | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2C | |||||
| Differentiators | What are the key differentiators for your company? | Users currently spend 110 hours/year deciding what to watch. We deliver personalized recommendations in 1-3 minuts through smart onboarding (vs TasteRay's 25+ Q to get started) plus real-time mood matching. The free trial lets users experience full value immediately and AI learns from precise one-click ratings. Competitors either overwhelm whitout personalization (JustWatch) or under-deliver (MovieWiser). CluWatch hit the sweet spot: fast, smart, AI learnable. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | new product from 0 to 1 | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | -11- | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | -11- | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | US aged 25-50, with 2-5 streaming services. Busy lives (work, family), high decision fatigue. Value time-saving tools. 49% would cancel due to poor discovery/browsing. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | 1. Trigger — "I want to watch something tonight" 2. Discovery / Browsing "what to watch"? 3. Decision 4. Watching 5. Reaction/ (Habit/Retention) | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | A) Group coordination - Watch something with other (partner/kids) with different taste B) Fear of Wasting Time (Will this match my mood?) C) Scroll fatigue — too many options across too many platforms (streaming fragmentation) leads to frustration. D) Poor Personalization (Only recommendations in silos on each app eg. Netflix, HBO, ec) E) Specific show, where do I find it? F) FOMO/Cultural pressure - everyone's talking about it? G) What do my freinds watch/recommendate | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | - C+D) - B) - E) - A) - G) - F) | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | - B+C+D) A personalized AI layer on top of fragmented streaming ecosystems that reduces discovery/browsing time and match the users mood - E) Cross-platform search - A) Help groups agree faster, by AI Group recommendation optimization and Shared watch profiles - F+G) AI Friend/community insights & Trend summarization - F+G) network app with a feed on what you freinds watch/talk about/have rated | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | B) Fear of Wasting Time (Will this match my mood?) C) Scroll fatigue — too many options across too many platforms (streaming fragmentation) leads to frustration. D) Poor Personalization (Only recommendations in silos on each app eg. Netflix, HBO, ec) | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | FIRST TIME USER - flow 1: 1) User provides: streaming services, genre, favorite films + what they enjoy and dislike 2) To refine users taste, the system preview of 20 movies on one page to scroll through 3) The user scrolles through the movies (feedback) and rate them as: ❤️ Loved it 👍 Liked it 😐 Okay 👎 Didn’t Like Or haven't watch, whishlist, play trailer 4) Database stores feedback 5) The system call LLM to get refined recommendations based on all userinputs 6) LLM checks tools (API) to get better reccomendations (and not hallutinate about provider/titles ec.) 7) User gets 6 movies reccommendations with a describtion on the match reason 8) New User can save info by subscribing via mail and save data to next time for free 9) Next iteration (repeat 3-7): LLM gets better data, recommendations improves AFTER CREATING LOGIN - new session - flow 2 1) the user login 2) System shows cluster of the users taste/mood based on the LLMs analysis over time, and user can choose between existing one 4-6 taste/mood OR Try new. 3) User gets 6 movies reccommendations based on their mood 4) Next iteration (same like Flow 1 repeat 3-7): | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | [Insert your response here. Include link to visuals as appropriate.] | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | ## SCREEN 1 — Landing Page / Hero Headline: “Stop scrolling. Start watching.” Subtext: “AI-powered recommendations across all your streaming services.” - “Find something worth watching in minutes, avoid wasting time — personalized to your taste, mood, and streaming services, through 3-short steps... - Then examples on recommendation cards with poster of the recommended films/shows, their title, streaming provider and a short describtion, and why it is recommended to the user. Underneath is the onboarding flow: Headline: “Try Free” Inputs Genre toggles Thriller Comedy Drama Sci-Fi True Crime Feel-good Slow burn Action Documentary Card toggles: Netflix Hulu Max Disney+ Prime Video Apple TV+ Peacock Paramount+ Text Box: “What do you love or hate” In the text box: "Describe what you enjoy or tell us about your favorite movies, shows, actors or what you dislike" - No worries if you skip this — your picks above are enough to get started. Avoid: Sexual violence, abuse, zombie movies, Romantic comedies CTA: “Continue” When Continue is pressed, show a animation: • “Analyzing your taste…” • “Finding patterns…” • “Matching similar viewers…” ## STEP 2: Show 20 recognizable titles. For each: • Poster • Title • Short descriptor • Actions Row 1- Already Watched: ❤️ Loved it 👍 Liked it 😐 Okay 👎 Didn’t Like ROW 2: ❓ Haven’t Watched ➕ Wishlist ▶ Watch Trailer ## STEP 3 — Save Taste Profile and Get final recommendations Timing Matters Do NOT ask login first. Ask after value is shown. CTA: “Save your taste profile and recommendations” Options: Email a link (magic link) This will improve conversion massively. Shows label: “Your recommendations are ready“ Show 6 titles. For each: • Poster • Title • Short descriptor • Streaming provider • Length (serie episotes, x time) | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | I have 2 LLM's with each specific job: ## 1. LLM-Prompt for enrichment layer Developer prompt: You will output a single JSON object containing a structured analysis of a movie. Grounding and scope: - Use ONLY the provided Title, Release date, and Overview/description. Do not use outside knowledge or franchise familiarity. - Prefer tags explicitly supported by the Overview/description or Title; if merely implied, keep them slightly conservative. Avoid speculation or contradictions. # Steps - Read only the given Title, Release date, and Overview/description fields for the movie. - Identify meaningful, explicitly stated or implied attributes directly supported by the overview or title. # Formatting requirements: - For each array or list in the JSON, provide 2–6 short, lowercase phrases per array. - Avoid duplicates within any array/field. - Output strictly a single JSON object—do not include any additional text, commentary, or code fences. # Notes - Strictly prohibit use of any detail not present implied in the Overview/description or title. V6_ USER PROMPT USER PROMPT: You are a film analyst. Extract grounded, concise tags from the movie metadata below, following the rules above. Use only evidence from the Overview and Title; if a detail isn’t clearly supported, omit it. Return only the JSON object with 2–6 short lowercase phrases per array. Title: {{title}} Release date: {{year}} Overview: {{description}} Output exactly this JSON shape: { "emotional_tones": [], "pacing": [], "themes": [], "storytelling_style": [] } --------------------------------- 2. LLM-prompt for FINAL/RECOMMENDATION: ## SYSTEM PROMPT You are a personalized movie recommendation assistant for streaming subscribers. Your job is to select 6 titles from a pre-filtered candidate list that best match this user's demonstrated taste. # Recommendation strategy: Learn from their top taste signals, and what the user described to enjoy and avoid. Prioritize themes, emotional tone, pacing, and storytelling style OVER genre. Never base a match solely on an avoid signal. There should always be some specific strong signals from preferred patterns, “loved” (4/4), “likes” (3/4), or user's preferred describtions. If 6 or more rated movies are in the history, apply combination-pattern reasoning in two steps before selecting: Step 1 — Derive combination patterns (internal reasoning only, do not output): Find similar tag pairs or triples that co-occur across multiple loved (4/4) and liked (3/4) films — these are the strongest signals. Identify tags that appear exclusively in disliked (1/4) films — treat these as hard avoids Flag tags that appear in both highly-rated and disliked films — these are context-dependent. Reason about which other tags they pair with in each case. Never treat a context-dependent tag as a flat avoid. Step 2 — Match candidates against combination patterns: Assess each candidate against the patterns derived in Step 1, not individual tags in isolation A candidate with one disliked tag is not automatically excluded if its overall tag combination matches a loved pattern If fewer than 6 rated movies are in the history: rely on top taste signals, and what the user stated preferences they enjoy and what the user avoid. Do not attempt combination pattern detection. REASON SPEC: 2 short sentences, second person. Referance the recommendated movie as "this" or "it" not by the title. The explanation should feel conversational, personal, and friendly, as if it was a friend who knows your movie taste. Focus on a single strong reason/pattern for the match. Sentence 1: If possible, start with connect it to a previously loved or liked movie and state the strongest reason this recommendation stood out based on the user's taste. Sentence 2: Briefly explain why that single connection makes this feel like a good fit for the users taste. Keep the explanation simple and easy to read. Prefer one clear insight over a complete explanation. Why a user liked a movie should be reflected in the profile or in repeated patterns across matadata from either the USER PREFERENCE PROFILE, HISTORY OF RATED MOVIES as well as from the recommended specific movies metadata in the CANDIDATE LIST. Avoid: - long or complex sentences - long detailed descriptions on the recommended movie. - multiple movie comparisons - phrases such as "matches your profile" or "preference pattern" - Parentheses, em dashes, semicolons # USER PROMPT> USER PREFERENCE PROFILE {{general_user_preference_profile}} HISTORY OF RATED MOVIES: {{rated_movies_history}} CANDIDATE LIST: {{candidate_movies}} TASK: Select exactly 6 highly personalized recommendations. OUTPUT: Return JSON: [{"title": "", "reason": " "}] | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | For the 1st LLM: I made these 2 evals (here just the initial task, not the whole is presented below): EVALS 1 - Relevant_checker: You are a film expert evaluator. Your job is to verify whether the details from the dynamic generated output tags (emotional_tone, pacing, themes, storytelling_style) are relevant based on the movie metadata (title and describtion). You are checking the relevance, if the final responded tags actually match the movie. (contines.....) --------- EVALS 2 - Hallucination: Fact checker System prompt: Assess the final response from an assistant in a conversation based on provided criteria, determine a classification label and determine if it is a "pass" or "fail." You are a hallucination and fact detector. Your job is to verify whether the dynamic generated output tags (emotional_tone, pacing, themes, storytelling_style) does align or support the movie metadata (title and describtion) or not. You are checking facts/hallucinations, if the final responded tags actually matches the movies metadata and provide a Classification lable. ------------------------ LLM Final-recommendation ---------------- For the 2nd LLM: I made these 2 eval: EVALS 1 — Reason matches User taste – fact check System Prompt: Eval-Hallucination: Assess the final response from an assistant in a conversation based on provided criteria, determine a classification label and determine if it is a "pass" or "fail." You are a hallucination and fact detector. Your job is to verify whether the dynamically generated recommendation REASON aligns with or is supported by the user's metadata — specifically their preference profile and their history of rated movies. You are checking facts/hallucinations: does the reason actually reflect this user's demonstrated taste, or does it reference patterns, preferences, or connections that are not grounded in the provided user data? -------- EVALS 2 - Movie Reason-fact detector EVAL 2: System prompt: Eval-Hallucination: Assess the final response from an assistant in a conversation based on provided criteria, determine a classification label and determine if it is a "pass" or "fail." You are a hallucination and fact detector. Your job is to verify whether the dynamically generated recommendation REASON aligns with or is supported by the recommended movie's metadata from the candidate list. You are checking facts/hallucinations: does the reason accurately describe something about this specific movie, or does it introduce details that have no grounding in what is known about it? | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | Inputformat is quite strict, the LLM is more based on extracted taste from users ratings, so edgecases aren't that oviously, primary are: • Weak taste profiles with fewer than 4 rated movies and 6 rated movies. (Handled by fallback strategy which both combine lower confident taste patterns with genre (in the extract-user-taste mechanism). • User with not consice taste are the difficult one, but is handled by confidensescore, as well as the LLM makes pair-pattern compatison | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | Two LLMs with distinct roles: 1) Enrichment: GPT-4.1 (temp 0.7, 1000 tokens) — structured tag extraction using function calling. Selected for high factual accuracy (93% hallucination-free in evals, 6.54/7 relevance score across 15 test cases) 2) Recommendations: gpt-5-mini via Responses API (effort: medium) — final 6-pick selection with personalised reasons. Selected for superior reasoning quality vs gpt-4o-mini which hallucinated across all test cases | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | For recommendation LLM (required): It matches the 1. users extract taste (including 2. rated movie history) with a 3. candidate list of 20 movies: 1) The users taste have been analyzed: • into structured inputfields. preferred_emotional_tones, pacing, themes, storytelling_style) • into structured inputfields. disliked_emotional_tones, pacing, themes, storytelling_style (string arrays) • details_enjoy_text, details_avoid_text (free text — is optinal but sent to the LLM) • favorite_genres (string array) 2) The users rated movie history: • rated_movies_history: loved/liked/disliked arrays with title, genres, tag arrays (max 15 total) 3) Candidate list: • candidate_movies: 20 pre-filtered movies with title, genres, tag arrays | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | Optional inputs (improve quality but not required): • details_enjoy_text / details_avoid_text: free-text user descriptions — enhance LLM reasoning if provided, ignored if empty • user_rating score per movie with loved/liked/disliked is optional how many to score - the more the better the LLMs recommendations get. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | • Reason references a specific signal from the user's taste that is correct (not hallusinated) - in the EVALS • Reason references a specific pattern from the movie that is correct (not hallusinated) - in the EVALS • Are exactly 6 movies returned? • Are all 6 titles drawn from the candidate list (no hallucinated titles)? • Does the LLM avoid recommending movies the user has already rated? • For users with disliked tags: are those tags respected without over-exclusion? • Is the JSON output parseable without errors? | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Subjective criteria (human review): • Reason feels conversational and personal — not generic or formulaic • Reason is in 2 sentences and easy to understand • Tone feels like a knowledgeable friend, not a formal recommendation system • Enrichment tags feel grounded in the movie's actual content | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | Initial enrichment prompt (V1): Plain text instruction asking for emotional_tones, pacing, themes, storytelling_style as JSON. Used gpt-4o-mini. Output was inconsistent — tags too generic, some hallucinations on lesser-known films. Initial recommendation prompt (V1): Simple instruction to pick 6 movies based on rated movies and no guidelines about language tone on the text output reasoning. LLM defaulted to highest matching_tags score, ignoring nuance. Reasons were formulaic (""matches your love of thrillers"")." | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | Enrichment iterations: - Added grounding rules ("Extract grounded, concise tags from the movie metadata below, following the rules above. Use only evidence from the Overview and Title; if a detail isn’t clearly supported, omit it.") — reduced hallucination clearly. - Used 3 different LLM-model (gpt-4.1 and gpt-4.1-mini, gpt-40-nano, and adjusted temp=1 and 0,7). -> Ended with: model to GPT-4.1 — relevance score improved from 5.2 to 6.54 (on a 7-scale ) across 15 eval cases. Recommendation iterations (did 22 iterations): -> Switched gpt-4o-mini → gpt-5-mini — eliminated 50% hallucinations that occurred in all gpt-4o-mini test cases → Added taste-extracted enrichment layer (so all profiles have structured tags on preferred_emotional_tones, pacing, themes, storytelling_style, and disliked_emotional_tones, pacing, themes, storytelling_style. -> Added combination-pattern reasoning that fixed over-exclusion of context-dependent tags → Added REASON SPEC with tone (2 sentences, second person, "this/it" rule) | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | Data sources: TMDB API with metadata: title, year, poster, description, duration, genres, streaming providers US. I used two LLM, with distint tasks. The first LLM enrichs the movie catalog from the TMDB movies-API. Backend calls GPT-4.1 to generate 4 tags per movie with movie_enrichment (emotional_tones, pacing, themes, storytelling_style). This is stored in a movie_enrichment table of ~400 movies. RAG implementation: Two-stage retrieval before LLM call: Stage 1: (get-candidates): The user's taste profile (extracted by ratings using an affinity-weighted scoring model) acts as the retrieval query. The system filters the movie catalogue (the movie_enrichment layer) by matching the user's preferred tags, excluding disliked tags and already-rated movies, and returns 20 candidates. Stage 2 (get-final-recommendations): Those 20 candidates are then passed directly into the LLM prompt as structured JSON — this is the augmentation step that makes the final recommendation specific and grounded rather than generic | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Choosing a streaming provider and genre, mabye the user writes what they like: a specific movie, a specific theme e.c. AND the most importent they rate 2-10 movies with loved, liked, disliked and mixed. And by that the backend finds 20 candidates. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Edge cases tested: • The ambiguous taste pattern and fewer than 4-6 rated movies: LLM prompt switches to simplified mode (no pattern detection) - Also same move serie duplicates: Spider-Man variants + Greenland/Greenland 2 appearing together in candidates → fixed by normalize the Title in my backend before passed to the LLM as a candidate. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | Manual review of 1 LLMs enrichmentlayer - used 15 test examples. ✗ Early enrichment (gpt-4o-mini): generic tags on the films that wasn't specific, occasional factual inaccuracies -> Switching to GPT-4.1 (effort: medium, not low) for enrichment 2. LLM for final recommendation, showed strong performance: ✓ Correct 6 titles from candidate list in all cases ✓ Reasons reference to many specific taste signals in one sentense (not generic praise) making the tone very algoritm-alike. Adjusted the prompt to get more conversational easy understandable reason I can display in the UI. -> gpt-5-mini for recommendations resolved both failure modes. | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | Enrichment eval (15 test cases, GPT-4.1): • Relevance score: 6.54/7 average (relevance_checker eval) • Hallucination rate: 7% (93% correctness on hallucination_detector eval) Recommendation eval (4 JSONL test cases, OpenAI Platform): • Eval 1 (reason matches user taste): Pass rate tracked per test case — hallucinated user preference references = Fail • Eval 2 (reason matches movie metadata): Subjective language accepted; concrete unsupported claims = Fail • gpt-4o-mini: failed multiple cases with hallucinated comparisons; replaced by gpt-5-mini | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | Edge cases tested: • The ambiguous taste pattern and fewer than 4-6 rated movies: LLM prompt switches to simplified mode (no pattern detection) | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Key prompt/system adjustments made: 1. Added grounding rules to enrichment prompt → reduced hallucination from ~40% to 7% 2. Added combination-pattern reasoning to recommendation prompt → fixed context-dependent tag over-exclusion 4. Added REASON SPEC to recommendation prompt → eliminated generic/formulaic reason text 5. Replaced gpt-4o-mini with gpt-5-mini for recommendations → eliminated all hallucination failures 6. Replaced gpt-4o-mini with GPT-4.1 for enrichment → increased relevance score +25% 7. Reduced candidate list from 30 to 20 → token reduction ~35%, faster LLM response | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | Evaluation approach: I have 2 seperat sets of evals on OpenAI Platform Evals on each LLM, . • 1. LLM - Enrichment (for the movie catalog): I have separate 15-case eval with relevance_checker and hallucination_detector scorers • 2. LLM - RecommendationLLM :4 automated grader prompts (reason-user-taste checker + reason-movie-fact checker) • For booth LLM - grader model: I use LLM-as-judge, classifying each claim as Supported/Subjective/HALLUCINATED or with a score 1-7. • Manual checklist: criteria checked per output (count, valid titles, JSON validity, no repeat ratings, etc.) | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | MVP stage: manual re-evaluation after each major prompt change or model switch. Post-launch plan: re-run automated evals on OpenAI Platform after any prompt update, model version change, or catalogue expansion (>100 new movies). Trigger conditions: hallucination rate >10%, relevance score drops below 6.0/7, or user feedback indicates poor recommendation quality. Long-term: schedule monthly eval, and expanded JSONL test set covering more user profiles and edge cases. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | MVP status: ✓ Supabase Edge Functions deployed (extract-user-taste, get-candidates, get-final-recommendations, enrich-movies) ✓ All 8 database tables created and tested ✓ OpenAI API key configured in Supabase secrets ✓ TMDB API key configured ✓ End-to-end pipeline tested with a users (STEP 1→2→3) ✗ Rate limiting and API quota monitoring not configured ✗ Rollback strategy not documented ✗ Performance testing under concurrent users not done | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | MVP / Capstone stage — solo development, no internal team. Documentation: this PRD and a product report is the primary documentation. Not yet completed: legal review of TMDB API terms for commercial use. Post-capstone: if scaled, would require engineering handover doc, privacy policy, and support channel setup before public launch. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | Phase 1 — Capstone demo (current): internal test with 5-10 real users through full flow. Manual review of recommendation quality and pipeline stability. Phase 2 — Soft launch: Could be a invite-only beta via magic link. Collect feedback via 7-day email survey. Iterate to improve | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | Current state: single-user, no concurrency testing. To scale: Supabase Edge Functions auto-scale horizontally. Primary bottleneck is OpenAI API rate limits and response time. Monitoring plan: track Edge Function execution time in Supabase logs. Catalogue scaling: enrich-movies function supports batch enrichment (up to 500/run) — The catalog should be replaced with a larger-scale catalog (more than 500+ movies) before an actual public launch. Cost scaling: gpt-5-mini per-call cost to be monitored. We could implement caching to avoid re-running LLM. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | MVP assets (capstone): • Live prototype via Lovable deployment • Product report (Word document) • AI PRD (this document) • Demo walkthrough for capstone presentation Post-launch assets needed: • Landing page copy emphasising "3 minutes to personalised picks" • FAQ covering data privacy, how recommendations work, supported streaming services • Short demo video of the 3-step flow | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | Capstone stage: updates communicated via course submission and this documentation. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | Data collected: user email (optional, magic link only), streaming service selections, genre preferences, free-text enjoy/avoid, movie ratings, taste profile tags. Storage: Supabase (PostgreSQL), EU region. No data sold or shared with third parties. User deletion: SQL deletion scripts implemented for all tables keyed by user_id. GDPR: right to deletion (by SQL quiry) | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | TMDB API: currently used under non-commercial terms. Commercial use requires separate licensing agreement — must be resolved before any kind of publich launch. OpenAI API: usage within terms of service. . Post-launch: log all LLM inputs/outputs for 30 days for audit trail. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | User metrics: • Pipeline completion rate: % reaching STEP 3 (target >60%) • Profile save rate: % signing up via magic link (target >30%) • Recommendation acceptance rate: % of 6 picks rated Loved/Liked on STEP 3 (target >5% at MVP) • User satisfaction: % reporting they stopped wasting time scrolling (target >15% via 7-day survey) Business metrics (post-launch): • Free-to-premium conversion rate (target >10% after month 1) • Monthly active users • Retention rate at 30 days | ||
| AI Metrics | How will you measure AI performance and accuracy? | Enrichment LLM (GPT-4.1): • Tag relevance score: currently 6.54/7 (target >6.0/7) • Hallucination rate: currently 7% (target <10%) Recommendation LLM (gpt-5-mini): • JSON validity rate: target 100% parseable outputs • Title hallucination rate: % of outputs containing non-candidate titles (target 0%) • Reason grounding pass rate: % passing automated eval (target >90%) • Pipeline latency: time from STEP 2 Continue to STEP 3 display (current ~30-40s, target <15s post-optimisation) | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | Implementet in-app feedback button (thumbs up/down on recommendations). MVP: no formal support channel. User can contact via email linked in the app footer. Escalation: critical pipeline failures (LLM errors, DB write failures) flagged via Supabase Edge Function error logs. Product owner (Louise) is sole owner at MVP stage. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | MVP feedback channels (planned, not yet implemented): • Thumbs up/down on overall recommendation set • Free-text feedback field on STEP 3 • 7-day email follow-up survey Triage: product owner reviews weekly. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | Current monitoring - manually checking: • Supabase Edge Function logs: all three functions log errors, LLM response status, candidate counts, and method used • recommendation_candidates_log table: per-run audit trail of 20 candidates and selection method • final_recommendations table: timestamp and user_id confirm pipeline completion Not yet implemented: alerting on error rate thresholds, dashboard for success metrics, latency monitoring. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | Post-launch improvement cycle: 1. Monthly expand the eval-dataset: re-run JSONL evals after any prompt/model change 2. Catalogue expansion: enrich 500+ new movies monthly via admin pipeline 3. User feedback analysis: review thumbs up/down patterns to identify systematic recommendation failures | ||||




