← All capstone projects

Media

Q

Built by Peter Leschenko Cohort 9 Media / smart TV content discovery

Q is a mood-first content discovery layer for smart TVs that recommends what to watch based on how one or more viewers feel, rather than forcing them through genres and app rows. Users choose who is watching and enter a mood, and Q returns exactly three titles with one-line reasons tied to that mood. The product is designed especially for shared viewing, where standard recommenders often fail to resolve conflicting preferences well.

The problem

Choosing what to watch on a smart TV fails hardest when more than one person is on the couch. Home rails, genres, and keyword search aren't mood-native, so viewers browse, veto each other's picks, scroll in silence, and eventually give up — abandoning the session for YouTube or a rewatch. Genre is not the same as mood, and history-based recommenders serve up stale "because you watched" rows that ignore how a household feels tonight. Couples end up negotiating; families face the gap between adult-tolerable and kid-safe.

The solution

Q is a mood-first content discovery layer built into the smart TV OS. Viewers choose who is watching — solo, two adults, or family — pick a vibe, and select a content type, and Q returns exactly three titles with a warm one-sentence reason tied to that mood. For shared viewing, the defining rule is overlap, not average: rather than splitting the difference into a mediocre compromise, Q finds genuine common ground that keeps everyone in the session. It is pre-installed, cross-app, D-pad friendly, and works from an explicit mood without depending on long watch history — solving cold start.

How it works

Q uses GPT-4.1 as its recommendation engine, chosen for reliably following complex structured instructions: strict JSON output, exact mood-interpretation logic, multi-person intersection rules, and hard platform constraints, while drawing on broad film and TV knowledge to suggest real, findable titles. It integrates via the OpenAI SDK with `response_format: json_object` to guarantee parseable responses, and caches results in MongoDB for 24 hours to control cost. Structured mood inputs resolve into a shared emotional territory, then a ranked set. OMDb is connected post-generation for posters and metadata as partial content grounding; a full catalog RAG over Canadian availability is planned for v2. Quality is judged on 20 fixed test cases with a ship bar of ≥80% overall pass and 100% no-hallucination.

Who it's for

The end users are CA/US smart TV viewers — adults and families — across three personas: solo viewers matching tonight's energy, couples who need to agree fast, and families balancing family-safe with adult-tolerable. The business model is B2B2C: TV OEMs license the OS and content partners pay for placement, while Q serves the viewer and delivers value back to OEMs and partners. Q sits on a mature, at-scale OS platform as a product extension, so success is measured by feature adoption, session quality, and partner outcomes rather than pre-revenue startup metrics.

Why it matters

The global smart TV market is often cited at roughly 12–17% CAGR through the late 2020s, and streaming now accounts for a majority of U.S. TV time, with roughly three in four Canadian households on connected TV. Discovery is where OS platforms lift content-partner placement and subscription revenue — and where they lose viewers to rival operating systems or short-form. By reducing browse abandonment and resolving conflicting moods on the TV itself, Q turns lost sessions into confident choices. The current build is a solo capstone prototype, honest about gaps like auth, rate limits, and a public privacy policy before any real launch.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Peter Leschenko
Your Product:CUE
Your Industry:Smart TV OS/Entertainment
Date:[Please insert the Date your started the AI PRD]
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Smart TV OS + living-room video. Cue = first-party discovery on the OS.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds: piracy/grey viewing; short-form (TikTok, Shorts, etc.); multi-app catalog fragmentation; D-pad/lean-back limits; browse-before-play abandonment; rival OS AI (Google TV, Roku, webOS, Tizen, Whale). Tailwinds: LLMs for mood→content without fine-tuning; strong NA CTV; partners want long-tail surfacing; AI as OS differentiator. Competitors: (1) OS-level discovery (2) in-app apps. Netflix (verified): Help Center documents limited opt-in beta, small member set, iPhone/iPad only, natural-language search refinable by mood/tastes (e.g. “Something funny and upbeat”) — https://help.netflix.com/en/node/349963919363172 . vs Cue: Netflix = single app, mobile, conversational; Cue = OS, TV, tiles, cross-app, pre-install — differentiate on surface + scope + modality, not “Netflix has nothing similar.”
What is the projected growth rate of your target market segment over the next 3-5 years?Global smart TV market often quoted ~12–17% CAGR late 2020s (e.g. Technavio 16.8% 2024–29; TechSci 11.6%; Grand View 13.9%). US: Nielsen May 2025 — streaming 43.8% of TV time (Mar 2025), ~+10 pts vs 2y — https://www.nielsen.com/insights/2025/connected-tv-transforming-advertising-trends/ . Canada: MTM 18+ — trade summaries cite ~76% anglophone / ~78% francophone households with CTV (Fall 2024 wave) — https://mtm-otm.ca/en/mtm18 . One-liner: ~3 in 4 CA households CTV; US viewing majority streaming.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Mature at-scale OS platform (licensing, partners, engagement ops). Cue = product extension, not pre-revenue startup. Success = feature adoption, session quality, partner outcomes.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)A: (1) Content partner placement (2) App/subscription rev share via TV store (3) OEM OS licensing. Typical TV OS stack. Cue lifts (1)(2) via engagement.
Who is your primary customer base (B2B, B2C, B2B2C)?B2B2C: OEMs license OS → consumers buy TVs; content partners B2B; Cue serves viewer, value to OEM/partners.
DifferentiatorsWhat are the key differentiators for your company?A: (1) Pre-installed on OS (2) Cross-app discovery vs one-app silo (3) Mood-first tiles → exactly 3 explained picks on TV vs history rails + vs Netflix iOS NL beta (single-app, mobile) (4) Cold start — explicit mood without depending only on long history.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?TV OEMs licensing OS; content partners (placement). Indirect: households buying TVs.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?CA/US smart TV viewers (adults/families). Daily actives drive partner metrics; churn to rival OS or phone/short-form hurts placement revenue.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?Today: home rails, genres, keyword search, continue watching, launcher — not mood-native. Cue: mood-first entry → 3 picks + reasons → faster confident choice.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)P1 Solo — Jordan, 39, CA/US metro. Goal: match tonight’s energy. Behavior: scrolls, rewatches. Cue: OS → mood+type tiles → 3 picks. P2 Couple — Sam 33 + Alex 36. Goal: agree fast. Cue: 2 mood tiles → overlap (not average) → 3 picks. P3 Family — Taylor 37, Morgan 40, kids 8+6. Goal: family-safe + adult-tolerable. Cue: family path tiles → 3 picks.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?| Stage | Action | Emotion | Pain | |-------|--------|---------|------| | 1 | Sit down, remote | Tired | — | | 2 | Home rows | Mild frustration | No shared “vibe” | | 3 | Genre pick | Overwhelmed | Genre ≠ mood | | 4 | Suggest → veto | Frustrated | Social + content friction | | 5 | Silent scroll | Resigned | Samey tiles | | 6 | YouTube / rewatch | Done | Session lost | Link---
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Browse abandonment (all, very high/high). Mood≠genre UI (all, vh/vh). Conflicting couple moods (couple, h/h). Family adult+kids gap (family, h/h). Partner negotiation fatigue (couple+family, h/med). Stale “because you watched” (solo+couple, h/med). Long-tail buried (all, med/med).
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.** **1** Mood≠genre — tiles + LLM → ranked picks + reasons (**primary**). **2** Browse abandonment — shorter path to play. **3** Stale recs — mood resets context vs history-only. **4** Long-tail — mood surfaces non-obvious titles *(guardrails: no fake titles)*.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.(1) **Free text** — good for web demo; bad TV remote v1. (2) **Mood tiles** — v1 input. (3) **Passive inference** — out (data/privacy/accuracy). (4) **Voice** — v2. (5) **Multi-mood overlap** — v1 with (2).
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.** **Multi-viewer mood tiles + emotional overlap + exactly 3 results + 1-line reasons.** D-pad only on Smart TV OS; same flow solo/couple/family; evaluable; LLM translation+reasoning without fine-tuning. v1 definition — activation, flow, AI, output, fallback? Activation: optional unobtrusive surface after ≥5 min browse, no play, no account required (cold-start story). Flow: (1) Who: Solo / Two adults / Family (+ 1b youngest age band if family) (2) Vibe: 10 adult mood tiles per cue-web-mood-model.md (+ kids 4-tile grid if family) (3) Content type: Movie / Series ep / Doc / Short / Surprise. AI: structured inputs → shared emotional territory (not average) → signals → ranked set. Output: 3 titles + platform + warm 1-sentence reason. Fallback: weak overlap → anchor lowest-energy mood. What rule governs multi-viewer recommendations? A: Overlap, not average — genuine common ground. Example couple: Relaxed + Excited → lively-but-low-stakes (e.g. Abbott-style zone), not mid compromise drama. Example family: Tired adult + upbeat kids → gentle for adult + fun for kids (e.g. Bluey-zone). Logic: tired viewer drops first — overlap keeps everyone in session. Prompt: Master prompt encodes overlap;
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?LinkPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Link
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Link
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Prompt Iterations
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Link
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Link
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Model choice: GPT-4.1 Cue uses GPT-4.1 as its recommendation engine because it reliably follows complex structured instructions — the system prompt requires strict JSON output, exact mood interpretation logic, multi-person intersection rules, and hard platform constraints. GPT-4.1 handles all of this consistently while drawing on broad film and TV knowledge to suggest real, findable titles. It integrates via the OpenAI Python SDK with response_format: json_object to guarantee parseable responses, and results are cached in MongoDB for 24 hours to minimize API costs. The model is configurable via the OPENAI_MODEL environment variable, making it straightforward to swap to a lighter model (e.g. GPT-4o-mini) or an alternative provider as needed.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Required and optional fields
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Required and optional fields
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Objective Criteria We score Cue with pass/fail checks plus 1–3 scales on 20 fixed test cases. Ship bar: ≥80% overall pass; 100% no-hallucination. Structure: Valid JSON; required keys (profiles, resolved_state, avoid, age_constraint, recommendations); exactly 3 recommendations; each with title, format, streaming_platform, and one-line reasoning. Factuality: Every title is real and findable on the named platform (no invented titles or fabricated plots). Relevance: All three picks match the mood(s) or resolved state (mood relevance 3/3); for multi-viewer, overlap serves each mood, not an average (overlap accuracy 3/3); reasoning references mood language, not generic plot summary (reasoning quality 3/3). Constraints: Canada platform allow-list only (Netflix, Crave, Prime Video, Disney+, Apple TV+, CBC Gem, Paramount+, Curiosity Stream); family age band enforced when applicable; no duplicate titles; excluded_titles honored on “Try again”; model never asks clarifying questions; always returns 3 picks.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Subjective Criteria Yes — these need human review on the eval set and in live demos: Tone — feels like a thoughtful friend, not a catalog algorithm (pass/fail). Borderline mood fit — when 2/3 titles are “almost right” (1–3 scale judgment). Overlap fairness — both viewers feel seen in picks and resolved_state.summary (couple/family cases). Compromise honesty — low-energy anchor feels fair, not like one mood was ignored (Tests 13–14). Free-text interpretation — vague phrases (“cozy but not rom-com”) interpreted sensibly. Group acceptance — would the group stop negotiating after seeing 3 picks? Discovery — at least one pick feels new vs what they already had in mind.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.Prompt Iterations
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?Prompt Iterations
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Cue v1 uses GPT-4.1 plus prompt/eval for recommendations, with OMDb connected post-generation for posters and title metadata (partial content-DB grounding). MongoDB caches responses. Full catalog RAG is planned v2: ingest OEM/partner catalog, TMDB-style metadata, and JustWatch-style availability for Canada; chunk one title per document; embed into a regional vector index; retrieve top-k using resolved_state + age/platform filters; LLM selects exactly 3 titles only from retrieved candidates to eliminate hallucinations and reflect real availability.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Typical: Solo/couple/family with 10 preset moods + optional free text; output = JSON with profiles, resolved_state, exactly 3 recommendations. Flagship cases: Test 1 (Cozy solo), Test 10 (Cozy + Excited couple), Test 15 (family). Real couple input with free text is in the file .
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Edge / negative: Tests 13, 14, 16 (strong conflict / hardest family), 18–20 (Surprise me), 19 (duplicate moods), refresh with excluded_titles, free-text-only mood, hallucination/platform violations.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Manual: Run all 20 inputs through cue-tv (GPT-4.1 + v4.1 prompt); scored mood 1–3, overlap 1–3, reasoning 1–3, hallucination pass/fail, tone pass/fail. Test 13 failed when output was pure action (ignored tired)
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Automated (v1): Script checks JSON parse, 3 recommendations, Canada platforms, no duplicate titles, excluded_titles honored. Mood/overlap/tone stay human-scored. Targets: ≥80% overall pass · 100%.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Identified: overlap averaging (13–14), family compromise (16), Surprise me breaking mood, free-text misread, refresh repeats, hallucinations. Fixes logged: v4 overlap/anchor rules · v4.1 few-shots + inline schema. Retest 13, 14, 16 after each prompt edit
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?I iterated from v3 JSON output through v4 INTERPRET→RESOLVE→RECOMMEND (overlap not average, low-energy anchor, family age hard constraint, Canada platforms, excluded_titles) to v4.1 (inline JSON schema, three few-shot user messages for couple+free text, family age 6, and refresh, plus 2–4 sentence resolved_state.summary). Changes targeted eval failures and edge cases: couple mood blending (Tests 13–14), family compromise (16), Surprise me vs mood (18–20), refresh repeats, hallucinations, and generic reasoning. Production uses GPT-4.1 with json_object; optional safety text was removed in v4.1 pending full eval results on family/sad cases.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Primary evaluation is manual against the 20-case rubric in Develop-EvalSet.md. Structural checks use API json_object plus manual review of platforms, title count, and refresh excludes. A batch Python validator is planned (schema, 3 picks, Canada allow-list, excluded_titles) to run all 20 cases against /api/v1/recommend before submission.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Prompt changes: re-run after every edit — failed IDs plus priority tests 13, 14, 16 before deploy. Model or major release: full 20-case golden set (Develop-EvalSet.md). New test data: full run when the golden set changes; otherwise run only new cases. Catalog/RAG (v2): re-baseline all 20 on each index version before ship. Post-launch: monthly sample of ~10 sessions with human rubric; immediate deep-dive on any hallucination or refresh failure; expand the golden set when new edge cases appear in prod.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Personal prototype live on Render (FastAPI + Next.js): /healthz, Pydantic validation, 30s OpenAI timeout, 24h MongoDB cache, async session logging, CORS whitelist. Documented in repo (CUE_PROJECT_SUMMARY.md). Not production-grade: no auth, rate limits, uptime alerting, or automated rollback—acceptable for a solo capstone demo.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Solo personal project—no support, legal, or comms teams. Documentation = GitHub repo, PRD markdown, master prompt, eval rubric, and technical handoff doc. Sufficient for PF review and portfolio demo; not sufficient for a public consumer product without privacy policy and accessibility work.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Open prototype URL for anyone with the link (capstone / portfolio). No accounts, no pilot cohorts, no feature flags. main branch deploys to Render on push. If this ever became a real product, I would add staging, gradual rollout, and a kill switch—out of scope for v1.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Sized for low traffic: 24h response cache, parallel poster fetch, Render free/starter + MongoDB/OMDb free tiers (~$7/mo). I monitor usage informally via Render dashboard and OpenAI billing. High traffic would need paid tiers, rate limits, warm instances, and CDN—documented as known gaps, not built yet.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?Primary asset = live demo URL + 4-min capstone video (TODO). Supporting: journey deck, PRD paste doc, GitHub README. No public FAQ or partner materials—personal project, not GTM.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Stakeholders = Product Faculty evaluators and anyone I share the demo with. Progress tracked in GitHub commits, this PRD folder, and submission pack. No internal Slack, SLA, or partner briefings.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?No accounts or PII collected. Mood tiles, content-type choice, and API responses may be logged to MongoDB for cache + debugging—no name, email, or watch history sent to the LLM. Secrets in env vars only. No published privacy policy yet (appropriate gap for a personal demo; required before public launch).
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Recommendations are read-only (no purchases, no auto-play). Safety via prompt rules: family age ceiling, Canada platform allow-list, eval bar on hallucination, optional safety addendum for sad/family cases. No formal legal review, content moderation team, or regulatory audit—conscious design choices for a capstone, not compliance-certified product.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?Capstone: demo completes end-to-end; time to 3 picks; PF eval pass rate. Aspirational (if product): time-to-first-play after Cue, % sessions where user picks a title, Try again rate, dismiss rate.
AI MetricsHow will you measure AI performance and accuracy?20-case golden eval: ≥80% overall pass, 100% no-hallucination; mood relevance, overlap (multi-viewer), reasoning 1–3; priority Tests 13, 14, 16; JSON/schema checks; latency & cache hit rate on prototype.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?No helpdesk—personal project. Support = GitHub issues (self), PF feedback, direct email for demo viewers. Escalation = owner (me); critical = API down or unsafe family output → fix prompt/deploy same day.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Manual: eval spreadsheet notes, demo observations, failed test IDs → prompt edit in Design-MasterPrompt-v1.md → retest. No in-app thumbs yet. Bugs tracked in GitHub; prompt regressions logged in develop/Develop.md.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?/healthz, Render logs, OpenAI billing dashboard, MongoDB session logs (debug). No Datadog/alerts. After prompt deploy: spot-check live URL + priority eval cases.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Re-run eval set on every prompt change; full 20 before PF submission; monthly informal prod check if URL stays public; add failed prod cases to golden set; v2 = catalog RAG + automated structural eval script.
Download the .xlsx ↓