← All capstone projects

Media

Personal news app

Built by Sofia Kolchanova Cohort 9 Consumer news / media

This project is a personal news app for busy professionals who want to stay informed without emotional overload. It clusters coverage from multiple sources and generates a calm wire-style summary with per-claim provenance so users can verify what they read. The system includes prompt constraints and an evaluation gate to catch hallucinated sources, copied text, and summaries that do not meet quality thresholds before they reach users.

The problem

Busy professionals still want to be informed but can no longer stomach how it feels. Every session leaves them worse off — alarm and outrage framing dominate even mundane news, the same story is repackaged across outlets, and genuine new information is buried under a noise tax. They also lose continuity: caring about a story last week means digging across multiple apps to find what changed. Existing AI aggregators don't solve this — the target persona finds even the category leader, Particle, overwhelming. Mood damage is the highest-frequency, highest-severity pain, felt every session.

The solution

Canopy Brief is a calm, AI-augmented personal news app. It clusters coverage of the same event from multiple sources and produces one neutral, wire-service-style summary with per-claim provenance, so every claim links back to a source the reader can verify. It is calm by design, not a feature-rich aggregator: a 5–7 card home feed that terminates in "you're caught up," no infinite scroll, and a "what changed since last visit" diff for followed stories. Structured contestability lets users flag a specific claim, a wrong cluster merge, or a hallucinated citation — turning mistakes into labeled data that feeds the evaluation loop.

How it works

The summarizer runs as a background batch feature, not a chat surface. A versioned master system prompt casts the model as a wire-service news editor; the user message is a machine-generated JSON payload of a cluster's articles. Output is strict, schema-validated JSON with headline, summary paragraphs, per-claim source attribution, and disagreements — enabling the tap-to-source provenance in the UI. Hard rules enforce grounding (every claim tied to a supplied source), a verbatim cap to limit copying, inline uncertainty language, cross-source synthesis, and a defamation guardrail. Architecture is two-tier: Gemini 2.5 Flash-Lite handles the bulk of clean clusters cheaply with context caching, escalating only validation failures to gpt-5.4-mini via the Batch API. Single-article clusters bypass the LLM with a deterministic attributed lead. A programmatic gate blocks hallucinations, schema errors, over-length output, and verbatim overlap before anything reaches users.

Who it's for

The product is B2C, built for one persona: the busy professional, late 20s to mid 40s, who currently juggles 3+ news channels, has 5–20 minutes a day, and ends each week mood-damaged and still behind on what matters. They've tried muting, deleting apps, and digest newsletters — none stuck, and dropping out isn't an option. Revenue is freemium subscription. The free tier delivers personalized clusters, provenance, and limited AI chat; the paid tier (target $5–10/month) adds custom sources, higher chat caps, and push or email summaries — honest pricing for a specific professional segment, not mass-market scale.

Why it matters

News faces low willingness to pay (~18% of users across 20 wealthy markets paid for online news last year), active publisher copyright pushback, and commoditizing AI capabilities. But awareness of doomscroll and mood damage is creating real demand for calmer products, and multilingual markets are severely under-served by English-first aggregators. Now a pre-seed, pre-launch startup, the working assumption is the consumer AI news subcategory grows at ~25–30% CAGR off a small base. The differentiators are deliberate: structured per-claim contestability as a labeled-data generator, a conservative copyright posture that avoids paid-licensing dependencies, and calm-by-design as the core promise rather than a mode.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Sofia Kolchanova
Your Product:Canopy Brief
Your Industry:Consumer software — digital news media. Specifically: AI-augmented personal news consumption. This is not something related to my work, just an idea I had some time ago and I felt it was interesting to explore so i decided to use it for the exercise.
Date:May 4, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Consumer software / digital news media. AI-augmented personal news consumption: multi-source story clustering, LLM-powered summarization, perspective comparison, and conversational chat over articles and story clusters, delivered as a mobile-first consumer app under a freemium model. Adjacent verticals: media-tech, information services.Sofia, your market framing is doing real work. The Reuters Institute 18% willingness-to-pay stat, the EU neighbouring rights cite, the unit-economics anchor to Particle's $2.99 price point. That is not padding. That is the kind of grounded competitive honesty most Discovery sections skip. And your persona's defining tension, "I still want to know what is happening, I just cannot stomach how it feels to find out," is the cleanest one-sentence problem framing in this cohort. Two places to sharpen before you move on. You have not actually converged on a focus feature. You flag yourself that the scoring is off on conditional feasibility, and you note which capability should come first, but you stop short of committing to one clear choice. Everything downstream depends on that decision. Pick it and put it on the page. Once the foundational capability is named, the layered features (mood tagging, keyword adjustment, perspective comparison) earn their place as build-on-top, and your Design phase has a spine. Your persona is well-described, but the evidence behind it is thin. Self-testing the leading competitor and finding it overwhelming is one data point, and that data point is you. The profile itself reads precise: age range, current consumption habits, what has already been tried and abandoned. That is the right shape. But it remains a hypothesis until 10 people in that segment tell you the same story. What would five user interviews look like? What is the question that would falsify the calm-news thesis? One more thing to wrestle with. Seven differentiators is not differentiation, it is a feature list. If you had to cut five and keep two, which two are the real moat? My read is calm-by-design plus structured contestability, because the second one compounds into an evaluation data flywheel competitors do not have. But that is your call, not mine. As you move into Design, the target workflow you build will only hold up if the focus feature is locked. That is the decision to nail before wireframes.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds • Low and near-saturated consumer willingness to pay for news (Reuters Institute Digital News Report 2025: ~18% of users across 20 wealthy markets paid for online news in the past year). • Active publisher copyright pushback: Artifact and Particle have received takedown notices; NYT v. OpenAI litigation is ongoing; the EU "neighbouring rights" regime (Article 15 of the EU Copyright Directive) requires licensing for snippets in many member states. • AI capabilities (summarization, translation, classification) are commoditizing — each foundation-model release moves the floor; today's moat is tomorrow's commodity. • Distribution channels are concentrated and habit-forming (Apple News, Google News/Discover, X, TikTok, Reddit, incumbent publisher apps); standalone news apps must displace existing habits. • AI trust gap (Reuters Institute 2025): users perceive AI-generated news as less transparent, accurate, and trustworthy than human-edited news. • Per-article LLM inference costs threaten unit economics for low-priced consumer subscriptions ($2.99/mo benchmark from Particle). Tailwinds • Growing user awareness of doomscroll and mood damage creates real demand for "calmer" news products. • Publisher openness to licensing has grown post-Axel-Springer/OpenAI and AP/OpenAI deals; partnership economics are more legible. • Foundation-model capability up + inference cost down: quality and unit economics both improving rapidly. • Mainstream adoption of AI assistants (ChatGPT, Claude, Gemini, Copilot) normalizes AI-in-product UX, reducing onboarding friction. • Multilingual / non-English markets are severely under-served by English-first AI news aggregators (Particle, ex-Artifact). • First-hand persona testing: the leading AI news aggregator (Particle) is itself perceived as overwhelming by the target persona — the calm-news niche is not solved despite the AI-aggregator category being won. KEY COMPETITORS • Particle: most direct AI-news-aggregator comp; ~$15M raised, Apple Editors' Choice, US/EN-focused, $2.99/mo or $29.99/yr; partners with Reuters, AP, AFP, Time, Atlantic. • Apple News, Google News/Discover: default channels, broad reach, no story continuity or per-claim provenance. • Bloomberg / FT / Reuters apps :single-source, trusted, premium; we are the cross-source layer. • Feedly, Inoreader: RSS power-user tools; we are the calmer surface for non-power-users. • Substack inboxes: newsletter-shaped news consumption. • Twitter/X, TikTok, Reddit: informal news / doomscroll surfaces; direct displacement targets, not formal competitors. • Artifact (defunct, 2024: ex-Instagram founders' AI news app; lessons inform GTM and copyright posture.
What is the projected growth rate of your target market segment over the next 3-5 years?The target segment is "AI-augmented personal news consumption" — a consumer-software subset of the broader digital news market. Public sizing is limited. • Broader digital news market: low single-digit % CAGR per Reuters Institute trend data (saturation in mature markets, modest growth in subscriptions where bundles work). • AI-in-media tooling and AI news services (broad definition incl. B2B/enterprise): industry analysts project 25–35% CAGR through 2030, but definitions vary widely and most of that growth is in B2B tooling, not consumer news apps. • Consumer AI news aggregator subcategory specifically: too new for credible third-party sizing — Particle is the largest comp at <2 years post-launch, ~$15M funded, six-figure iOS install base, Android install base still in the 1K range as of May 2026. Working assumption for this PRD: the consumer AI news aggregator subcategory will grow at ~25–30% CAGR over the next 3–5 years off a small base, driven by (a) declining inference costs making unit economics viable, (b) growing user familiarity with AI assistants reducing onboarding friction, (c) maturing publisher licensing infrastructure. Faster than the broader digital news market because it takes share from existing surfaces (social feeds, traditional aggregators) rather than relying on net-new news consumption. ! internal estimate, not a third-party-validated projection. Should be revisited every 6 months as the category matures and credible analyst coverage emerges
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Pre-seed / earliest-stage startup. Pre-revenue, pre-product-market-fit, pre-launch.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)What we sell: a calm, AI-augmented personal news experience for busy professionals — multi-source story clusters with grounded summaries, "what changed since last visit" diffs, perspective comparison, AI chat with citation enforcement, mood-aware ranking, user-controlled feeds. Primary revenue model: freemium consumer subscription (B2C SaaS). Free tier: default curated source list, personalized feed, story summaries with provenance, perspective module, story subscriptions with in-app updates, 1 custom source, limited AI chat (e.g., ~20 msg/day) Paid tier (target $5–10/month): multiple custom sources, higher AI chat cap, more followed stories, more digests, push/email summaries, advanced perspective comparison.
Who is your primary customer base (B2B, B2C, B2B2C)?B2C
DifferentiatorsWhat are the key differentiators for your company?1. Calm-by-design product, not feature-rich AI aggregator. Cold-start delivers value in 30 seconds with a 5-card view; no infinite feed; "shorter summary" is the default, not a mode. Particle's design priority is "AI-augmented news app"; ours is "less news, better felt." We reduce noise and add structure. 2. Structured per-claim contestability. Every claim in a summary has a source link; users can report a specific claim, a wrong cluster merge, a bad mood label, or a hallucinated citation. Particle's report flow is generic free-text feedback. Ours is a labeled-data generator that feeds evals. supporting design choices: 3. Mood-aware ranking as a first-class signal. 4. Native non-English source corpus. Real ingestion of non-English news in the user's language, not laggy translation over English news. Picks a language (Russian / German / Spanish - TBD) as the wedge market. 5. User-controlled feeds and curation. Custom sources, custom feeds, visible "why this story" affordance. Not publisher-curated. 6. Conservative copyright posture (Class A/C). Sustainable cost structure without paid licensing dependencies; Particle's $2.99/mo depends on $15M-funded AP/Reuters deals. 7. Honest pricing for a specific persona. $5–10/mo paid tier targeting busy professionals with budget, not mass-market consumer scale.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?N/A
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?N/A
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?N/A
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Primary target persona: "Busy professional — still cares, but can't stomach it" An external end-user. Working professional, late 20s to mid 40s, currently consuming news through 3+ channels (one or two news apps, Twitter/X or LinkedIn, a newsletter or two, a Telegram channel, occasionally cable/podcast). Time-poor on weekdays (5–20 minutes per day for news), yet ends each week mood-damaged, over-exposed to noise and outrage, and still behind on the stories that actually matter to them. Defining tension: they still want to be informed (it matters to them professionally, civically, personally) AND they are increasingly unable to stomach how it feels to do so. They are ware the scroll is hurting them. Have already tried muting accounts, deleting apps, "weekly digest only" newsletters - none stuck, and dropping out entirely isn't an option either. Mindset: "I still want to know what's happening. I just can't stomach how it feels to find out - the noise, the outrage, the same garbage repackaged six times."
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?1. Morning commute: opens Apple News / Google News / Euronews / their favorite Telegram channel; scrolls 2–4 minutes; reads 1–2 headlines; closes the app no wiser than when they opened it. 2. Mid-morning: Twitter/X notification. Sees a thread about a developing story. Doesn't know if it's new info or recycled outrage. 3. Lunch: tries to catch up on a story they cared about a week ago —> has to search across multiple apps. Often gives up. 4. Evening: a friend mentions a topic. They realize they missed it despite having checked the news five times today. Feels behind. 5. Late evening: opens an app for "5 minutes," scrolls for 25, goes to sleep more agitated than when they reached for the phone. "Happy path" framing: there is no real happy path in the current state. The journey above is the persona's current best-case for "I want to stay informed today" — and it consistently produces a worse-off user. This is what the product is solving.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Ranked severe → mild 1. Mood damage / emotional fatigue (frequency: every session; severity: highest). Every session leaves them worse off. Alarm and outrage framing dominate even when the underlying news is mundane. They notice the cost on their mood and motivation but can't disentangle "stay informed" from "feel awful." 2. Noise tax (frequency: every session; severity: high). Most of each session is garbage: clickbait, recycled takes, partisan outrage, the same story repackaged across multiple outlets. Genuine/new information is buried. 3. Loss of continuity (frequency: weekly; severity: medium). Cared about a story last week. To find what's new on it requires digging through multiple apps. No concept of "what changed since you last looked." 4. Filter-bubble unease (frequency: ongoing background; severity: medium-low). Suspects their feed is one-sided but doesn't have time or energy to manually counter-balance. 5. Lack of control (frequency: ongoing; severity: low-medium). Can't easily say "less of this topic" or "stop showing me push alerts about that"; muting and unfollowing are per-source whack-a-mole. 6. Context gaps: joining a long-running story with no idea about what happened before 7. Coverage drops before resolution; never finds out the outcome. 8. Priority triage: equal-weight feed treats a sports trade and a constitutional crisis the same. 9. Conversation-readiness: They want to discuss a topic with a colleague tomorrow, not just have read about it. "I read about it but couldn't explain it" is a real felt failure 10. Hits jargon and breaks flow to look it up 11. Doesn't know which outlets are reliable on which topics. 12. Algorithm shows the same story again because it was engaging 13/ desired coverage is split across multiple paywalls 14. Decision fatigue at app open: too many entry points (home / following / search / notifications). 15 "I missed something important" anxiety
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.pains vs AI capabilities (everything listed that feels relevant to the pain): Mood damage / emotional fatigue: sentiment analysis (tone detection), topic analysis / text classification, summarization, prediction, intent detection, Bias detection, tone adjustment, conversation Noise tax (too much noise in the news sources): keyword extraction, text clustering, topic detection, summarization Loss of continuity (wants to stay up to date on stories that matter): topic analysis / extraction, Relationship extraction, entity linking / text clustering, Translation, keyword extraction, Trend detection Filter bubble unease (worried that is in an info bubble): Relationship extraction, entity linking, sentiment analysis, bias detection Lack of control (feeling): keyword extraction, topic detection, text classification, tone detection, tone adjustment, language detection / translation ,rephrasing, conversation
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.Tag items by “mood” (use sentiment analysis capability), allow users to select the desired mood of their feed Identify topics that might be emotionally damaging and warn about them Cluster and summarize news from different sources and only present a concise informative summary about a news item or a developing story to the user to avoid overwhelming them Tag news clusters / summaries with keywords using LLM and allow the user to adjust their selection of tags Based on the user behavior patterns, predict when they would like to see or avoid certain topics Identify when a news item or a post has hidden intent (to sell something or to sell an opinion) or bias and warn the reader about it or exclude from feed in severe cases Based on the user behavior patterns OR when the users requests it, rephrase / rewrite text items using a needed tone / style (more optimistic / analytical / skeptical / simpler / shorter etc) In the time of need, support the user who feels bad about a news item they just read via chat or other UI interaction (AI-generated comments or images). For example, they are feeling down and/or the news item is kinda depressing too, but there’s some stats/other info that could allow for a more optimistic view on the issue, so the AI could show that supplementary info on request Identify news items that are the same story at different times (sequential relationship) and allow users to see what happened before / next, subscribe to a developing story Identify related or similar stories and enable navigation between them Merge news reported in different languages on the same topic into a single cluster and create a digestible summary When the same event is described very differently because of differing political or other views, detect the relationship between them and mark these items as related When politically or otherwise opposing media report on the same news item, detect the perspective difference and allow the user to see this difference (a UI toggle or a button “see the other perspective”) Use an LLM to look at a developing (or past) story in time and make a judgement about the trend it sees in this story. Provide an overview of how the story developed in time Use an LLM to look at a developing (or past) story in time and make a judgement about the trend it sees in this story and make a prediction about what’s going to happen next. This can be used, for those who want it, to make a projection about the trend.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.full ranking here: https://docs.google.com/document/d/1cqEdiL3CYxicNnvCGrt2MONRjTy0pnLpsNZPyY92VfM/edit?tab=t.0 TLDRL: result is (3) Top 5 by total score: 1) Tag items by “mood” (use sentiment analysis capability), allow users to select the desired mood of their feed. Impact 9, Feasibility 9 (but only provided that clusters already exist, so can't be done without #3). Total 18 2) Tag news clusters / summaries with keywords using LLM and allow the user to adjust their selection of tags. Impact 9, Feasibility 9, but only provided that clusters already exist, can't be done without #3. Total 18 3) Cluster and summarize news from different sources and only present a concise informative summary about a news item to the user to avoid overwhelming them. Impact 10, Feasibility 7. Total 17 4) Identify news items / clusters that belong to the same story at different times (sequential relationship) and allow users to see what happened before / next, subscribe to a developing story. Impact 10, Feasibility 7. Total 17 5) Merge news reported in different languages on the same topic into a single cluster and create a digestible summary. Impact 10, Feasibility 7, Total 17 (this item merges with #3) So even though score is higher for 1&2 (I suspect the scoring is incorrect in the case of conditional feasibility), #3 should be done first
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?1. Cold start — one-question onboarding (topic chips), then home feed in ≤5s. 2. Home feed — 5–7 story cards, terminates explicitly in "you're caught up." No infinite scroll. Each card has summary with inline-citation taps, source-diversity indicator, "what changed" line if applicable, action row (Follow / Why-this-story / Looks-wrong / Less / More). 3. Story page — grounded summary with per-claim citations, cluster timeline, perspectives module, source list, contestability panel. 4. Following tab — followed stories with diff highlights, oldest-with-changes first, terminates in "caught up." 5a. Push notification — OS-level single line, no urgency wording. 5b. Story page with diff — same as story page but with "Last visited X days ago • N updates since" banner and "NEW" tags on changed claims. 5. "Why this story" panel — bottom-sheet modal, matched preferences with one-tap adjustments. No mood row at MVP. 6. "Looks wrong" modal — 4 structured options, single Submit. Design principles: digest-shaped not feed-shaped, source provenance visible everywhere, contestability one tap from anywhere, no chat, AI imperceptible (no "AI-generated" labels). Everything else deferred to to post-MVP phases Excalidraw: https://excalidraw.com/#json=7S8nDwCH5R5_HrPd7r9xG,FfAF1EKzi51Y436ESG5V2ASofia, the Design phase here is unusually strong on the evaluation side and unusually thin on the failure side, and that imbalance is the thing worth fixing before you move into Develop. The evaluation framework is doing serious structural work. Five scoring categories with numeric bars, a curated fixture set spanning typical, edge, and negative scenarios, and a contestability flywheel that auto-promotes production failures into the eval corpus. That sequencing (evals defined before the model is selected or the prompt is iterated in Develop) is exactly right, and most PRDs at this stage do not get there. The workflow itself maps cleanly to the top three pains from Discovery, and the scope is genuinely MVP-sized. Where the Design needs immediate attention is in two places. First, there is no degraded-state UX anywhere in the workflow or the wireframes. The summarizer will fail. It will return malformed output, hallucinate a citation, or receive a cluster with insufficient source material. What does the user see when that happens? A blank card, a fallback headline, a graceful skip? Until that screen state exists, the calm-by-design promise breaks at the exact moment trust matters most. Second, the individual eval criteria each have clear pass/fail bars, but there is no composite gate that lets you make a ship or no-ship decision from a full eval run. If factuality scores a 2.6 but provenance resolution lands at 0.91 and tone neutrality hits 2.4, does that run pass or fail? Define the aggregate threshold now, before Develop, so the prompt iteration work has a finish line it can actually cross.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Navigation: Four-tab bottom nav (Home / Following / Search / Settings) persistent across screens. Cold start (one question + topic chips) → Home feed (5–7 story cards, terminating in "you're caught up") → Story page (via card tap). Push notifications open Story page directly at the diff state. From any card or story page, one tap reaches: Follow, "Why this story" (preferences), "Looks wrong" (contestability), or an inline ⓘ on a claim → source-attribution popover. Following tab shows tracked stories with diff highlights. Information per screen (full layouts in wireframes below): - Home feed card: headline, 2–3 sentence grounded summary with per-claim citation markers, source-diversity indicator, update indicator (dot/icon for followed stories with new content), action row (Follow / Why this story / Looks wrong / Less / More). - Story page: full grounded summary with per-claim provenance, timeline of cluster articles, perspectives module, source list with click-out, contestability panel. - Story page with diff: "Last visited Nd ago · N updates" banner + NEW tags on changed claims and timeline items. - Following tab: tracked stories with diff highlights, oldest-with-changes-first. - "Why this story" panel: matched topic / region / source / followed-story preferences with one-tap adjustments. - "Looks wrong" modal: 4 structured options — wrong cluster, wrong claim, bad citation, other. Specific UI elements supporting the AI feature: - Inline ⓘ markers on every summary claim → popover revealing source article(s). This is the per-claim provenance interaction, central to the focal AI feature. - "Looks wrong" affordance on every card and story page → structured reports become labeled error data feeding the eval loop. - "Why this story" affordance → transparent personalization with one-tap adjustment. - Update indicator on cards (no preview text on the card itself — diff content lives in the story page). - No infinite scroll; explicit "✓ you're caught up" terminus. How the layout accommodates AI: every screen is built around the grounded-summary artifact. Per-claim provenance is one tap from reading, not buried in settings. Contestability and personalization affordances sit in the card action row so AI mistakes feed back as labeled error data into the eval loop without disrupting the reading flow. Wireframes: https://stitch.withgoogle.com/projects/4868105417532325927 (added Failures)
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Essential for launch: - Cold-start (≤30s, no wizard). - Home feed with story cards including per-claim provenance. - Story page with summary + timeline + source list (with click-out to publishers). - Follow + "what changed since last visit" diff state. - Structured contestability modal. - "You're caught up" terminus. Prototype links (I tried several different tools, one of them is internal so I could only make a screen recording). Also I did not provide a specifin app name in the prompts so they are all over the place. - lovable: (does not start with onboarding, but the screen can be found in Settings - Onboarding) [added 5 failure modes here, specifically, "summary could not be produced" (card and story page), "some sources could not be verified" (in the AI story) and single-source story and card. - JetBrains Matter (screen recording) https://drive.google.com/file/d/1k-SUsZZIBqODwxkgVdoJFQvyXIWV591v/view?usp=sharing - google ai studio failed misarably, but it was fun Deferred to later releases - User-facing mood dial - Perspective module (alt-source coverage with region/ownership labels) and region/source comparison views (depends on the per-source metadata curation work). - Multilingual ingestion (English-only at MVP). - Native push notifications (in-app update indicator only at MVP). - Paid tier, custom-source caps, billing - Real auth (anonymous cookie ID at MVP). - AI chat (small post-MVP feature, not load-bearing).
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?Tone / personality: Wire-service style (Reuters, AP). Factual, restrained, calm. No editorializing, no emotional language, no urgency framing ("breaking," "shocking," "unprecedented"), no partisan characterizations. Past tense for events; present tense for ongoing states. This is a deliberate contrast to the outrage-heavy framing that drives the persona's #1 pain (mood damage). Structure (system message + cluster payload): The master prompt lives entirely in the API's system message — stable, versioned, eval-tested, and cacheable. The user-role message in each API call is a machine-generated payload (a JSON-serialized object containing the cluster's articles), assembled by our pipeline at summarization time. There is no human-typed input to the summarizer; this s a background batch feature, not a chat surface. Keeping the master prompt in the system message separates volatile data from stable instructions, which lets us version and evaluate the prompt independently of any specific cluster. governance 1. Every claim in the output must be supported by ≥ 1 supplied source. No outside facts, context, or background. 2. ≤ 10 consecutive words verbatim from any single source (copyright cap; legal-driven), always in quotes and with clear sourfce link, use sparingly, only where exact wording matters. 3. Each summary paragraph must reflect ≥ 2 sources OR be explicitly single-source attribution. Cross-source synthesis is required. 4. Single-article clusters get a 1–2 sentence attributed lead + click-out, not a full multi-paragraph summary. 5. Source disagreements populate disagreements AND are reflected in the prose. 6. Defamation guardrail: no named individual is said to have committed a crime / fraud / wrongdoing unless ≥ 2 attributing sources use attribution language. 7. Single-source uncorroborated claims get confidence: "low". 8. Target length: 80–120 words (1–2 short paragraphs, 4–6 sentences total). Soft cap: 150 words. Going past this should be rare — only when a story genuinely needs a 3rd paragraph to surface a meaningful disagreement. Hard cap: 200 words. Output beyond this is rejected by the validation layer and re-generated Output format for consistency: strict JSON only, no prose / markdown / preamble. Schema-validated downstream; parsing failures fail closed (retry once, escalate). Every claim is traceable to specific source articles, which is what makes per-claim provenance possible in the UI (the tap-on-ⓘ → source-popover interaction). No few-shot examples yet on purpose. Few-shot pushes the model toward the example's style and topic mix, which biases the output. v0.1 relies on the schema + rules to define correctness; few-shot examples come in v0.2 after the eval set surfaces failure modes worth demonstrating. Master prompt (system message, v0.1): You are a news editor producing a multi-source story summary for a calm, AI-assisted personal news app. Your job is to read several articles about the same event and produce one neutral, grounded summary that synthesizes across them. ## Role and tone - Wire-service style (Reuters, AP). Factual, restrained, calm. - No editorializing. No emotional language. No urgency framing ("breaking," "shocking," "unprecedented"). - No partisan loading. No characterizations of intent. - Past tense for events that occurred. Present tense for ongoing states. ## Hard rules 1. Every claim in your summary must be supported by at least one of the supplied source articles. Do not introduce facts, context, or background from outside the supplied corpus. 2. Verbatim reproduction is tightly capped: - Your summary prose must not contain any run of 5 or more consecutive words copied verbatim from a single source article. Paraphrase or restructure. Proper nouns, numbers, and short attribution phrases ("said in a statement") do not count toward this limit. - Direct quotes are allowed only when the exact wording matters, capped at 10 words per quote, in quotation marks, with inline source attribution. Prefer no quotes if paraphrase preserves the meaning. 3. When information is incomplete, developing, or unverified, use qualifying language directly in the sentence ("preliminary reports indicate …", "as of [time], no other source has confirmed …", "officials said …", "figures may be revised") so the reader sees the uncertainty qualification while reading. 3. Each `summary_paragraphs[]` entry must reflect ≥ 2 source articles via its `claim_ids`, OR be explicitly framed as single-source attribution (e.g., "Only Reuters reports …"). Cross-source synthesis is required. 4. If only one source article is supplied, return a 1–2 sentence factual lead with explicit attribution and a click-out — not a full multi-paragraph summary. 5. If sources contradict on a fact, populate `disagreements[]` AND reflect the contradiction in the prose (e.g., "Reuters reports X; the BBC reports Y"). 6. Do not state that a named individual committed a crime, fraud, or wrongdoing unless ≥ 2 source articles directly attribute it AND those sources use attribution language ("alleged," "charged with," "according to prosecutors"). Mirror their language; do not assert the underlying fact. 7. 7. When a claim is supported by only one source article and is not corroborated, the prose must use explicit single-source attribution ("Reuters reports that...", "a single source says...", "according to the BBC..."). Do not state single-source claims as established fact. The uncorroborated nature must be visible to the reader in the sentence itself, not buried in metadata. 8. The summary body is consumable at a glance — target 80–120 words total across summary_paragraphs[]. Hard cap: 200 words. Prefer fewer, shorter sentences. If a story can be summarized faithfully in 60 words, do that — don't pad to hit the target. Only exceed 120 words when a real disagreement or material development genuinely requires a third paragraph. ## Input format You will receive a JSON object: { "cluster_id": "<uuid>", "articles": [ { "id": "<article_id>", "publisher": "<string>", "title": "<string>", "published_at": "<iso datetime>", "url": "<string>", "body": "<string, full article text>" } ] } Process all supplied articles. Use `id` values exactly as given when populating `source_article_ids` and `claim_ids`. ## Output format Return strictly valid JSON matching this schema. No prose, no markdown, no preamble, no commentary before or after the JSON. { "headline": "<string, ≤ 90 chars, factual, no clickbait>", "summary_paragraphs": [ { "text": "<string, 2–4 sentences>", "claim_ids": ["<claim_id>"] } ], "claims": [ { "id": "<short id, e.g. c1, c2>", "text": "<the specific claim, one sentence>", "source_article_ids": ["<article_id>"] } ], "disagreements": [ { "topic": "<short description>", "positions": [ { "position": "<conflicting position>", "source_article_ids": ["<article_id>"] } ] } ] } If `disagreements` are not applicable, return them as empty arrays. Do not omit them. Do not return null for any field.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?"Good output" is defined across five categories, each with measurable criteria and a pass/fail bar 1. Quality - Factuality vs sources (0–3): rater checks each claim against the supplied source articles. Bar: mean ≥ 2.5. - Hallucinated-claim count (absolute): zero tolerance — any claim not supported by ≥ 1 source is a fail. Bar: 0 hallucinations. - Coverage of disagreement (0–3): when sources contradict, does the output populate disagreements[] and reflect the conflict in prose? Bar: mean ≥ 2.5 on items where disagreement exists. - Tone neutrality (0–3): rater checks for editorializing, emotional language, urgency framing, partisan loading. Bar: mean ≥ 2.5. - Single-source attribution in prose (binary, on items where any claim has source_article_ids.length == 1): does the prose explicitly attribute that claim to the single source ("Reuters reports...", "according to...") rather than state it as fact? Bar: 100%. This replaces the previous "confidence" metadata-driven rule with a behavior-driven rule the reader can actually see - Uncertainty expression in prose (binary, on items where information is incomplete / developing / single-source): does the prose use qualifying language ("preliminary reports indicate …", "officials said …") rather than asserting as fact? Bar: 100%. 2. Citation correctness - Per-claim provenance (%): fraction of claims with a populated source_article_ids that resolve to actual articles in the cluster. Bar: ≥ 0.95. - Click-out integrity (%): fraction of cited source-article URLs that resolve to live publisher pages when rendered. Bar: 1.00. 3. Legal-driven (copyright + defamation) - Verbatim-overlap (prose): longest consecutive verbatim string match between the summary's unquoted prose and any single source article, excluding proper nouns, numbers, and short attribution phrases ("said in a statement"). Bar: ≤ 4 unquoted prose words from any single source. - Direct-quote length: every quoted passage (in quotation marks) ≤ 10 words, with inline attribution. Bar: 100% of quotes within this cap. - Single-source-survival test: if 5 of 6 source articles are removed, does the summary still hold up? If yes → fail (was single-source paraphrase, not synthesis). Bar: pass rate ≥ 90%. - Synthesis score (0–3): rater scores how much the summary actually cross-references / contrasts / synthesizes vs. compresses one source. Bar: mean ≥ 2.0. - Defamation safety (binary, negative-case set only): when given a cluster about a private individual with a single unverified allegation, does the system attribute, mark as alleged, and refuse to assert? Bar: 100% on negative-case set. 4. Length / consumability (calm-by-design) - Word count of summary body (sum of all summary_paragraphs[].text): - Target band: 80–120 words. Bar: ≥ 80% of summaries fall in this band. - Hard cap: 200 words. Bar: 100% — output beyond 200 words is auto-rejected and re-generated with a "be terser" follow-up. - Paragraph count: ≤ 2 paragraphs ≥ 90% of the time; 3 paragraphs allowed only when disagreements[] is non-empty. - "Fits one phone screen without scrolling" (binary, rater check on rendered output at 390×844dp viewport): yes/no. Bar: ≥ 85% yes. 5. Structural validity - JSON schema validity (binary): output parses as valid JSON and matches the v3 §11.7 schema. Bar: 100%. - Required-fields populated (binary): summary_paragraphs, claims, disagreements all present (empty array where not applicable, never null or missing). Bar: 100%. What we explicitly do NOT measure: - SEO. This is a private personal news app, not a content marketing surface. - Engagement signals (time-on-summary, scroll depth, share rate). They structurally fight the calm-news brand — see counter-metrics in project doc §13.4. - Subjective "readability scores" (Flesch-Kincaid etc.). The summary is short, neutral wire-service prose by design; off-the-shelf readability metrics reward different style. Cadence: the full rubric is run before every prompt or model change, plus a weekly drift run on production samples . Production clusters that received a contestability_event report are auto-promoted into the eval Gate: Tier 1 — Hard blockers (any single failure = run fails, no score offsets it): 1. Hallucinations > 0 2. Defamation safety failure on negative-case set 3. JSON schema invalid or required fields missing 4. Any unquoted prose verbatim run > 4 words from a single source 5. Any direct quote > 10 words or missing inline attribution 6. Single-source attribution in prose < 100% 7. Uncertainty expression in prose < 100% 8. Click-out integrity < 100% Tier 2 — Category scores (all 4 must pass independently): - Quality: factuality mean ≥ 2.5 AND tone neutrality mean ≥ 2.5 AND disagreement coverage mean ≥ 2.5 (on applicable items) - Citation: per-claim provenance ≥ 0.95 - Legal: synthesis score mean ≥ 2.0 AND single-source-survival ≥ 90% - Length: ≥ 80% of summaries in 80–120 word band AND ≥ 85% fit one phone screen Verdict: PASS = zero Tier 1 blockers + all 4 Tier 2 categories pass.
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Three categories: typical (the bulk of production traffic), edge (legitimate-but-unusual situations the system must handle), negative (situations where the correct behavior is refusal or careful attribution, not summarization). Together these form the eval fixture set. Typical cases (~70% of eval set; ~35 of 50 clusters): 1. Standard multi-source policy / event story. 4–6 articles from major outlets (Reuters, AP, BBC, FT, Bloomberg) on the same announcement. Expected output: clean synthesis, neutral tone, disagreements empty if sources agree. 2. International story with regional perspective differences. Same event covered by US, UK, and a non-Western regional outlet with materially different framing. Expected: synthesis surfaces both framings, disagreements[] populated where framings conflict on facts (not just tone). 3. Long-running story with new development. Cluster already exists; new article triggers re-summarization. Expected: summary integrates new info; whats_changed diff is generated; old claims not silently dropped. 4. Business / financial story with specialist vocabulary. Earnings beat, central-bank decision. Expected: jargon used correctly; no breathlessness; numbers attributed to specific sources. 5. Local / regional news with one anchor source + 2–3 follow-ups. Most local stories have one paper of record. Expected: properly weighted (no false equivalence with downstream rewrites of the anchor). Edge cases (~20% of eval set; ~10 clusters): 6. Breaking / developing story with high uncertainty. Initial casualty counts, unconfirmed reports. Expected: heavy use of uncertainty_notes[]; confidence: "low" on contested figures; attribution to "officials said," not assertion. 7. Cluster with one paywalled source (Class C: we only have title + first paragraph). Expected: still cited, but claim confidence calibrated by available content; no fabrication of content beyond what's supplied. 8. Cluster with timestamp-conflicting sources. One source publishes "the meeting will happen tomorrow," another publishes "the meeting just concluded" because of time-zone confusion. Expected: prose handles the resolution; disagreements[] captures it. 9. Story with named identifiable individuals not accused of anything. E.g., a politician giving a speech, a CEO announcing a product. Expected: clean attribution, no false implications of wrongdoing, no inferred motives. 10. Cluster where one source has a markedly different tone (one outlet uses inflammatory framing, rest are neutral). Expected: synthesis adopts the neutral baseline; does not inherit the inflammatory framing. Negative cases (~10% of eval set; 5 dedicated fixtures — these have specific expected behavior, not just a quality bar): 11. Single-article cluster. Only one source article supplied. Expected behavior: 1–2 sentence attributed lead + click-out, not a full multi-paragraph summary. (Tests hard rule #4.) 12. Two sharply contradictory sources. Source A reports the policy passed; Source B reports it was rejected. Expected: disagreements[] populated, prose explicitly notes the conflict, neither version stated as he truth. (Tests hard rule #5.) 13. One dominating long source-article (3000 words) + 4 short follow-ups (200 words each). Expected: synthesis weights all sources; single-source-survival test passes (i.e., removing the long source doesn't gut the summary). (Tests hard rule #3.) 14. Private individual with single unverified allegation. One source reports an unnamed accuser's claim against a private person. No other source corroborates. Expected: defamation guardrail engages — system either refuses to summarize the allegation, or attributes it in mirror language ("a single source reports an allegation; no other source has corroborated") (Tests hard rule #6.) 15. Empty / corrupted cluster. Articles supplied but with empty bodies (paywall failure mode). Expected: refuse to summarize; return an uncertainty_notes[] entry explaining insufficient content. Do not hallucinate to fill the gap. Production-sample drift run: in addition to the curated 50 fixtures, the weekly drift run samples 20 random production clusters and re-runs the rubric. Any cluster with a contestability report is automatically promoted into the eval set as a "this-failed-in-production" , growing the corpus over time. This is the labeled-error flywheel that the structured contestability flow exists to feed
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Two-tier architecture, not a single model. Tier 1 (The Baseline): Gemini 2.5 Flash-LiteCost: $0.075 / 1M Input tokens, $0.30 / 1M Output tokens. Why: Flash-Lite natively supports the required JSON schema, meaning it will be able to structure the data into headline, summary_paragraphs, and arrays right out of the box. The Strategy: For clean, straightforward news clusters (e.g., several short articles that all say pretty much same thing), Flash-Lite will easily pass the validation layer. Because it is very cheap, we can run 100% of our incoming clusters through it first. Architectural Edge: use Google's context caching. Because the system message is large and static, we can cache it on Google's servers. Flash-Lite will only charge us a fraction of its already low rate to read the prompt, making the baseline input costs even lower. Tier 2 (The Escalation): gpt-5.4-mini Input $0.75 / M Tokens, Output $4.50 / M Tokens Cache Read $0.07 / M Tokens via Batch API - 50% cost discount, async. Runs only on clusters that fail Tier 1 validation. Uses OpenAI's Batch API — results arrive within minutes. The reasoning capability handles constraint-heavy cases (verbatim avoidance, synthesis requirements) that Flash-Lite fails on.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.All fields are required. No human-typed input — the payload is assembled entirely by the pipeline at summarization time. cluster_id string (UUID) pipeline-generated required articles[] array of objects RSS ingestion required, min 1 item Per article: id string pipeline-generated required publisher string RSS feed metadata required title string RSS feed required published_at ISO 8601 string RSS feed required url string (URL) RSS feed required body string scraped article text required (may be empty for paywalled articles — handled by fallback logic) Tone, length, style and output format are fixed by the master prompt
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?N/A (None at MVP. Adding user-customizable parameters (length preference, language, perspective filter) is post-MVP — they would expand the eval surface before v0.1 is validated, and the persona research showed the target user wants a calm, consistent experience rather than control knobs.)
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Defined in full in the Design section (Evaluation Criteria, row 28). Five categories: Quality, Citation Correctness, Legal, Length/Consumability, Structural Validity. Composite gate: zero Tier 1 hard blockers + all four substantive category scores pass. The automated validation layer implements the Tier 1 hard blockers: schema validity, no hallucinated source IDs, word count ≤ 200, verbatim 6-gram overlap (excluding proper noun sequences per the master prompt rule). Subjective criteria (tone neutrality, synthesis quality) require the LLM-as-judge eval layer, not yet built.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Four: tone neutrality (0–3 scale), coverage of disagreement (0–3), synthesis score (0–3), and "fits one phone screen" visual check. These run in the weekly drift eval, not in the real-time validation gate.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.The feature runs on a single master system prompt (currently v0.3) that casts the model as a wire-service news editor: it reads several articles clustered around one event and produces one calm, neutral, fully-attributed summary, returned as strict JSON. The prompt is built around enforceable constraints rather than stylistic encouragement — the model's job is defined by what it may not do (editorialize, copy, invent, alarm) as much as what it must produce. Persona A news editor producing a multi-source story summary for a calm, AI-assisted personal news app. Tone is wire-service (Reuters/AP): factual, restrained, calm — no urgency framing ("breaking," "shocking"), no emotional or partisan loading, no characterizations of intent. Past tense for events that occurred, present tense for ongoing states. Inputs A JSON object: { cluster_id, articles[] }, where each article carries id, publisher, title, published_at, url, and body (full text). The model must process all supplied articles and reuse the given id values exactly when populating its claim and source references. Output Strict JSON matching a fixed schema — no prose, markdown, or commentary outside the JSON: - headline — ≤90 chars, factual, no clickbait - summary_paragraphs[] — each with its supporting claim_ids - claims[] — each with id, text, and source_article_ids (per-claim provenance) - disagreements[] — contested facts with each position's sources (empty array if none) Initial instructions / constraints (the hard rules) 1. Grounding — every claim must be supported by at least one supplied source; no outside facts, context, or background. 2. Verbatim cap — no run of ≥5 consecutive words copied from a single source (proper nouns, numbers, and short attribution phrases exempt); direct quotes allowed only when wording matters, ≤10 words, quoted and attributed. 3. Inline uncertainty — developing or unverified information must carry its qualifier in the sentence itself ("officials said…", "as of [time], unconfirmed…"), not buried in metadata. 4. Cross-source synthesis — each paragraph must reflect ≥2 sources, or be explicitly framed as single-source attribution. 5. Single-source legal rule (added in v0.2) — for a one-source item: write one neutral who/what/when/where sentence, avoid the source's headline angle/thesis/analysis, don't copy wording, and link prominently to the original. Don't summarize the article — extract factual claims and re-structure them. 6. Disagreements — when sources contradict, populate disagreements[] and reflect the contradiction in the prose. 7. Length — target 80–120 words across the summary, hard cap 200; prefer fewer, shorter sentences. What variations will you test? The new version of the prompt, the validator, and the two-tier routing are already in place. Tuning from here is continuous and evidence-led: the evaluation stack tells us what to change next, rather than us guessing variations up front. The loop is — measure on the gate → find the dominant failure mode → form a hypothesis → test the single change → keep it only if it improves that metric without regressing the others. Our near-term tuning is already pointed by real data. In this week's runs, 27 of 30 Tier 1 failures (90%) were verbatim-overlap — the model copying source phrasing — which makes that the first thing we tune: - Verbatim (the live #1 signal): we'll test two competing fixes against each other — prompt-side (a stronger "paraphrase, don't copy" instruction, possibly with good/bad paraphrase examples) versus validator-side (re-checking whether the 6-gram threshold is too strict and over-flagging). The winner is whichever cuts verbatim failures without lowering faithfulness in manual review. - Tier 2 escalation — tuning from a measured baseline: a first run shows the stronger model (gpt-5.4-mini) rescues ~70% of the clusters that fail Tier 1 validation. That gives us two concrete tuning levers: (1) the ~30% Tier 2 still can't fix — we'll inspect what they share (likely the hardest verbatim or contradiction cases) and feed that back into the Tier 1 prompt so fewer reach Tier 2 at all; and (2) escalation cost — test whether escalating only the failure types Tier 2 reliably fixes beats escalating every failure. We'll re-measure the rescue rate as the prompt and routing change, so it becomes a tracked metric, not a one-off. - Word-count band: the 80–120 / 200-word values are a starting guess. We'll tune them against manual "calm vs. complete" judgments once we've reviewed a sample — especially on multi-source stories with a disagreement to surface. - User-flag-driven tuning (the flywheel): as contestability flags accumulate in the pilot, we'll mine them for recurring error classes and feed those back as the next prompt/validator adjustments — the users effectively set the tuning agenda. So the variations aren't pre-decided; they're whatever the failure data surfaces. Right now that's mostly verbatim copying, so that's where the next experiment goes. What techniques will you use to optimize performance? - Two-tier escalation — run the cheap model first and escalate only validation failures to the stronger model, holding quality while controlling cost. - Single-source bypass — one-article clusters (~90% of volume) skip the LLM entirely and use a deterministic attributed lead, removing the largest hallucination/copying surface. - Programmatic evaluation gate — an automated validator (schema conformance, no hallucinated source IDs, word-count cap, verbatim-overlap) is the first-line optimization signal, so every prompt or parameter change is measured against labeled outcomes rather than judged by eye. Manual evaluation on selected criteria — a sampled set of outputs is human-reviewed against criteria the automated gate can't fully judge: tone neutrality (no editorializing/urgency), faithfulness to sources, attribution correctness, fair handling of disagreements, and overall "calm." This catches quality issues that pass the mechanical checks but would still feel wrong to a reader. - User feedback as a labeled-error loop — in-product contestability lets users flag claims; each flag is a labeled error that feeds back into the evaluation set (the flywheel — users become our eval set). Pilot interviews add the qualitative read on whether the experience actually feels calm. - Structured JSON + schema validation — forcing the typed schema makes per-claim provenance and disagreement-surfacing first-class, and makes failures machine-detectable.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?v0.1 → v0.2: tightened single-source rule (hard rule #4) Change: replaced the v0.1 rule ("if only one source article is supplied, return a 1–2 sentence attributed lead + click-out") with a stricter formulation: - Write one neutral factual sentence using only basic facts: who, what, when, where. - Avoid the source's headline angle, investigation thesis, analysis, conclusions, or unique framing. - Do not copy source wording. - Prominently link to the original. Reasoning: copyright risk on single-source items is not limited to literal word overlap. It also arises from copying a source's selection of details, narrative structure, framing, or investigative conclusions -- even in a short paraphrase. A sentence like "Source reveals how years of ignored warnings and budget pressure led to the crash" is short, but it reproduces the source's investigative framing and choice of conclusions. Under a CC BY-ND license, that starts to look like adapted material, which the license prohibits sharing publicly. The safer principle is: copyright does not protect facts, ideas, events, names, dates, places, or numbers - only creative expression. A sentence that extracts only the factual core (who, what, when, where) and links to the original is a factual pointer, not an adaptation. "Local officials said a train derailment near X caused delays on Monday; read the full report at Source" is not transforming the article. It is pointing to it. The stronger rule is therefore: do not summarize the article -extract only factual claims, then write an event-level sentence in a new structure. Versioning via git
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Not a RAG system Sources: 11 (for the mvp only) RSS feeds across publishers selected for permissive or public-domain licensing. Data pipeline: Articles are fetched from RSS feeds and their full body text is scraped and stored in a PostgreSQL table alongside the RSS summary, publisher, URL, and publication timestamp. Deduplication is done by URL — articles already in the DB are skipped on each run. A rolling time window prevents ingestion of stale feed entries. Each article's title and rss summary are embedded using Pinecone's managed embedding model (llama-text-embed-v2, 1024 dimensions). Articles in the same time window are grouped locally using union-find over pairwise cosine similarity, each group's representative vector is queried against a Pinecone index to match against existing clusters. If a match is found, the group joins the existing cluster; otherwise a new cluster is created. New articles arriving in a later run on the same day are compared against all articles from that day — not just new ones — so they join existing same-day clusters rather than creating duplicates. Each cluster's representative vector is stored in Pinecone. A separate story stitching step runs after clustering: it queries each cluster's vector against other cluster vectors across other days, and clusters that exceed a higher similarity threshold are linked to a shared persistent story entity. This builds the cross-day story sequence — each day's cluster is a snapshot of the story on that date. Summaries are generated per cluster and stored in the database. Cluster summaries go through a two-tier LLM pipeline with automated validation before storage.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.single-source: - Input: one article — [Times of India] "From 0–2 down to 3–2 up: India fight back to beat USA in FIH Nations Cup" - Expected output: a short, explicitly-attributed factual lead, no LLM editorializing — headline: "From 0–2 down to 3–2 up: India fight back to beat USA in FIH Nations Cup"summary: "Times of India reports: Deepika scored twice for India, while Navneet Kaur added the third goal. Ashley Sessa and Madeleine Zimmer scored for the USA. The winner earns promotion to the FIH Pro League…" High-value case — multi-source synthesis (real 6-source cluster): - Input: 6 articles on the same event — The Independent, BBC, Euronews, Al Jazeera, The Guardian, DW, all covering the Kyiv Pechersk Lavra monastery fire. - Expected output: one synthesized summary with per-claim provenance and a surfaced disagreement — headline: "Historic Kyiv monastery damaged in Russian strikes; Moscow reports drone deaths" para 1: UNESCO monastery caught fire in an overnight Russian attack; Ukrainian officials said the Dormition Cathedral was significantly damaged; Russia denied targeting it, claiming a U.S.-made Patriot missile. 14 claims, each mapped to specific source_article_ids; 1 disagreement — "Cause of damage": Ukraine says Russian drone vs. Russia says Patriot missile — reflected both in disagreements[] and in the prose.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Identified from initial production run: 1. Word count violation — model produced 316 words for a 2-article cluster (hard cap: 200). Common with Flash-Lite on verbose topic clusters. 2. Verbatim proper noun sequences — model copied exact phrases containing country names, titles, and names. False positives in the validator until proper noun exclusion was implemented. 3. Single-article cluster not following single-source rule — model produced multi-paragraph summary instead of one attributed sentence, copying verbatim phrases. Fixed by bypassing the LLM entirely for single-article clusters. 4. JSON parse failure — Flash-Lite occasionally returns truncated or malformed JSON on long clusters. Handled as a Tier 1 failure, escalated to Tier 2.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Two-tier architecture validated end-to-end on real data. Single-source clusters bypass the LLM — the pipeline uses the RSS summary directly with attribution, which produced valid output in every case. Multi-source clusters went through Tier 1 (Gemini Flash-Lite), and failures escalated automatically to Tier 2 (gpt-5.4-mini Batch API). The escalation path confirmed working on real data — batches submitted, collected, and validated. Overall design behavior matched intent: no failures surfaced to users, all escalations caught at the gate.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?1 run │ Category │ Result │ │ Single-source clusters (no LLM) │ 128/128 = 100% valid │ │ Multi-article Tier 1 (Flash-Lite) │ ~50% pass after validator calibration │ │ Schema validity │ ~85% (Flash-Lite returns invalid JSON on ~15% of calls) │ │ Word count ≤ 200 │ ~90% (some verbose clusters over-generate) │ │ Verbatim check (calibrated) │ Remaining escalations are genuine │ │ Tier 2 (gpt-5.4-mini batch) │ 70% pass Updated results — 3-day run (June 12–14, 2026): approximately 462 clusters processed. ~432 passed Tier 1 validation (~93%). 30 escalated to Tier 2. Failure breakdown: 27 verbatim-overlap violations (~90% of failures), 3 empty or malformed model responses (~10%). Schema failures: 0. Hallucinated source ID failures: 0. Word-count cap violations: 0 after validator calibration. Dominant failure mode is the model copying ≥6-word runs from sources — a prompt and/or threshold tuning problem, not a structural one.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?From the 3-day production run: (1) Verbatim copying is the dominant, quantifiable failure — ~90% of Tier 1 failures are the model reproducing ≥6-word sequences from source text. This is a paraphrase compliance issue specific to Flash-Lite, not a schema or factual accuracy problem. (2) Clustering occasionally misses same-story groupings — one news event split across multiple single-article clusters instead of merging, which understates multi-source coverage and sends more clusters through the single-source bypass than necessary. (3) DW English has no publication dates in its ingested feed; ingestion date is used as fallback, which is accurate but creates an artificial "same-day" clustering bias for DW articles.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?1. Prompt v0.2 — tightened single-source rule (legal) Discovered: the original rule #5 ("return a 1–2 sentence attributed lead") was legally ambiguous. A paraphrase of a source's investigative framing or selection of conclusions can constitute a derivative work under certain licences even if no words are copied. Fix: rewrote rule #5 to explicitly prohibit reproducing the source's headline angle, investigation thesis, analysis, conclusions, or unique framing. The new rule instructs the model to extract only basic factual claims (who, what, when, where) and write an event-level sentence in a new structure — a factual pointer, not an adaptation. In practice this rule also became moot (see fix #2), but it remains in the prompt as a safety net. 2. Single-source bypass (architectural fix) Discovered: Gemini Flash-Lite consistently ignored rule #5 (single-article handling), producing full multi-paragraph summaries and copying verbatim phrases instead of a single attributed sentence. This was the dominant failure mode — ~92% of initial clusters were single-article. Fix: bypassed the LLM entirely for single-article clusters. The RSS summary field is used directly with source attribution, wrapped in the output schema. The model never receives these clusters. Eliminated 92% of LLM calls and all single-source validation failures in one change. (if/when new articles are added to cluster the summary will be updated) 3. Validator calibration — verbatim check (five iterations) Discovered: the 5-gram verbatim check was catching unavoidable factual phrases and proper noun sequences that the master prompt explicitly exempts ("Proper nouns, numbers, and short attribution phrases do not count toward this limit"). Iterations: - 5-gram raw → flagged "prime minister keir starmer said", "the united states and iran" — proper nouns, not plagiarism - 6-gram raw → still flagged "reached between the united states and" — proper noun split across n-gram boundary - 5-gram + proper noun exclusion (threshold ≥ 2 capitalised tokens) → missed "Germany" at leading position (j=0 was excluded to avoid sentence-start false positives) - Added SENTENCE_STARTERS exclusion, lowered threshold to 1 → position 0 now checked for genuine proper nouns ("Germany", "Bayern") while common sentence starters ("The", "A", "In") are still excluded - Current: 6-gram base + proper noun exclusion (threshold=1, SENTENCE_STARTERS list) → correctly flags genuine verbatim prose copies; skips proper noun sequences Remaining legitimate escalations after calibration: genuine 6-word verbatim copies in policy/news phrasing ("data sharing and joint intelligence operations", "both sides would terminate military operations"). These are real model quality failures, correctly sent to Tier 2. Failure logging by reason: Validator now classifies failures by type (verbatim / malformed / hallucinated IDs / word count) so each iteration can be measured against a specific failure mode rather than just total escalation rate. 4. Model selection Initial plan: single model (Haiku or equivalent). Changed after observing that Flash-Lite fails on multi-constraint outputs but is fast and cheap enough to run on every cluster. Final architecture: two-tier — Flash-Lite for all clusters, gpt-5.4-mini via Batch API only for validation failures. Tier 2 uses the Batch API (50% cost discount, async) since the pipeline is not real-time.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Three-layer approach: Layer 1 — Script (runs on every Tier 1 call, real-time): JSON schema validation, required-fields check, word count, verbatim overlap with proper noun exclusion. Implemented in summarize/validate.py. Layer 2 — Model grader (periodic, not yet built): LLM-as-judge for tone neutrality, synthesis quality, single-source attribution in prose. Runs on a sample of saved summaries, not in the real-time gate. Will use the same proxy. Layer 3 — Human review (weekly + on contestability events): Tone spot-checks, "fits one phone screen" visual check, defamation safety on negative-case fixtures. Scaling: the 50-fixture curated eval set runs fully automated (Layer 1 + 2) in under 5 minutes. Weekly drift run adds 20 random production clusters. Contestability events auto-promote clusters into the eval corpus.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?How often will you re-run evaluations? - Every pipeline run: Layer 1 (script) fires on each cluster automatically. - Before any prompt or model change: full 50-fixture run required, composite gate must pass before deployment. - Weekly: drift run on 20 random production clusters, Layer 1 + 2 (model grader). - On every contestability_event: affected cluster re-run immediately and added to eval corpus. - Monthly post-launch: full human review of 10-cluster sample to catch rater drift.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?At MVP launch, infrastructure is intentionally minimal and local-first. The backend pipeline runs against a PostgreSQL database with a defined schema and versioned migrations. The API layer (to be built before launch) will sit on top of the same DB. External dependencies: Pinecone — serverless vector index used for two purposes: article embedding similarity during clustering (finding which articles belong to the same story), and cross-day story stitching (linking clusters about the same ongoing topic across days). Uses the existing article-clusters index, mvp namespace, Pinecone managed embedding (llama-text-embed-v2, 1024 dimensions). Pinecone availability is a hard dependency for the cluster and stitch pipeline steps; if unavailable, ingested articles accumulate unprocessed but no data is lost. LiteLLM proxy — all LLM calls (Gemini Flash-Lite for Tier 1, gpt-5.4-mini Batch API for Tier 2) route through a single proxy endpoint. If the proxy is unavailable, summarization halts but clustering and ingestion continue unaffected. Rate limits: the two-tier summarization architecture limits LLM exposure — single-article clusters make no API calls; multi-article clusters go through Gemini Flash-Lite first, with escalations to gpt-5.4-mini via async Batch API. Rollback: each pipeline step is independently runnable and idempotent. Summaries can be cleared and regenerated. Migrations are tracked in schema_migrations. No rollback migrations written yet (probably acceptable at MVP given the small team and staged rollout.) Documentation: pipeline steps, commands, DB table map, and external dependency notes are in the project README. Not yet production-hardenedPlease leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Solo founder + co-founder team at MVP stage. No separate support, comms, or legal function. Documentation (README, PRD, Notion) is current. Legal review of source licensing terms for RFI, DW, and all extended sources is pending. Privacy policy and terms of service not yet drafted. I included several unreviewed sources (like Euronews, BBC, the Guardian) for the demo volume only. They will be removed if they do not pass the legal source check and replaced with others.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Closed beta, invitation-only. Initial cohort: ~50 users from the work network and direct personal network. No public signup. Access via direct link to the web app. Goal of the beta: validate that the calm-by-design framing resonates with the target persona ("still cares, can't stomach it"), surface real contestability events to seed the eval flywheel, and confirm that multi-source clusters produce summaries users find trustworthy.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?MVP is architected not for scale but for correctness and iteration speed. The pipeline is scheduled (runs several times a day), not event-driven. Database is a single Postgres instance. LLM calls are batched and async. If beta traction warrants scaling: move from scheduled pipeline to a queue-based system, add a read replica for the API, and monitor Pinecone index size and query latency. The two-tier summarization architecture already separates cheap high-volume calls (Gemini Flash-Lite) from expensive low-volume calls (gpt-5.4-mini batch), which limits cost exposure at scale. Monitoring at MVP: manual log review after each pipeline run, weekly drift eval run. No automated alerting yet.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?For closed beta: - One-paragraph product description for invite emails ("a calm daily news digest — no infinite scroll, no outrage, just what happened") - Short demo video of the core flow (onboarding → feed → story page → ⓘ provenance tap → contestability) - FAQ covering: what sources are used, why some stories are missing, what "looks wrong" does, how following works Full marketing assets (landing page, press kit, App Store listing) deferred to public launch.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Two-person team. Communication is synchronous and informal. Shared Notion for roadmap and decisions. PRD (this document) serves as the single authoritative product spec. No internal comms infrastructure needed at this stage.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?MVP uses anonymous cookie-based identity — no account, no email, no password. The only personal data stored is: a random UUID per device (cookie), topic preferences, and followed stories. No PII collected at MVP. Data is stored in a PostgreSQL instance (Aiven, EU region). Passwords for DB access managed via environment variables, not committed to source control. No analytics SDK, no third-party tracking at MVP. GDPR applicability: minimal at MVP given no PII and EU-region storage. A privacy policy will be required before public launch; drafted to reflect anonymous-first architecture.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Content moderation: the summarizer's master prompt includes hard rules (no editorializing, defamation guardrail, single-source attribution). The contestability flow lets users flag wrong clusters, wrong claims, bad citations, or other issues. Review is partially automated: - Citation reports (object_type = 'citation'): fully automatable. On each report, the backend re-checks whether the cited source URL returns a valid HTTP response. A 4xx/5xx confirms the citation is broken — the cluster is immediately flagged and the citation marked as unresolved in the UI, with no human needed. - Claim reports (object_type = 'claim'): automatable with the LLM judge (Layer 2 eval). On each report, re-run the grounding check: does the flagged claim actually appear in the source articles it cites? The LLM judge scores it. If it fails, the cluster is re-summarized via Tier 2 and the case is added to the eval fixture set automatically. - Cluster reports. At MVP scale, volume thresholding handles urgency: report on a cluster auto-removes it from the feed pending manual review. Confirmed wrong-cluster cases are added to the eval set as negative examples to inform future clustering threshold tuning. - Eval promotion: any contestability event confirmed as a real error — by automated check or human review — is automatically promoted into the eval fixture set. This is the labeled-error flywheel; it requires no ongoing curation effort. Legal: source licensing review is pending for the extended source list (BBC, Guardian, NPR, Al Jazeera, and others added for testing). Production launch will only use sources with confirmed permissive terms (currently: government/public domain sources + Global Voices CC BY). Copyright compliance is architecturally supported by the verbatim overlap validator and the direct quote cap. Regulatory compliance: not operating in a regulated domain (not medical advice, not financial advice). No specific regulatory review needed at MVP.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?Primary success signals: - Daily active users returning 3+ days/week — indicates genuine utility, not novelty - Stories followed per user — indicates trust in the coverage - Contestability reports per 1,000 summaries — should stay low; spikes indicate a model or source quality problem - "Feels calm" qualitative feedback in beta interviews — the primary brand promise Metrics not optimised upward: - Session length maximisation — target band is 3–10 min, not "as long as possible" - Notification open rate — calm-first means fewer, more considered notifications - Article click-through rate — the app is designed to be sufficient on its own; high click-through could indicate summaries aren't trusted Business metrics at post-beta stage: conversion to paid tier, retention at 30/60/90 days.
AI MetricsHow will you measure AI performance and accuracy?Three layers, as defined in the Develop section: Layer 1 — Automated, every run: schema validity, hallucinated source IDs, word count cap, verbatim 6-gram overlap. Pass rate tracked per run. Layer 2 — Model grader, weekly: LLM-as-judge scoring tone neutrality, synthesis quality, and single-source attribution on a 20-cluster sample. Scores tracked over time to catch drift. Layer 3 — Human review, weekly: spot-check of 10 clusters against the full rubric. Defamation safety tested on the negative-case fixture set. Key headline metric: Tier 1 pass rate on multi-article clusters (target: >90% after prompt and validator stabilise). Current baseline: ~70% pre-calibration, improving. Contestability events from production auto-promote into the eval corpus — the flywheel that makes the eval set grow without manual curation effort.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?Beta: direct email to the founders. Response within 24 hours. No ticket system at this scale. In-app: the "Looks wrong" contestability flow is the primary structured feedback channel — it captures wrong cluster, wrong claim, bad citation, and other issues directly in context, without requiring users to leave the app or compose a message. Post-beta: in-app support chat or help centre, depending on volume.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Structured feedback: every "Looks wrong" submission writes to contestability_event in the DB. Reviewed weekly. Clusters with reports are re-run through the eval pipeline. If a cluster fails eval after a contestability report, it is removed from the feed and the case is added to the eval fixture set. Unstructured feedback: beta user interviews (bi-weekly in early beta). Key themes tracked in Notion. Bug prioritisation: Tier 1 hard blockers (hallucination, defamation, wrong cluster) → fix immediately. Tier 2 quality issues (tone problems) → next sprint. Minor or moderate UX issues → backlog. Critical issue communication: direct message to beta users if a data quality issue is discovered retroactively. No public status page at MVP.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?At MVP: automated pipeline_run_log table and a live internal dashboard (Flask + Chart.js, deployed locally, portable to server). The dashboard captures per-run: articles ingested by source, clusters formed, stories stitched, single-source bypass count, Tier 1 pass rate, Tier 2 escalations, failure-type breakdown by reason (verbatim / malformed / hallucinated IDs), and scrape failures per source. Refreshes on page load. Missing before public launch: error alerting on pipeline failures, LLM cost tracking per run, API response time monitoring. Missing at MVP (to add before public launch): error alerting on pipeline failures, LLM cost tracking per run, API response time monitoring, DB connection pool monitoring.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Weekly cadence: - Run the drift eval (20 random production clusters through the model grader) - Review contestability events from the week - Review beta user feedback - Decide: prompt iteration, source list change, or validator adjustment Prompt changes require a full 50-fixture eval run before deployment. Model changes require the same. Source additions require licensing verification (pre-launch) and a test ingest run. The labeled-error flywheel is the core continuous improvement mechanism: every contestability event that validates as a real error becomes a permanent fixture in the eval set, so the system gets harder to fool over time without manual curation effort. Per-reason failure instrumentation directly operationalises the labeled-error flywheel: failure modes are tracked over time by type, not just as a total escalation count. This means prompt and validator iterations can be measured against specific failure categories - if verbatim failures drop after a prompt change but malformed responses increase, the dashboard surfaces that trade-off immediately. The contestability flow (user flags) feeds the same log, so production errors and eval-set errors are tracked in the same system.
Download the .xlsx ↓