Media
Signal Brief
Signal Brief is a daily media intelligence tool for Canadian fintech communications teams. It pulls signals from Google News, RSS, and Reddit, classifies stories by severity, and produces both a full brief and an executive-ready 'so what' summary with suggested communications actions. The product is built to replace a manual 45-minute workflow with a repeatable, context-aware briefing process that runs in under two minutes.
The problem
The communications lead at a Canadian fintech starts each morning with low-grade dread, piecing together a media brief across four or five open tabs — Google News, Reddit, Twitter/X — in roughly 45 minutes. There is no single starting point, no severity framework beyond gut feel, and no way to tell an isolated customer complaint from an emerging journalist narrative. Worse, the brief is descriptive, not actionable: it explains *what happened* but never *what to do about it*. By 10am it is already stale, and leadership follow-ups expose the gaps. As the persona puts it, confidence never peaks — residual doubt is the default even after the brief is sent.
The solution
SignalBrief is a daily AI media intelligence tool for Canadian fintech comms teams. Its core differentiator is the interpretation layer: every story is delivered as *what happened / why it matters / suggested action*, tagged by severity — Noise / Watch / Act / Crisis — on a consistent rubric rather than gut feel. One trigger produces two outputs simultaneously: a full email brief grouped by severity, and a separate executive "so what" story short enough to forward to a CEO. It replaces the manual 45-minute ritual with a repeatable process that runs in under two minutes, and understands Canadian fintech regulatory context — OSFI, PIPEDA, Bank of Canada — and the local media hierarchy.
How it works
After a one-time company-profile setup, the app pulls live signals server-side from Google News RSS and Reddit public JSON, deduplicates and filters for relevance, then passes them to Claude Sonnet 4.6 via a single API call. The model filters, classifies severity with one-line reasoning, writes the "so what" layer per story, and assembles both outputs. All relevant context fits in Claude's 200K-token window, so no vector database is needed for the MVP. Claude was chosen for strict template adherence, low hallucination (every claim must trace to a provided source), and long context. Prompt iteration was disciplined: a Canadian-context block lifted "why it matters" usefulness from 3.8 to 4.3, headline-only safeguards eliminated hallucination on thin inputs, and tier-specific tone guidance raised tone ratings to 4.6. Testing hit 87% severity-calibration agreement and a 0% hallucination rate.
Who it's for
SignalBrief is B2B, sold to communications and growth teams at digital-first Canadian fintechs of 50–500 employees — neobanks and financial apps. The persona is "Jane," a Director or Manager of Communications at a ~250-person fintech who can expense a $35/mo tool without CMO approval. Jane isn't buying media-monitoring software; she's buying confidence, speed, executive trust, and protection from being blindsided. Winning looks like a brief ready in under five minutes, fewer leadership follow-up questions, and "so what" framing in every story.
Why it matters
The global AI media monitoring market is projected to nearly triple from ~$4B in 2025 to $9–11B by 2030, with the mid-market segment growing fastest as enterprise tools like Meltwater and Brandwatch ($15–30K/yr) stay out of reach. Post-SVB and post-crypto-contagion, brand trust is a board-level concern. SignalBrief's wedge is a price-to-value gap: $35/mo replacing ~15 hours of manual work a month, at ~97% gross margin, with no vertical-specific AI brief tool for Canadian fintech today. It replaces the DIY ritual — Google Alerts plus manual search — not Meltwater, and in a tight community where word of mouth travels fast, that focus is the moat.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Khushboo Sharma | |||||
| Your Product: | SignalBrief - AI Media Intelligence Tool | |||||
| Your Industry: | Canadian Fintech - specifically digital-first banks, neobanks, and financial apps at the 50 - 500 employee stage | |||||
| Date: | 5th May, 2026 - 15th June, 2026 - Video link - https://preview--md-masterpiece-maker.lovable.app/ | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | B2B SaaS — AI Media Intelligence, vertically focused on Canadian Fintech (50–500 employee digital-first banks, neobanks, and financial apps). SignalBrief sits at the intersection of two markets: • Primary: AI-powered media intelligence / brand monitoring SaaS • Vertical: Canadian Fintech — the domain it serves and understands TAM / SAM / SOM (at $99/mo): • TAM — global mid-market AI brand intelligence: ~$800M–1B • SAM — Canadian fintech, 50–500 employees, ~120 realistic buyers: ~$142K ARR • SOM — Year 1 target (10 customers): ~$12K ARR; US fintech expansion multiplies opportunity ~10x | Khushboo, the product framing in your attached document carries real clarity. The gap you identified between expensive enterprise monitoring platforms and free alternatives with no intelligence layer is a genuine market observation, and anchoring the MVP at a specific price point for a specific company size range shows you are thinking in product decisions, not abstractions. Two gaps to close before this Discovery section is Demo Day ready. The first is that the attached document reads as a strong executive summary, not as the completed Discovery work. The persona, journey map, pain-point ranking, AI opportunities, headwinds and tailwinds, and competitive landscape analysis are all referenced in the PRD as living in a detailed document, but neither the PRD cells nor the Word file contain them. On Demo Day, a judge evaluates what is in the artifact. A product pitch paragraph, no matter how sharp, cannot substitute for the structured Discovery work that shows you understand the user's daily friction, ranked the severity of each pain point, and traced a logic chain from those pains to an AI-powered solution. Build out each of those sections with the same specificity you brought to the overview. The second is AI necessity. The overview describes an interpretation layer that classifies severity and attaches a recommendation to each story. Severity classification on its own could run on keyword rules and source-authority scoring without a language model. The place where generative AI earns its seat is in synthesizing unstructured media mentions into a contextual narrative that explains why a particular story matters to this specific company and what the comms team should do about it. That reasoning chain, taking raw coverage from multiple sources, inferring reputational relevance, and producing a plain-language action recommendation, is not something a rules engine can replicate. Isolate that moment explicitly in your AI opportunities section and make it the anchor of the hypothesis. One forward note: the competitive landscape currently names only the two price extremes. Map the middle tier. Tools in the low-hundreds-per-month range serving similar teams likely exist, and a judge will ask where they fit. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | HEADWINDS & TAILWINDS HEADWINDS: • Market: Meltwater/Brandwatch have strong enterprise lock-in • Technology: Slower AI adoption in fintech; trust is fragile • Regulatory: PIPEDA limits some data sources • Buyer behaviour: Comms budgets at small fintechs are lean • Competition: Google Alerts exists — free and "good enough" for some • Distribution: Small market (~200–300 companies) limits scale TAILWINDS: • Market: Mid-market fintechs are priced out of enterprise tools • Technology: LLM interpretation layer is now feasible at low cost • Regulatory: Open banking regulation increasing fintech media coverage — more signal to monitor • Buyer behaviour: Post-SVB, post-crypto contagion — brand trust is a board-level concern • Competition: No vertical-specific AI brief tool exists for Canadian fintech today • Distribution: Tight community = word of mouth travels fast (Fintech Cadence, DMZ, Volta) COMPETITORS Enterprise (not direct — their pricing gap is the opportunity): • Meltwater (~$15–30K/yr), Brandwatch, Sprinklr — serve companies SignalBrief's buyer cannot afford Mid-market tools (same price point, but they give data — SignalBrief gives meaning): • Mention.com ($99–300/mo), Brand24, Talkwalker • These tools aggregate mentions and show volume; they do not classify severity, generate "so what" framing, or understand Canadian fintech regulatory context DIY / Workarounds (real competition today): • Google Alerts + manual search, ChatGPT + copy-paste, the comms person doing it manually 45 min/day • SignalBrief replaces this ritual, not Meltwater | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | Global AI media monitoring market projected to nearly triple from ~$4B (2025) to ~$9–11B (2030), growing ~18–22% CAGR. The mid-market SaaS segment is the fastest-growing cohort as enterprise pricing remains out of reach for the majority of companies. Generative AI is accelerating the shift from data-aggregation tools to interpretation-layer tools — the category SignalBrief is entering. Key drivers: increased media fragmentation, board-level demand for brand risk management post-SVB/crypto contagion, and declining cost of LLM inference making AI interpretation commercially viable at the $35/mo price point. | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Startup | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | B2B SaaS subscription — 30-day free trial (no credit card required), then $35/month. Pricing rationale: • $35/mo is deliberately below the mid-market tier (Mention.com at $99–300/mo, Brand24 at $79–239/mo) — removes the "let me check with my manager" conversation • At $35/mo Jane expenses it instantly, no procurement process needed at any Canadian fintech • Replaces ~15 hours of manual work per month — effective hourly rate of $2.33/hr for the value delivered • 30-day free trial removes the biggest B2B SaaS adoption barrier: "what if it doesn't work for us?" • Path to expansion: $35/mo solo → $79/mo team (3 seats + competitor layer) → $149/mo growth (Slack alerts + history) Unit economics at $35/mo: • LLM cost per user: ~$0.80/mo (20 briefs × $0.04) • API cost per user: ~$0.20/mo • Gross margin: ~97% at MVP scale • MRR targets: $350 (10 users, Month 2) → $1,050 (30 users, Month 4) → $3,500 (100 users, Month 8) | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2B | |||||
| Differentiators | What are the key differentiators for your company? | SignalBrief's core differentiator is the INTERPRETATION LAYER — every brief delivers not just what happened, but why it matters to this brand and what to do about it. Structured as: what happened / why it matters / suggested action. Key differentiators: 1. "So what" framing: Every story includes an LLM-generated implication + action recommendation. Mid-market tools (Mention, Brand24) show data. SignalBrief gives meaning. 2. Severity classification: Every story is tagged Noise / Watch / Act / Crisis using a consistent AI-driven rubric — not gut feel. 3. Canadian fintech vertical context: Prompts are trained on Canadian regulatory context (OSFI, PIPEDA, Bank of Canada), the difference between a customer complaint and an emerging journalist narrative, and the specific media landscape (The Logic, BetaKit, Globe & Mail Fintech coverage). 4. Briefing memory: Week-over-week trend detection so Jane can say "this is the third time this theme surfaced" — not possible with stateless monitoring tools. 5. Price-to-value gap: $35/mo vs. $15–30K/yr enterprise tools. Replaces ~15 hrs/mo of manual comms work — expensable without a budget conversation. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | [Insert your response here] | |||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | [Insert your response here] | ||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | [Insert your response here] | ||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | PERSONA: Jane — Comms Lead Role: Director / Manager of Communications or Growth at a 250-person digital-first Canadian fintech Company size: 50–500 employees User type: External — primary daily user Buyer role: Champion (can expense at $35/mo without CMO approval; CEO/CMO holds the formal budget) Tools today: Google News, Reddit, Google Alerts, Slack, Email Opening quote: "I spend 45 minutes every morning piecing together a media brief from five different tabs — and by the time I send it, I'm already worried I missed something." Goals: • Keep leadership informed before the 9am standup • Catch bad PR before it escalates to the CEO • Look sharp and credible to executives • Spend less time on research, more time on strategy Frustrations: • No single place to start — 4–5 apps open simultaneously • No framework to judge severity — uses gut feel, which doesn't scale • Brief is descriptive, never actionable — doesn't answer "what should we do?" • News cycle moves faster than her workflow — brief is stale by 10am What winning looks like: Brief ready in under 5 min · leadership asks fewer follow-up questions · never blindsided · "so what" framing in every story Strategic insight: Jane isn't buying "media monitoring software." She's buying confidence, speed, executive trust, and protection from being blindsided. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | CURRENT-STATE JOURNEY MAP (6 steps) Step 01 — TRIGGER | 7:45am Jane opens her laptop thinking "did anything bad happen overnight?" She opens Google News with low-grade dread. Pain: No single starting point. No overnight alert. Walks in blind. Emotion: Anxious 😰 Step 02 — GATHERING | 8:00–8:20am Searches Google News, Reddit, Twitter/X across 10–15 open tabs. Pain: 20 min of pure manual labour. No prioritisation — a Bloomberg article = a 4-like tweet. Misses niche forums, G2, HackerNews. Emotion: Overloaded 😤 Step 03 — INTERPRETATION | 8:20–8:35am Reads articles, decides what matters, judges severity alone with no framework. Pain: No rubric. Can't tell if something is a one-off complaint or an emerging journalist narrative. No one to pressure-test with. Emotion: Uncertain 😟 Step 04 — DRAFTING | 8:35–8:50am Copy-pastes headlines with one line of context each. Blank page every day. Pain: Inconsistent format. No "so what." 9am deadline pressure. Brief describes events, doesn't explain implications. Emotion: Rushed 😓 Step 05 — DISTRIBUTION | 8:50am Sends to Slack or email, moves on. Pain: No confirmation she got it right. No feedback loop. Residual doubt even after sending. Emotion: Uneasy relief 😮💨 Step 06 — AFTERMATH | Throughout day Answers follow-up questions from leadership. Reacts to late-breaking stories the brief missed. Pain: Brief didn't answer "what should we do?" Leadership follow-ups expose the gap. No trend context — every day starts fresh. Emotion: Exposed 😬 KEY INSIGHT: Jane's confidence never peaks. Even after sending, residual doubt is the default, not the exception. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | TOP PAIN POINTS — ranked by frequency × severity CRITICAL (Daily / High Severity): P1 — No single starting point: 4–5 apps to begin each morning. [Daily / Critical] P2 — No overnight alert: walks in blind every day. [Daily / Critical] P3 — 20-min manual search: zero intelligence layer. [Daily / Critical] P4 — No prioritisation: Bloomberg article = a 4-like tweet. [Daily / Critical] P5 — No severity framework: gut feel, doesn't scale or audit. [Daily / Critical] P6 — Can't distinguish complaint vs. journalist narrative. [Weekly / Critical] P7 — Brief is descriptive only — no "so what" framing. [Daily / Critical] P8 — No sentiment baseline — can't say better or worse vs. last week. [Daily / Critical] P9 — Brief doesn't answer "what does this mean for us?" [Weekly / Critical] P10 — No trend memory — every day starts fresh. [Daily / Critical] HIGH (Frequent / High Impact): P11 — Coverage gaps: niche forums, G2, HackerNews missed. [Weekly / High] P12 — Blank page every day — no template or starting structure. [Daily / High] P13 — Sends with residual doubt — no confidence signal. [Daily / High] P14 — Stories break after 9am — brief stale within hours. [Weekly / High] THREE ROOT CAUSES: 1. Fragmentation of sources (no single intake point) 2. Missing judgment scaffold (no severity/priority framework) 3. No forward value (no trend memory, no week-over-week context) | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | AI OPPORTUNITIES — ranked by impact (LLM-addressable pain points only) Rank 1 — P7: No "so what" framing → LLM generates what-happened / why-it-matters / suggested-action per story → AI capability: Contextual text generation + vertical reasoning → WHY THIS REQUIRES GENERATIVE AI: A rules engine cannot infer reputational implication from unstructured text. The LLM reads raw coverage, infers relevance to this specific company's brand position, and produces a plain-language action recommendation — a reasoning chain no keyword system can replicate. Rank 2 — P5: No severity framework → LLM classifies each story as Noise / Watch / Act / Crisis with reasoning → AI capability: Classification + chain-of-thought reasoning Rank 3 — P3: Manual gathering (20 min) → LLM summarises + filters raw feeds; removes duplicates and irrelevant mentions → AI capability: Summarisation + relevance filtering Rank 4 — P4: No prioritisation → LLM scores signal by source credibility, reach, and brand proximity → AI capability: Ranking + contextual reasoning Rank 5 — P6: Complaint vs. narrative detection → LLM detects pattern across sources (single complaint vs. emerging media narrative) → AI capability: Pattern recognition across unstructured sources Rank 6 — P9: Doesn't answer leadership → LLM pre-generates "implication for us" section for executive read → AI capability: Audience-aware contextual reasoning Rank 7 — P12: Blank page every day → LLM outputs a structured, formatted brief draft → AI capability: Structured generation with consistent template Rank 8 — P8: No sentiment baseline → LLM scores + aggregates daily sentiment; flags directional shift → AI capability: Sentiment analysis + longitudinal comparison Rank 9 — P10: No trend memory → LLM references prior briefing context to flag recurring themes → AI capability: Session-context injection + longitudinal reasoning Rank 10 — P11: Coverage gaps → LLM infers what's likely missing from the current signal set → AI capability: Gap detection + inference | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | 32 IDEAS ACROSS 5 CATEGORIES AUTOMATION & AGGREGATION (ideas #1–8): 1. Overnight news digest auto-pulled at 6am 2. Multi-source aggregator (News + Reddit + Twitter/X + G2) 3. Real-time Slack alert for Watch/Act stories 4. Auto-tagging by topic (regulatory / product / leadership / competitor) 5. Duplicate detection and deduplication across sources 6. Source credibility scoring (tier 1 media vs. fringe blog) 7. Brand mention frequency counter 8. Overnight silence confirmation ("nothing significant, here's why") INTELLIGENCE & INTERPRETATION (ideas #9–16): 9. Competitor intelligence layer (top 3–5 fintechs monitored in parallel) 10. Severity classification: Noise / Watch / Act / Crisis per story 11. Sentiment delta vs. prior week ("brand sentiment up 12%") 12. Complaint vs. narrative detection (single tweet vs. journalist pickup) 13. Share-of-voice tracker vs. competitors 14. "So what" framing layer — why it matters + what to do 15. Regulatory radar: OSFI / Bank of Canada / FINTRAC mentions 16. Influencer + analyst detection (is this person credible?) BRIEF GENERATION (ideas #17–24): 17. Structured brief template auto-filled by AI 18. Executive "so what" story (separate from the full brief) 19. Tone selector (neutral / cautious / urgent) 20. Length selector (brief / standard / deep dive) 21. Custom focus instruction ("focus on acquisition impact today") 22. One-click brief regeneration with new instructions 23. Week-over-week brief comparison report 24. Month-end brand health summary COLLABORATION & WORKFLOW (ideas #25–28): 25. "Send to my email" button 26. Copy-to-clipboard for brief sections 27. Shareable brief link (view-only for leadership) 28. Team annotation layer (Jane adds context before sharing) PROACTIVE & PREDICTIVE (ideas #29–32): 29. Early warning: "this topic is gaining velocity" 30. Crisis playbook trigger: auto-suggests response framework if Act/Crisis tagged 31. Story trajectory prediction: "this complaint pattern typically escalates in 48h" 32. Pre-brief alert if overnight signal crosses severity threshold | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | TOP 3 SOLUTIONS — ranked by impact × feasibility 🥇 SOLUTION 1 — The AI-Generated Daily Brief [SELECTED FOR MVP] Combines ideas: #14 (so what framing) + #10 (severity classification) + #2 (multi-source aggregation) + #17 (structured brief generation) + #21 (custom focus instruction) + #25 (send to email) Pain points solved: P3, P4, P5, P6, P7, P12, P13 Why it wins: Solves Jane's daily pain directly and completely. Buildable without a developer (NewsAPI + Reddit + Claude API). Solutions 2 and 3 depend on it existing first. Clearest AI value proposition. Definable evaluation criteria. Feasibility: High — public APIs, copy-paste interface, no database required for MVP demo. 🥈 SOLUTION 2 — The Competitor Intelligence Layer [Phase 2] Combines ideas: #9 (competitor monitoring) + #13 (share-of-voice) + #15 (regulatory radar) Pain points solved: P9, P11 + latent competitive blind spots Why Phase 2: Requires Solution 1 as foundation. Adds complexity (competitor config per account). High value but not MVP-necessary. 🥉 SOLUTION 3 — The Proactive Slack Alert [Phase 2] Combines ideas: #3 (real-time Slack alert) + #30 (crisis playbook trigger) + #29 (velocity warning) Pain points solved: P2, P9, P13, P14 Why Phase 2: Requires Slack integration and persistent monitoring infrastructure. Meaningful value but not the core daily-use case. SELECTION RATIONALE: Solution 1 is the standalone daily value prop. Without it, Solutions 2 and 3 have no context to amplify. It is the only solution where AI necessity is unambiguous — a rules engine cannot write a contextual "so what" interpretation that understands Canadian fintech brand risk. | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | TARGET STATE WORKFLOW — SignalBrief AI-Generated Daily Brief PHASE 0 — First-time setup (runs once; editable anytime) Jane configures: Company profile (brand name, stage, core products, key markets) → Competitor list (e.g., Koho, EQ Bank, Wealthsimple) → Sensitivity rules (priority topics, exclusion keywords) → Delivery preferences (email, tone, brief length). Saved as persistent system context injected into every brief run. Jane never repeats this setup. PHASE 1 — Jane triggers the brief Jane clicks "Run today's brief." System loads her saved company profile as context automatically. No copy-pasting, no configuration required. PHASE 2 — AI builds the brief (fully automated; no human input) Step 1: Data pull — Live NewsAPI call (company brand + key terms) + Reddit search (relevant fintech subreddits) Step 2: Filter + rank — AI scores each result for signal vs. noise; removes duplicates and off-topic mentions Step 3: Severity tag — AI classifies each story: Noise / Watch / Act / Crisis with one-line reasoning Step 4: "So what" layer — AI writes what-happened / why-it-matters / suggested-action per story Step 5: Assembly — Two outputs generated simultaneously: ① Full email brief (summary paragraph + categorised stories with headlines, 3-line summaries, links) ② Executive "so what" story (severity + implication + action, short-form for leadership) Error branches: Data pull fails → error shown, prompt retry. Low-confidence story → flagged with ⚠ warning for Jane to review. PHASE 3 — Jane reviews, edits, sends Both outputs displayed side-by-side on screen. Decision: Happy with output? NO → Jane types instruction ("focus more on regulatory angle") OR edits directly → AI regenerates that section → loop back YES → Copy to clipboard (either or both sections) → Click "Send to my email" → brief delivered to saved inbox Fallback: email fails → clipboard fallback displayed END STATE ① Full email brief in Jane's inbox — she forwards to leadership in one click ② Executive "so what" story copied separately — she pastes into her CEO/CMO update Architecture note: The LLM is stateless — it reads a prompt and responds, then forgets. The application layer orchestrates everything: assembles the prompt package (company config + raw data + session context), manages state, and decides when the session is complete. | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | WIREFRAMES — 4 key screens SCREEN 1: Onboarding — Company Setup Layout: Single-column form, 4 sections Fields: Company name, industry vertical, key brand terms to monitor (tags), competitor names (up to 5), exclusion keywords, preferred brief length (Short / Standard / Detailed), email delivery address CTA: "Save & Generate My First Brief" AI element: None — pure configuration Design note: One-time setup with a progress bar (4 steps). Stored profile shown as badge on main dashboard ("Profile: Wealthsimple | 3 competitors"). SCREEN 2: Main Dashboard — "Run Today's Brief" Layout: Clean single-action screen Center: Large "▶ Run Today's Brief" button with last-run timestamp ("Last run: yesterday 8:02am") Status panel: Data sources status (NewsAPI ✓ / Reddit ✓ / Error ✗) Profile summary: Company, competitors, monitored terms (editable inline) Processing state: Animated progress bar with status labels ("Pulling sources... Classifying severity... Writing brief...") SCREEN 3: Brief Output — Review & Edit Layout: Two-column view Left column (full brief): Summary paragraph at top → Stories grouped by severity tier (Act / Watch / Noise) → Each story: headline + 3-line AI summary + "so what" framing + source link + severity badge Right column (executive story): Condensed "so what" narrative for leadership — one paragraph per Act/Watch story, plain language Actions: - "Edit this section" → inline text editor per story - "Re-run with instruction" → text input field → AI regenerates - "Copy brief to clipboard" button - "Copy exec story to clipboard" button - "Send to my email" primary CTA button - ⚠ warning badge on low-confidence stories SCREEN 4: Email Delivery Confirmation Layout: Modal overlay Content: "Brief sent to jane@company.com ✓" with preview of subject line Secondary: "Copy to clipboard instead" fallback link Note: No new screen — overlay on Screen 3 | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | PROTOTYPE — SignalBrief MVP (built in Lovable) Prototype link: https://id-preview--992ede1c-b216-4a11-ac23-ad59c519ebf2.lovable.app/ Tech stack: React + TypeScript + Tailwind CSS + shadcn/ui components + TanStack createServerFn (Lovable server-side functions) Live data sources (no API key required on frontend): • Google News RSS — `news.google.com/rss/search?q={brand_terms}&gl=CA` — fetched server-side via Supabase edge function (bypasses CORS) • Reddit public JSON — `reddit.com/search.json?q={company}&t=week` — fetched server-side via Supabase edge function LLM: Claude Sonnet 4.6 (`claude-sonnet-4-6`, Anthropic API) — called via TanStack createServerFn (Lovable server-side) with ANTHROPIC_API_KEY stored as a Lovable secret (never exposed to frontend) Architecture pattern: Frontend → Supabase edge functions → external APIs. All sensitive keys server-side only. 3 screens built: 1. Onboarding — company profile setup form with 30-day free trial banner; profile saved to localStorage 2. Dashboard — "Run Today's Brief" main action with animated loading states and cycling status messages 3. Brief Output — two-column layout: left = full brief grouped by severity tier (NOISE/WATCH/ACT/CRISIS) with colour-coded badges; right = executive "so what" story; copy + email send buttons; re-run with custom instruction MVP1 scope (prototype): ✓ Live news + Reddit data ✓ Claude AI "so what" interpretation ✓ Severity classification (Noise/Watch/Act/Crisis) ✓ Copy to clipboard + send to email (mailto) ✓ Re-run with custom instruction MVP2 (post-capstone): • Briefing history (localStorage then Supabase) • Competitor intelligence layer • Proactive Slack alerts • Week-over-week sentiment trends Prototype link: https://id-preview--992ede1c-b216-4a11-ac23-ad59c519ebf2.lovable.app/ | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | MASTER PROMPT — INITIAL DESIGN SYSTEM INSTRUCTION: You are SignalBrief, an AI media intelligence assistant built specifically for communications and growth teams at Canadian fintech companies. Your job is to read raw news and social media mentions about a brand and produce a structured daily brief that helps the comms lead understand what happened, why it matters, and what to do about it. Tone: Professional, clear, direct. Never alarmist. Never dismissive. Output format: Strictly follow the template below. Do not add sections that are not in the template. Audience: A comms or growth lead who needs to brief their CEO in under 5 minutes. Canadian fintech context: You understand the Canadian financial services regulatory landscape (OSFI, PIPEDA, Bank of Canada), the major Canadian fintech media outlets (The Logic, BetaKit, Globe and Mail), and the difference between an isolated customer complaint and an emerging journalist narrative. INPUT STRUCTURE: [COMPANY PROFILE] Brand: {{company_name}} Industry: Canadian Fintech Products: {{products}} Competitors to monitor: {{competitor_list}} Priority topics: {{priority_topics}} Exclusion keywords: {{exclusion_keywords}} Brief style: {{brief_length}} [RAW SIGNALS — from NewsAPI + Reddit] {{raw_article_titles_and_snippets}} TASK: 1. Filter out irrelevant results and duplicates. 2. For each relevant story, classify severity: Noise (informational only) / Watch (worth monitoring) / Act (requires comms response today) / Crisis (immediate escalation required). 3. For each story, write: - HEADLINE: [original or paraphrased title] - SUMMARY: 2–3 sentences of what happened - WHY IT MATTERS: 1–2 sentences on reputational implication for {{company_name}} - SUGGESTED ACTION: One clear recommendation (e.g., "Monitor for 48h" / "Prepare holding statement" / "Escalate to CMO") 4. Group stories by severity tier (Act first, then Watch, then Noise). 5. Write a 2–3 sentence EXECUTIVE SUMMARY at the top. 6. Write a separate EXECUTIVE "SO WHAT" STORY: One short paragraph summarising the most important stories for leadership, framing the business implication. CONSTRAINTS: - Do not hallucinate sources. Only reference articles included in the raw signal input. - Flag any story where you are uncertain about severity with ⚠. - If no relevant stories found, say so explicitly: "No significant brand signals detected today." - Never include competitor stories that are not directly relevant to {{company_name}}'s brand position. OUTPUT TEMPLATE: --- SIGNALBRIEF — [DATE] Company: {{company_name}} Brief generated: [timestamp] EXECUTIVE SUMMARY [2–3 sentences] EXECUTIVE "SO WHAT" STORY [1 paragraph for leadership] --- ACT --- [stories requiring response today] --- WATCH --- [stories to monitor] --- NOISE --- [informational only] --- | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | EVALUATION CRITERIA — what makes a "good" SignalBrief output OBJECTIVE CRITERIA (pass/fail or scored): 1. Relevance accuracy (target: >90% of stories are actually about the brand) Measure: Count irrelevant stories included ÷ total stories in brief Pass: Fewer than 1 irrelevant story per brief 2. Severity classification accuracy (target: >85% alignment with human reviewer) Measure: Human reviewer independently classifies same stories; compare to AI output Pass: AI and human agree on severity tier for ≥85% of stories 3. "So what" quality — completeness (target: 100%) Measure: Every story includes all 3 components: Why it matters + Suggested action + Headline Pass: Zero stories missing any required component 4. Hallucination rate (target: 0%) Measure: Any AI-generated claim not traceable to a provided source URL Pass: Zero hallucinated facts in any brief 5. Format compliance (target: 100%) Measure: Output matches template structure (Executive Summary → "So What" Story → Act / Watch / Noise grouping) Pass: All required sections present and correctly ordered 6. Brevity — executive summary (target: 2–3 sentences, ~50–70 words) Measure: Word count of executive summary section Pass: Within range SUBJECTIVE CRITERIA (human judgment required): 1. Tone appropriateness: Is the language appropriately cautious for an Act story vs. measured for a Watch story? (Reviewed by Jane) 2. "So what" usefulness: Would Jane actually use this suggested action? (Rated 1–5 by Jane) 3. Executive readiness: Could the exec story be forwarded to a CEO without editing? (Rated 1–5 by Jane) 4. Canadian fintech context: Does the AI correctly understand the regulatory or media nuance? (Reviewed by domain expert) | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | EXAMPLE CASES for test prompts TYPICAL EXAMPLES (happy path): Case 1 — Watch story: Reddit thread on r/PersonalFinanceCanada mentions slowdowns with a payment app. 12 upvotes, no news pickup yet. Expected output: Watch tier. Why it matters: Early Reddit signal before journalist pickup. Suggested action: Monitor for 24h; prepare holding statement if thread exceeds 50 upvotes or gets media mention. Case 2 — Noise story: BetaKit article profiles Canadian fintech sector growth. Company name mentioned once as an example. Expected output: Noise tier. Why it matters: Positive sector association, no brand risk. Suggested action: Share internally for team morale; no external response needed. Case 3 — Act story: Globe and Mail article questions data practices of Canadian neobanks, names company explicitly. Expected output: Act tier. Why it matters: Tier 1 media + data privacy angle = high reputational risk. Suggested action: Escalate to CMO today; prepare factual response addressing data handling practices. EDGE CASES: Case 4 — No relevant results: NewsAPI and Reddit return zero results for brand terms today. Expected output: "No significant brand signals detected today. Sources checked: NewsAPI, Reddit. Consider expanding search terms if this persists." Case 5 — Ambiguous severity: Article mentions competitor Koho's data breach. Company not named. Expected output: Watch tier with ⚠ flag. Why it matters: Competitor breach elevates sector scrutiny; journalists may broaden coverage to include similar companies. Suggested action: Monitor for journalist outreach; review your own data handling comms proactively. Case 6 — Extremely long raw input (many articles): 50+ articles returned from NewsAPI. Expected output: AI filters to top 8–10 most relevant; notes "X additional low-signal results excluded" at bottom. NEGATIVE CASES: Case 7 — Hallucination risk: Raw input contains only a headline, no article body. Expected output: AI summarises only what is visible; does not invent article content. Flags with ⚠: "Summary based on headline only — full article not available." Case 8 — Off-topic injection: User's competitor list accidentally includes a non-fintech brand. Expected output: AI includes competitor stories only if they have a direct connection to the user's brand position; otherwise excludes with note. | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | AI MODEL SELECTION: Claude Sonnet 4.6 (`claude-sonnet-4-6`, Anthropic) — Primary JUSTIFICATION: SignalBrief requires a model that can (a) reason about reputational implication from unstructured text, (b) maintain consistent output structure across many stories, (c) apply Canadian fintech regulatory context, and (d) avoid hallucination on a zero-tolerance basis (no invented sources). Why Claude 3.5 Sonnet: • Superior instruction-following: Consistently adheres to strict output templates (critical for structured brief format) • Low hallucination rate: Anthropic's Constitutional AI training reduces confabulation — essential when every claim must be traceable to a source • Long context window (200K tokens): Handles large raw news payloads (50+ articles) in a single prompt without chunking • Strong reasoning: Chain-of-thought reasoning for severity classification ("this is Watch, not Act, because...") is reliable • Cost-effective at scale: ~$3 per million input tokens — a daily brief with 10K tokens of input costs ~$0.03/run at $35/mo pricing Capabilities utilised: • Summarisation of long-form article text • Classification with structured reasoning (Noise/Watch/Act/Crisis) • Contextual generation (why-it-matters + action recommendation) • Format compliance (strict template adherence) Limitations and mitigations: • No real-time web access → mitigated by pre-fetching via NewsAPI + Reddit before LLM call • Stateless → mitigated by injecting session context and prior brief summary in each prompt • Occasional over-caution (classifies Watch as Act) → mitigated by calibration examples in system prompt Integration approach: Anthropic API (REST) → called by application layer after data fetch. Single API call per brief run. Response parsed and rendered in UI. | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | REQUIRED INPUT FIELDS Field 1: company_name Format: Text string Source: User-configured company profile (setup screen) Requirement: Required Example: "Wealthsimple" Notes: Injected into every prompt as primary search term and brand anchor Field 2: products Format: Comma-separated text Source: User-configured company profile Requirement: Required Example: "Wealthsimple Trade, Wealthsimple Cash, Wealthsimple Tax" Notes: Helps AI understand brand scope; prevents off-topic product stories being excluded Field 3: raw_news_feed Format: JSON array of {title, description, url, source, published_at} Source: NewsAPI live call — query: company_name + brand terms, last 24h Requirement: Required Example: [{"title": "Wealthsimple expands to US market", "description": "...", "url": "...", "source": "The Logic", "published_at": "2026-06-13T07:30:00Z"}] Notes: Fetched immediately before LLM call; passed as structured data Field 4: raw_reddit_feed Format: JSON array of {title, selftext, score, url, subreddit} Source: Reddit API — subreddits: r/PersonalFinanceCanada, r/CanadianInvestor, r/fintech Requirement: Required Example: [{"title": "Anyone else having issues with Wealthsimple transfers?", "selftext": "...", "score": 47, ...}] Notes: Filtered to posts with >5 upvotes; last 24h only Field 5: brief_date Format: ISO 8601 date string Source: System-generated (today's date) Requirement: Required Example: "2026-06-13" Notes: Included in output header; used for briefing history tracking | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | OPTIONAL INPUT FIELDS (MVP1 — as built in prototype) • competitor_list (comma-separated, max 5 names) e.g. "Koho, EQ Bank, Nesto" → AI monitors competitor mentions and flags stories where competitor news has brand implication for the user's company • brief_tone (enum: Executive / Strategic / Conversational / Cautious) Default: Executive → Adjusts the voice and framing of the entire brief — summaries, "why it matters", suggested actions, and executive story all adapt to the selected tone • custom_instruction (free text) e.g. "Focus more on the regulatory angle" / "Downplay the Reddit results today" → Entered on the brief output screen before re-running; appended to the prompt so AI adjusts emphasis without re-fetching news data Note: priority_topics and exclusion_keywords are planned for MVP2. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | OBJECTIVE OUTPUT EVALUATION CRITERIA 1. STRUCTURE COMPLIANCE Criterion: All 5 required sections present (Executive Summary, Executive "So What" Story, Act, Watch, Noise) Measure: Section header detection in output string Pass: All 5 headers present; sections in correct order Fail: Any section missing or out of order 2. STORY COMPLETENESS Criterion: Every story contains Headline + Summary + Why It Matters + Suggested Action Measure: Parse each story block for all 4 sub-fields Pass: 100% of stories have all 4 components Fail: Any story missing a required sub-field 3. HALLUCINATION RATE Criterion: Every factual claim is traceable to a URL in the input data Measure: Human spot-check 3 claims per brief; cross-reference to provided sources Pass: 0 unverifiable claims Fail: Any claim referencing a source not in the input 4. RELEVANCE RATE Criterion: Stories included are about the company's brand, products, or direct competitive context Measure: Human reviewer rates each story as Relevant / Marginal / Irrelevant Pass: ≥90% rated Relevant or Marginal Fail: >10% rated Irrelevant 5. SEVERITY CALIBRATION Criterion: Severity tier matches human reviewer's independent classification Measure: Human independently classifies same stories; compare to AI output Pass: ≥85% agreement on tier Fail: <85% agreement, or any Crisis-level miss (false Noise classification on a Crisis story = automatic fail) 6. WORD COUNT Criterion: Executive Summary 50–70 words; each story summary 30–60 words Measure: Automated word count Pass: Within range Fail: >20% over or under target | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | SUBJECTIVE OUTPUT EVALUATION — HUMAN JUDGMENT Method: Jane runs 5 test briefs across different days and companies. After each brief, she completes a one-page scoring sheet with 5 questions. Results are averaged across all sessions. SCORING SHEET (Jane completes after each brief): 1. TONE APPROPRIATENESS "Does the language match the severity of each story — urgent where needed, measured where appropriate?" Scale: 1 = tone mismatched or alarming | 5 = perfectly calibrated Target: ≥4.0 average across 5 sessions 2. "SO WHAT" USEFULNESS "Would I actually send this suggested action to my CEO as-is?" Scale: 1 = vague or generic ("monitor the situation") | 5 = specific and immediately actionable Target: ≥4.0 average across all stories 3. EXECUTIVE READINESS "Could I forward the Executive 'So What' story right now without editing a single word?" Scale: 1 = needs significant rewriting | 5 = forward-ready as-is Target: ≥4.0 average across 5 sessions 4. CANADIAN FINTECH CONTEXT ACCURACY "Did the AI correctly read the regulatory or media nuance? (e.g. flagging an OSFI statement as Act, not Watch; treating The Logic as higher-signal than a personal finance blog)" Scale: 1 = significant context error | 5 = accurately applied Target: ≥4.5 — context errors are a trust-destroying failure mode for this product 5. OVERALL BRIEF CONFIDENCE "After reading this brief, am I confident enough to send it without double-checking the original sources?" Answer: Yes / No Target: ≥80% Yes across 5 sessions PASS CRITERIA: All 4 scored criteria ≥4.0 (context ≥4.5) + confidence ≥80% Yes. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | MASTER PROMPT — VERSION 1 (initial) [SYSTEM] You are SignalBrief, an AI media intelligence assistant for Canadian fintech communications teams. Your job is to turn raw news and Reddit signals into a structured daily brief that tells Jane (the comms lead) what happened, why it matters, and what to do. Rules: - Only use sources provided in the input. Never invent facts. - Flag any uncertain classification with ⚠. - If no relevant stories exist, say so explicitly. - Always follow the output template exactly. [USER] Company: {{company_name}} Products: {{products}} Competitors: {{competitor_list}} Priority topics: {{priority_topics}} Brief style: {{brief_length}} Custom instruction: {{custom_instruction}} [omit if blank] RAW NEWS (last 24h from NewsAPI): {{raw_news_json}} RAW REDDIT (last 24h, r/PersonalFinanceCanada + r/CanadianInvestor): {{raw_reddit_json}} TASK: 1. Filter irrelevant and duplicate results. 2. Classify each relevant story: Noise / Watch / Act / Crisis. Add one-sentence reasoning. 3. For each story write: HEADLINE: [title] SUMMARY: [2–3 sentences] WHY IT MATTERS: [reputational implication for {{company_name}}] SUGGESTED ACTION: [specific recommendation] 4. Group by severity tier (Act → Watch → Noise). 5. Write Executive Summary (2–3 sentences, <70 words). 6. Write Executive "So What" Story (1 paragraph for CEO/CMO, plain language). OUTPUT FORMAT: --- SIGNALBRIEF — {{brief_date}} Company: {{company_name}} EXECUTIVE SUMMARY [text] EXECUTIVE "SO WHAT" STORY [text] --- ACT --- HEADLINE: ... SUMMARY: ... WHY IT MATTERS: ... SUGGESTED ACTION: ... --- WATCH --- [same structure] --- NOISE --- [same structure] --- INITIAL TEST RESULTS: Prompt v1 produced accurate structure and relevant filtering. Hallucination rate: 0 in 5 test runs. Main weakness: "Why it matters" sections were occasionally too generic ("this could affect brand reputation") without enough Canadian fintech specificity. → Led to v2 refinement (see Prompt Iterations row). | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | PROMPT ITERATIONS — changes made and rationale VERSION 1 → VERSION 2 (primary iteration) Change: Added explicit Canadian fintech context block to system prompt Rationale: V1 "Why it matters" sections were generic. V2 adds: "You understand Canadian fintech regulatory context (OSFI statements, PIPEDA, Bank of Canada announcements), the tier hierarchy of Canadian media (The Logic and Globe and Mail > BetaKit > personal finance blogs), and the difference between an isolated customer complaint and a media narrative gaining traction." Result: "Why it matters" sections became measurably more specific. Human reviewer agreement on usefulness rating improved from 3.8 to 4.3 average. VERSION 2 → VERSION 3 (edge case handling) Change: Added explicit handling for no-result and low-confidence cases Rationale: V2 sometimes generated plausible-sounding summaries when input was very thin (only a headline, no article body). Added constraints: "If article body is absent, base summary on headline only and flag ⚠ 'Summary based on headline only.'" Also added: "If fewer than 3 relevant stories exist, reduce to a shorter brief and note total stories reviewed." Result: Hallucination risk eliminated in headline-only inputs. Zero unverifiable claims across 10 test runs post-v3. VERSION 3 → VERSION 4 (tone calibration) Change: Added severity-specific tone guidance Rationale: V3 tone was uniform across severity tiers. Act stories sometimes sounded the same as Noise stories. Added: "For Act stories, use direct language with clear urgency but without alarm. For Crisis stories, be factual and specific. For Watch stories, be measured. For Noise stories, be brief and reassuring." Result: Jane's tone appropriateness rating improved from 3.9 to 4.6 average. TRACKING METHOD: Each version stored in a versioned prompt log. Test brief outputs saved with version tag (e.g., brief_v2_test_3.txt). Human evaluation scores recorded per version to track improvement. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | DATA SOURCES & PREPARATION DATA SOURCES: 1. NewsAPI (newsapi.org) — Primary news source - Query: company_name + brand terms, last 24h, English language, Canadian sources prioritised - Endpoint: /v2/everything?q={{company_name}}&from={{yesterday}}&language=en&sortBy=relevancy - Output: title, description, url, source name, publishedAt - Rate limit: 100 requests/day on Developer plan ($449/mo for production) - MVP: Developer plan sufficient for demo 2. Reddit API (via PRAW or Reddit JSON endpoint) - Subreddits monitored: r/PersonalFinanceCanada, r/CanadianInvestor, r/fintech, r/canada - Query: company_name, filtered to posts with >5 upvotes in last 24h - Output: title, selftext (post body), score, url, subreddit - Rate limit: 60 requests/minute — sufficient for daily brief - MVP: Public JSON endpoint (no auth required): reddit.com/r/PersonalFinanceCanada/search.json?q={{company_name}}&sort=new&t=day DATA PREPARATION STEPS: 1. Fetch: Parallel API calls to NewsAPI and Reddit at brief trigger 2. Deduplicate: Remove articles with >80% title similarity (string matching) 3. Filter: Remove results where company_name not in title OR description (reduces irrelevant results) 4. Truncate: Limit article body to first 500 characters (tokens) to fit context window affordably 5. Structure: Convert to consistent JSON schema before injection into prompt RAG APPROACH: No vector database required for MVP — all relevant context fits within Claude's 200K token context window in a single prompt. Future V2 will add RAG for briefing history retrieval (past 30 days of briefs stored and retrieved semantically to enable trend detection). MVP token budget per brief run: ~8,000–12,000 input tokens + ~2,000 output tokens ≈ $0.03–0.05 per brief run at Claude 4-5 Sonnet pricing. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | EXAMPLE INPUT/OUTPUT DATA — TYPICAL CASES TYPICAL CASE 1: Watch story (Reddit thread gaining traction) INPUT (raw Reddit): {"title": "Koho transfer stuck for 3 days — anyone else?", "selftext": "Been trying to transfer $2,000 from Koho to my TD account. Support hasn't responded. Seeing others in comments with same issue.", "score": 89, "subreddit": "PersonalFinanceCanada", "url": "reddit.com/r/..."} Company: "Koho" | Competitor list includes Koho EXPECTED OUTPUT: HEADLINE: Reddit thread flags multi-day Koho transfer delays — 89 upvotes SUMMARY: A Reddit post on r/PersonalFinanceCanada reports a 3-day delay on a $2,000 transfer from Koho to TD, with commenters reporting similar issues. No news coverage yet. Post gaining traction with 89 upvotes. WHY IT MATTERS: A competitor's payment reliability issue in the same customer segment elevates sector-level trust questions. Journalists monitoring Reddit may pick this up. Positions [Company] to emphasise its own transfer reliability proactively. SUGGESTED ACTION: Monitor for 24h. If thread crosses 200 upvotes or gets picked up by BetaKit/The Logic, prepare a proactive statement on [Company]'s transfer reliability and post on social channels. SEVERITY: Watch ⚠ TYPICAL CASE 2: Noise story (sector profile piece) INPUT (raw NewsAPI): {"title": "Canadian fintech sector raises $1.2B in Q1 2026", "description": "A report by MaRS Discovery District highlights record fintech investment in Canada. Several companies including [Company] cited as growth examples.", "source": "BetaKit", "url": "..."} EXPECTED OUTPUT: HEADLINE: BetaKit: Canadian fintech raises $1.2B in Q1 — [Company] mentioned as example SUMMARY: BetaKit reports record Q1 fintech investment in Canada based on MaRS research. [Company] mentioned as a growth example alongside sector peers. WHY IT MATTERS: Positive sector association and name recognition. No brand risk. Good signal for talent attraction and investor awareness. SUGGESTED ACTION: Share internally. Consider resharing BetaKit article on LinkedIn company page to amplify positive association. SEVERITY: Noise | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | EXAMPLE INPUT/OUTPUT DATA — EDGE & NEGATIVE CASES EDGE CASE 1: Ambiguous severity — competitor crisis with no direct company mention INPUT: NewsAPI article — "EQ Bank suffers data breach affecting 12,000 accounts" — Company not mentioned. EXPECTED OUTPUT: SEVERITY: Watch ⚠ (borderline Act) HEADLINE: EQ Bank data breach: 12,000 accounts affected — company not named WHY IT MATTERS: Competitor data breach in the same sector elevates regulatory and media scrutiny for all Canadian neobanks. OSFI and journalists may broaden inquiry to include [Company]. Pattern: sector breaches typically trigger 30-day wave of similar coverage. SUGGESTED ACTION: Proactively review [Company]'s data handling public comms. Prepare a brief internal FAQ for if journalists call. Do not issue public statement yet — monitor for journalist outreach. EDGE CASE 2: Headline-only input (no article body available) INPUT: {"title": "Wealthsimple under investigation by OSFI", "description": null, "url": "..."} EXPECTED OUTPUT: SEVERITY: Act ⚠ (pending confirmation) HEADLINE: ⚠ Wealthsimple OSFI investigation reported — limited source information SUMMARY: ⚠ Summary based on headline only — full article not accessible. Headline indicates a regulatory investigation by OSFI. Unable to confirm scope, timeline, or source credibility without full article. WHY IT MATTERS: If confirmed, OSFI regulatory action is a high-severity reputational and compliance event. SUGGESTED ACTION: Do not share externally until verified. Access full article directly at [URL]. If confirmed, escalate to CMO and Legal immediately. NEGATIVE CASE: No relevant results found INPUT: NewsAPI + Reddit return 0 results for "SignalBrief" in past 24h. EXPECTED OUTPUT: EXECUTIVE SUMMARY No significant brand signals detected for [Company] in the past 24 hours. Sources checked: NewsAPI, Reddit (r/PersonalFinanceCanada, r/CanadianInvestor). EXECUTIVE "SO WHAT" STORY No brand coverage or social mentions requiring attention today. This is a quiet news day — a good opportunity to proactively pitch a story or publish owned content. --- NOISE --- No stories to report. Note: If this is the third consecutive quiet day, consider expanding your monitored search terms or adding additional sources in your profile settings. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | MANUAL REVIEW RESULTS (5 test briefs run during develop phase) Test run methodology: 5 brief runs using real NewsAPI + Reddit data for a Canadian fintech (Wealthsimple used as proxy for testing). Human reviewer (Khushboo) independently assessed each output against the evaluation criteria. RESULTS SUMMARY: Test 1 — V1 prompt, standard input (12 articles, 8 Reddit posts) Structure compliance: PASS (all 5 sections present) Story completeness: PASS (100% of stories had all 4 components) Hallucination rate: PASS (0 unverifiable claims) Relevance rate: PASS (10/11 stories relevant; 1 off-topic excluded manually) Severity accuracy: PARTIAL PASS (1 Watch story classified as Noise — Reddit post with 67 upvotes should have been Watch) Tone: 3.8/5 — "Why it matters" felt generic → FAILED on tone/usefulness → triggered V2 prompt revision Test 2 — V2 prompt, same input Tone: 4.3/5 — Canadian fintech context noticeably improved "Why it matters" specificity Severity accuracy: PASS (4/4 human-AI agreement) Executive "So What" story: 4.1/5 — good but slightly long Test 3 — V3 prompt, headline-only input (3 articles with no body) Hallucination rate: PASS — AI correctly flagged ⚠ on all 3 headline-only stories Format compliance: PASS Test 4 — V4 prompt, high-volume input (28 articles) Filter performance: AI correctly excluded 19/28 as irrelevant; kept 9 brand-relevant stories Processing: Single prompt handled full input without truncation errors Test 5 — V4 prompt, no-results input No-result handling: PASS — correct "No significant brand signals" output with source confirmation FAILURES AND FIXES: - V1: Generic "Why it matters" → Fixed in V2 with Canadian fintech context block - V1: One severity miss (Watch classified as Noise) → Fixed in V3 with upvote threshold guidance in prompt - V3: Occasional verbose Suggested Actions → Fixed in V4 with explicit "one clear sentence" constraint | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | AUTOMATED EVALUATION APPROACH Automated checks implemented for scalable testing: 1. STRUCTURE CHECK (Python script) Method: Regex scan for required section headers in output string Checks: ["EXECUTIVE SUMMARY", "EXECUTIVE "SO WHAT" STORY", "--- ACT ---", "--- WATCH ---", "--- NOISE ---"] Pass/fail: All 5 present = PASS Score across 5 test runs: 5/5 PASS (100%) 2. STORY COMPLETENESS CHECK (Python script) Method: Parse output into story blocks; check each for ["HEADLINE:", "SUMMARY:", "WHY IT MATTERS:", "SUGGESTED ACTION:"] Score across 5 test runs: 43/43 stories = 100% completeness PASS 3. WORD COUNT CHECK (Python script) Method: Count words in Executive Summary section Target: 50–70 words Score: 4/5 within range; 1 brief had 78-word summary (minor over) → fixed with explicit word limit in V4 prompt 4. RELEVANCE SCORING (model-graded) Method: Secondary LLM call (Claude Haiku, lower cost) — pass each story with the company profile; ask "Is this story about [company_name]'s brand, products, or direct competitive context? Yes/No + reasoning" Score across 5 runs: 91% Relevant classification match between Haiku grader and human reviewer 5. PASS/FAIL SUMMARY ACROSS ALL TESTS: Structure: 100% pass rate Completeness: 100% pass rate Hallucination: 0 failures (0% rate) Relevance: 91% (target: 90%) — PASS Severity calibration: 87% human-AI agreement (target: 85%) — PASS Overall: Prompt V4 meets all objective criteria for MVP launch | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | EDGE CASES IDENTIFIED IN TESTING EC-1: Headline-only articles (no body text) Identified: Test 3 — 3 NewsAPI results returned title only; description field null Risk: AI generates plausible-sounding summaries from nothing (hallucination) Status: RESOLVED in V3 — explicit instruction added: "If article body absent, summarise from headline only, flag ⚠" EC-2: Very high-volume inputs (>25 articles) Identified: Test 4 — 28 articles returned from broad search term Risk: Token overflow; AI includes too many stories; brief becomes unreadably long Status: RESOLVED — pre-filtering step added in data preparation (limit to top 20 by relevancy score from NewsAPI); AI also instructed to cap at 10 stories EC-3: Competitor breach with no company mention Identified: Manual review — EQ Bank data breach article did not mention user's company Risk: AI either excludes (misses important signal) or includes (confuses scope) Status: RESOLVED in V2 — explicit competitor monitoring instruction added: "Include competitor stories if they have direct brand implication for {{company_name}}; explain the connection in Why It Matters" EC-4: Ambiguous company name (common word in name) Potential risk: A company named "Wave" or "Float" may return unrelated results (ocean wave news, float therapy) Status: MITIGATED by exclusion_keywords optional field; user can add "ocean, therapy, swimming" to filter out. Full resolution requires future entity disambiguation layer. EC-5: Crisis-level miss (false Noise on Act/Crisis story) Risk: Highest-severity failure mode — AI classifies a Crisis story as Noise, Jane doesn't see it Status: TESTED — no false Noise classification on Act or Crisis stories in 5 test runs. Severity prompt includes: "When in doubt between two tiers, always classify at the higher tier. A missed crisis is far more damaging than a false Act." EC-6: Custom instruction conflicts with safety Risk: User types "ignore all negative stories" as custom instruction Status: MITIGATED — system prompt precedence noted; custom instructions cannot override core brief structure or exclude stories flagged Act/Crisis | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | PROMPT AND SYSTEM UPDATES BASED ON TESTING UPDATE 1 — Canadian fintech context block (V1→V2) Trigger: Generic "Why it matters" sections in V1 — didn't demonstrate vertical expertise Change: Added to system prompt: "You understand Canadian fintech regulatory context: OSFI statements carry high authority and always warrant Act classification if related to data or consumer protection. Bank of Canada announcements affect sector sentiment. BetaKit and The Logic are tier-1 Canadian fintech outlets. Personal finance subreddits often surface customer complaints before media picks them up." Outcome: "Why it matters" usefulness rating improved from 3.8→4.3 UPDATE 2 — Headline-only safeguard (V2→V3) Trigger: Hallucination risk on thin inputs identified in Test 3 Change: Added explicit constraint: "If description field is null or <20 words, base your summary on the headline only. Begin the SUMMARY field with ⚠ 'Summary based on headline only — full article not verified.' Do not infer facts not present in the title." Outcome: Zero hallucinations in headline-only inputs; ⚠ flag correctly triggered on all 3 test cases UPDATE 3 — Tone calibration by tier (V3→V4) Trigger: Jane's tone rating plateau at 3.9; Act and Noise stories sounded too similar Change: Added to system prompt: "TONE GUIDE — Act: Direct, specific, no alarm words ('catastrophic', 'disaster'). Watch: Measured, 'this warrants attention'. Noise: Brief, reassuring. Crisis: Factual, immediate, no hedging." Outcome: Tone rating improved to 4.6 average UPDATE 4 — Word count constraint (V3→V4) Trigger: Executive Summary occasionally ran 75–90 words Change: Added explicit constraint: "Executive Summary must be 50–70 words maximum. Count and adjust before outputting." Outcome: 100% of V4 briefs within word count range | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | AUTOMATED EVALUATION APPROACH CHOSEN METHOD: Hybrid — script-based structural checks + model-graded semantic quality RATIONALE: Pure human review is not scalable beyond testing phase. Script-based checks cover structural/objective criteria (section presence, word count, completeness). Model-graded evaluation (Claude Haiku as grader) handles semantic quality (relevance, severity calibration) at low cost (~$0.001 per brief evaluation). EVALUATION PIPELINE (post-launch): Step 1: Every brief output passed through Python evaluation script automatically after generation Step 2: Script checks: structure compliance, story completeness, word count, hallucination (URL cross-check) Step 3: Haiku grader called for: relevance scoring, severity calibration check Step 4: Composite score computed (0–100) per brief; logged to evaluation dashboard Step 5: If score <80, brief flagged with ⚠ for human review before delivery Step 6: Jane's subjective ratings (1–5 on tone, usefulness, executive readiness) collected via quick in-app feedback after each brief send SCALING TO DIVERSE TEST SETS: - Synthetic test set: 50 pre-built input scenarios (typical, edge, negative cases) run on every prompt version change - Real-world validation: Random sample of 10% of production briefs reviewed by human monthly - Regression testing: All previously-failed test cases re-run on each prompt update to prevent regression | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | EVALUATION FREQUENCY AUTOMATED EVALUATION: Every brief (real-time) - Structural + completeness check runs on every output before display to user - Model-graded relevance and severity check runs on every output - Results logged to evaluation dashboard; composite score stored per brief HUMAN EVALUATION: Weekly (during beta) - Jane rates 3–5 briefs per week on tone, usefulness, executive readiness (in-app 1–5 rating) - Weekly aggregate scores reviewed by PM; any metric below 4.0 triggers prompt review PROMPT VERSION REVIEW: After every 50 production briefs or any failure event - Full regression test suite (50 synthetic inputs) re-run on current prompt version - If any objective criterion drops below target, prompt iteration triggered within 48h POST-LAUNCH CADENCE: Month 1 (beta): Daily monitoring of composite score; weekly human review; bi-weekly PM review Month 2–3: Weekly automated review; monthly human spot-check; prompt updates as needed Ongoing: Monthly automated review; quarterly full evaluation audit; major prompt version with each product update TRIGGER-BASED REVIEW (immediate, regardless of schedule): - Any user-reported "missed crisis" (false Noise on Act/Crisis story) → immediate investigation + prompt fix - Composite score drops below 75 for 3 consecutive briefs → same-day review - New competitor enters market or major regulatory change → prompt context block update within 5 business days | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | TECHNICAL READINESS CHECKLIST APIs & INTEGRATIONS: ✓ NewsAPI: Developer account active; endpoint tested; 100 req/day limit confirmed sufficient for MVP (1 req per brief run per user) ✓ Reddit API: Public JSON endpoint confirmed; no auth required for read-only search at MVP scale ✓ Anthropic Claude API: Account active; Claude 3.5 Sonnet endpoint tested; rate limits (1000 req/min) far exceed MVP needs ✓ EmailJS (or equivalent): Email delivery tested; "Send to my email" button functional with mailto fallback INFRASTRUCTURE: MVP: No-code / serverless deployment. Built in Claude Artifacts (HTML/JS) for capstone demo — zero infrastructure to manage. Production path: Vercel (frontend) + serverless functions (API calls) — cost <$20/mo at MVP scale. Rate limits: All APIs within free/dev tier limits at <50 users. MONITORING: MVP: Manual (PM checks daily). Production: Basic uptime monitoring via UptimeRobot (free tier); API error logging via console. ROLLBACK: MVP: No rollback needed — stateless demo. Production: Previous prompt version stored; API key can be rotated in <5 min. DOCUMENTATION: ✓ API setup documented in internal notion ✓ Prompt version log maintained ✓ Data flow diagram complete (Setup → Fetch → Process → LLM → Display → Deliver) | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | ORGANIZATIONAL READINESS (Note: SignalBrief is a pre-seed startup at MVP stage. Organizational readiness applies to the founding team context.) INTERNAL READINESS: Support: PM (Khushboo) handles all beta user support directly via email. Response SLA: 24h. No support tooling needed at <10 users. Documentation: User onboarding guide drafted (one-pager: how to set up profile, run a brief, interpret severity tiers). Legal: No employee data processed. All data sources are public (NewsAPI, Reddit). Privacy policy drafted (see Data & Privacy row). BETA USER PREPARATION: 5 beta users identified from Canadian fintech comms community (Fintech Cadence, DMZ network). Each beta user onboarded via 20-min 1:1 video call — walk through: profile setup, first brief run, rating flow. Feedback collection: Weekly check-in Slack DM + in-app rating (1–5 per brief). STAKEHOLDER COMMS: Instructor/Demo Day: Full PRD + 4-min video demo submitted by June 15, 2026. Beta users: Expectation set — "This is an MVP. Expect rough edges. Your feedback shapes V2." No external investors or board at this stage. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | LAUNCH APPROACH: Closed Beta → Controlled Launch PHASE 1 — Closed Beta (Weeks 1–4 post-capstone) Who: 5 handpicked Canadian fintech comms leads from PM's network (Fintech Cadence, DMZ, Volta community) Access: Direct invite only; 30-day free trial, no credit card Goal: Validate core value prop ("does the brief save Jane 40 min?"); identify top 3 product gaps; achieve ≥80% brief confidence rating Success criteria: 3 of 5 beta users run brief 5+ times in 4 weeks; average brief confidence rating ≥4.0/5 PHASE 2 — Soft Launch (Month 2) Who: Expand to 25 users via word-of-mouth + Fintech Cadence Slack community post Access: Waitlist signup; 30-day free trial; PM approves each account manually Pricing: $35/month after trial — positioned as "less than a lunch meeting, replaces 15 hours of work" Goal: Test willingness to pay after free trial; target 10 paid conversions Success criteria: ≥40% trial-to-paid conversion; churn <20% Month 1 PHASE 3 — Open Launch (Month 3+) Who: Any Canadian fintech comms or growth lead Access: Self-serve signup; 30-day free trial → $35/mo credit card Distribution channels: BetaKit Startup Spotlight, Fintech Cadence newsletter, LinkedIn (comms community), Product Hunt MRR targets: $350 (10 users) → $1,050 (30 users) → $3,500 (100 users) | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | SCALE READINESS CURRENT STATE (MVP / Closed Beta): Infrastructure: Serverless / no-code stack handles 1–50 concurrent users with zero cost increase. API limits: NewsAPI Developer plan (100 req/day) sufficient for <100 users running 1 brief/day. Upgrade to Business plan ($449/mo) if >100 daily active users. LLM cost: ~$0.03–0.05 per brief run. At 100 users × 20 briefs/month = 2,000 runs/mo = ~$60–100/mo in LLM costs. Covered at 1 paying customer ($35/mo). SCALE MONITORING: Volume tracking: Brief runs/day logged; alert triggered if >80% of daily API limit consumed. Performance: Brief generation time target <30 seconds; monitored via timestamp logging. Error rate: API error rate tracked; >5% errors triggers same-day investigation. SCALE-UP PLAN: 50 → 500 users: Migrate from serverless to lightweight backend (FastAPI on Railway.app); add Redis caching for repeated brand queries. 500+ users: NewsAPI Enterprise plan; dedicated Claude API rate limit agreement; database for briefing history and trend tracking. COST STRUCTURE AT SCALE: At $35/mo with 20 briefs/user/month: LLM cost: $0.04 × 20 = $0.80/user/month API cost (NewsAPI): $0.01/request × 20 = $0.20/user/month Total COGS: ~$1/user/month → 99% gross margin at MVP scale | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | GO-TO-MARKET ASSETS EXTERNAL MARKETING: 1. Landing page (signalbrief.com): One-page site with value prop, demo GIF, pricing ($35/mo), waitlist signup. Headline: "Your AI morning brief. Done in 5 minutes, not 45." 2. Demo video: 4-minute walkthrough — setup → run brief → review severity → send. Hosted on Loom; embedded on landing page. 3. BetaKit Startup Spotlight pitch: 200-word product description submitted for editorial coverage. 4. LinkedIn post series (3 posts): (1) "I spent 45 min every morning on a media brief. Here's what I built to fix it." (2) Brief output example (screenshot). (3) "We're opening 5 beta spots for Canadian fintech comms teams." 5. Fintech Cadence Slack post: Community-native post in #tools channel — short, direct, link to waitlist. PRODUCT EDUCATION ASSETS: 1. Onboarding guide (1-page PDF): How to set up your profile → how to interpret severity tiers → how to use custom instructions. 2. In-app tooltip layer: Tooltip on each UI element for first-time users (e.g., "⚠ means the AI is uncertain about this classification — review the original article"). 3. Severity guide (in-app modal): Visual one-pager: what Noise / Watch / Act / Crisis means and what Jane should do for each. 4. FAQ (5 questions): "How does SignalBrief find news?" / "What if it misses something?" / "Can I edit the brief?" / "How is this different from Google Alerts?" / "Is my data stored?" BETA USER ENABLEMENT: - 20-min onboarding call per beta user - Dedicated Slack DM channel for each beta user (direct PM access) - Weekly check-in template: "This week: ran X briefs. Confidence rating avg: X. Biggest friction: ___. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | STAKEHOLDER & INTERNAL COMMUNICATIONS (Note: SignalBrief is a pre-seed solo founding stage. Stakeholder comms applies to: beta users, course instructor, and potential future investors/advisors.) BETA USER COMMS PLAN: Pre-launch: Personal email invite with 1:1 onboarding call booking link. Expectation set: "MVP — expect rough edges. Your feedback directly shapes the product." Weekly: Slack DM check-in — "How's the brief working? Any misses this week?" + share what was improved since last week. Monthly: Product update email — what changed, what's coming next, how beta feedback was incorporated. Policy: All beta users know they are using a pre-revenue MVP. No SLA beyond "we'll fix issues within 48h." INSTRUCTOR / DEMO DAY COMMS: Capstone PRD submitted: June 15, 2026 Demo video submitted: June 15, 2026 Demo Day presentation: 4-min video + Q&A FUTURE INVESTOR / ADVISOR COMMS (post-capstone): Trigger: 3 paying customers achieved Asset: 1-page investor summary (problem, solution, traction, ask) Channel: Direct warm intros via Fintech Cadence and Product Faculty alumni network LAUNCH ANNOUNCEMENT: Day 1 (soft launch): LinkedIn post + Fintech Cadence Slack Day 7: BetaKit startup spotlight submission Day 30: Product Hunt launch (if 10 paying users achieved — provides credibility) | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | DATA & PRIVACY DATA COLLECTED: 1. Company profile data (company name, products, competitor list, brand terms) — stored in browser session only for MVP demo. No server-side storage. Post-MVP: stored in encrypted database, user-owned, deletable on request. 2. Brief outputs — not stored server-side in MVP. User responsible for saving. Post-MVP: optional briefing history stored in user account (30-day rolling window). 3. Email address (for brief delivery) — used only for brief delivery. Not sold. Not shared. Stored in-app only. 4. Usage analytics — brief run count, rating scores — aggregated and anonymised for product improvement. DATA NOT COLLECTED: - No customer data from the user's fintech company - No PII of the user's customers or end users - No browsing history or session recordings EXTERNAL DATA SOURCES: - NewsAPI: Processes publicly available news articles only. No personal data. - Reddit: Processes public posts only. No personal data accessed. Both sources are publicly available and do not require consent from individuals mentioned in articles. PRIVACY COMPLIANCE: PIPEDA (Canada): Canadian users' data handled under PIPEDA. No sensitive personal information collected. Brief outputs contain public information only. GDPR: Not applicable at MVP stage (Canadian market only). Will apply if expanding to EU. Data residency: MVP — no server-side storage. Production — data stored in Canadian AWS region (ca-central-1) to comply with PIPEDA. PRIVACY POLICY: Published at signalbrief.com/privacy before any user data collection begins. | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | POLICY & COMPLIANCE CONTENT MODERATION: SignalBrief processes publicly available news and Reddit posts. No user-generated content published by SignalBrief to public channels. Content moderation risk is low — the product reads public data and generates internal summaries, not public-facing content. AI output guardrails: System prompt explicitly instructs the model to only summarise factual, verifiable content from provided sources. No opinion generation, no content creation for publication. LEGAL: Terms of Service: Published at signalbrief.com/terms before launch. Key clauses: (1) AI-generated brief is for informational purposes only — not legal or regulatory advice. (2) User responsible for verifying AI output before external distribution. (3) SignalBrief not liable for missed stories or misclassified severity. Copyright: SignalBrief summarises and links to original articles; does not reproduce full article text. Aligns with fair use / fair dealing principles. Defamation: AI instructed to summarise factual content only; never generate editorial opinions about named individuals or companies. REGULATORY (Fintech-specific): SignalBrief is a media intelligence tool, not a financial advice platform. Not subject to OSFI, FINTRAC, or securities regulation. No financial transactions processed. No consumer financial data collected. AUDIT TRAIL: Every brief run logged with: timestamp, input token count, output token count, API sources called. Retained for 90 days for debugging and compliance purposes. Prompt version logged with every brief — ensures reproducibility if a brief is ever disputed. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | SUCCESS METRICS — USER & BUSINESS USER METRICS: 1. Brief Confidence Rating (primary UX metric) Definition: Jane's self-reported confidence after each brief (1–5 in-app rating: "Would you send this brief as-is?") Target: ≥4.0 average across beta users Baseline: Manual brief process — estimated 2.5/5 (sends with residual doubt) 2. Time to Send (core value metric) Definition: Time from "Run Brief" click to "Send to Email" click Target: <10 minutes (vs. 45 min manual baseline) Source: Timestamp logging This is the headline metric — if we nail this, everything else follows. 3. Brief Run Rate (engagement / habit formation) Definition: Briefs run per user per week Target: ≥5 briefs/week (daily habit) Source: Brief run event logging 4. 7-Day Retention Definition: % of users who run at least 3 briefs in their first 7 days Target: ≥60% 5. Trial-to-Paid Conversion Definition: % of 30-day trial users who convert to $35/mo paid Target: ≥40% (strong for B2B SaaS with a clear daily use case) Source: Billing event logging BUSINESS METRICS: 6. Monthly Recurring Revenue (MRR) at $35/mo Month 2 target: $350 MRR (10 paying users) Month 4 target: $1,050 MRR (30 paying users) Month 8 target: $3,500 MRR (100 paying users) 7. Churn Rate Target: <10%/month Early warning: Any churn in Month 1 triggers immediate user interview 8. Net Promoter Score (NPS) Survey sent after 30 days Target: NPS ≥40 9. PMF Signal Measure: "How would you feel if you could no longer use SignalBrief?" (Superhuman PMF survey) Target: ≥40% "very disappointed" — canonical PMF threshold | ||
| AI Metrics | How will you measure AI performance and accuracy? | SUCCESS METRICS — AI PERFORMANCE 1. HALLUCINATION RATE Definition: % of briefs containing at least one AI-generated claim not traceable to a provided source Target: 0% (zero tolerance) Measurement: Human spot-check of 10% of production briefs monthly; user-reported flags Alert: Any confirmed hallucination triggers same-day prompt review 2. RELEVANCE RATE Definition: % of included stories that are genuinely relevant to the company's brand (rated by user) Target: ≥90% Measurement: In-app "was this story relevant?" thumbs up/down per story (optional, non-blocking) Threshold: If relevance rate drops below 85% for a user, trigger profile review (are their brand terms broad enough?) 3. SEVERITY CALIBRATION ACCURACY Definition: % of severity classifications the user agrees with Target: ≥85% user agreement Measurement: In-app "severity wrong?" flag per story (optional) Critical failure mode: Act or Crisis story classified as Noise — tracked separately with 0% tolerance 4. BRIEF GENERATION LATENCY Definition: Time from "Run Brief" click to brief displayed on screen (includes API calls + LLM processing) Target: <30 seconds p95 Measurement: Timestamp logging; alert if >45 seconds Note: NewsAPI + Reddit fetch (~3–5s) + Claude API (~10–15s) = estimated total ~15–20s per brief 5. EVALUATION COMPOSITE SCORE Definition: Automated score (0–100) per brief = structure (20) + completeness (20) + relevance (20) + severity calibration (20) + word count (20) Target: ≥80/100 on every brief delivered to user Measurement: Automated evaluation pipeline (see Evaluation Method row) Alert: Score <75 flags brief for human review before delivery | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | USER SUPPORT & FEEDBACK — SUPPORT CHANNELS BETA PHASE (0–10 users): Primary channel: Direct email (khushboo@signalbrief.com) — PM handles personally Response SLA: 24 hours (Monday–Friday) Secondary: Dedicated Slack DM channel per beta user for quick questions Escalation: All issues handled by founding PM (no escalation path needed at this scale) SOFT LAUNCH PHASE (10–50 users): Primary: In-app feedback button ("Something wrong? Let us know") — routes to PM email + Notion issue tracker Secondary: Email support Response SLA: 24 hours for bug reports; 48 hours for feature requests Knowledge base: FAQ page at signalbrief.com/help (5 core questions; expanded with common support topics) PRODUCTION (50+ users): Primary: Intercom or Crisp in-app chat (live chat during business hours) Secondary: Email support Escalation: P1 bugs (brief not generating, data source failure) → PM + dev within 4 hours Knowledge base: Expanded Notion-based help center with video walkthroughs OWNERSHIP CLARITY: All support tickets owned by PM until first hire. First hire = Customer Success / community manager (at ~50 paying users). | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | FEEDBACK WORKFLOW FEEDBACK COLLECTION: 1. In-app brief rating: 1–5 stars after each "Send to email" click + optional comment field 2. In-app story flags: Thumbs down on individual stories (severity wrong / irrelevant / hallucination) 3. Monthly NPS survey: Sent at Day 30 and Day 90 via email 4. Weekly PM check-in: Qualitative conversation with each beta user (5–10 min Slack DM or call) 5. Feature request log: Notion page shared with beta users — they add requests directly TRIAGE PROCESS: P1 — Critical (product broken, brief not generating, data source failure): Investigate within 2 hours; fix within 24h; notify affected users within 4h P2 — High (hallucination confirmed, severity miss on Act/Crisis): Investigate within 24h; prompt fix within 48h P3 — Medium (relevance issues, tone complaints, UI bugs): Log in backlog; address in weekly sprint P4 — Low (feature requests, enhancement ideas): Add to roadmap; acknowledge in monthly update FEEDBACK → PRODUCT LOOP: Weekly: PM reviews all in-app ratings + flags. Any P1/P2 triggers immediate action. Monthly: Aggregate feedback review → top 3 themes → inform prompt iteration or feature priority Quarterly: Beta user group retrospective call — "what changed, what still needs fixing, what's next" COMMUNICATION: Users notified of fixes within 48h of resolution: "We fixed the issue you reported on [date]. Here's what changed." This closes the loop and builds trust. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | MONITORING & LOGGING — POST-LAUNCH OPERATIONAL MONITORING: 1. Uptime monitoring: UptimeRobot (free) pings signalbrief.com every 5 minutes; alert via email if down >2 min 2. API health: Each brief run checks NewsAPI + Reddit + Anthropic API response status; failure logged and surfaced to user ("NewsAPI is currently unavailable — try again in 5 minutes") 3. Brief generation latency: Timestamp logged at run_start, api_fetch_complete, llm_response_received, brief_displayed. p50/p95 tracked weekly. 4. Error rate: API errors logged per source per day; alert if error rate >5% in any 1-hour window AI PERFORMANCE MONITORING: 5. Evaluation composite score: Logged per brief (structure + completeness + relevance + severity). Weekly average tracked. 6. Hallucination flags: User-reported flags logged in Notion; reviewed within 24h 7. Severity miss flags: Any story flagged "severity wrong" logged with brief_id, story_id, user_id, AI classification vs. user rating 8. LLM cost per brief: Token counts logged per run; monthly cost summary tracked against revenue DASHBOARDS: Beta phase: Manual — PM reviews Notion log + email alerts daily Production: Simple Metabase dashboard (connected to logs database): daily brief count, avg composite score, avg rating, error rate, LLM cost ALERT THRESHOLDS: - Uptime <99% in any 24h period → page PM - Composite score <75 for 3+ consecutive briefs from one user → PM reviews user's profile setup - Any confirmed hallucination → immediate P1 escalation - LLM cost >120% of projected monthly budget → review token usage and add truncation | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | ONGOING IMPROVEMENT — CONTINUOUS LEARNING LOOP LEARNING COLLECTION: 1. Weekly: Review all in-app brief ratings + story flags. Identify patterns (e.g., "4 users this week flagged Reddit stories as irrelevant — Reddit filter may be too broad") 2. Bi-weekly: Run full evaluation suite (50 synthetic test cases) on current prompt version — check for drift 3. Monthly: Aggregate NPS data + qualitative interview synthesis → top 3 themes 4. Quarterly: Full prompt audit — review all 4 prompt versions, current production version, and user feedback patterns IMPROVEMENT TRIGGERS: Prompt update: Any of — (a) composite score drops below 80 on average, (b) user relevance rating drops below 4.0, (c) new regulatory event requires context update (e.g., new OSFI guidance), (d) new major Canadian fintech media outlet to add to context Data source update: If NewsAPI coverage gaps identified (missing outlets) → add supplementary source (e.g., Google News RSS for specific outlets) UI update: Driven by friction patterns in time-to-send data (if users consistently take >15 min, identify which step creates friction) POST-LAUNCH ROADMAP (driven by learnings): V2 (Month 3–4): Briefing history + week-over-week trend detection (resolves Pain Point P10) V3 (Month 5–6): Competitor intelligence layer — monitor top 3–5 competitors in same brief (resolves P9, P11) V4 (Month 7–8): Proactive Slack alert for Watch/Act stories breaking mid-day (resolves P2, P14) V5 (Year 2): US fintech expansion — new vertical context block; similar Canadian fintech community distribution strategy applied to US (Fintech Meetup, a16z fintech community) IMPROVEMENT CADENCE SUMMARY: Every brief → automated evaluation score Weekly → PM review of ratings + flags Monthly → aggregate analysis + prompt consideration Quarterly → full audit + roadmap reprioritisation Per-event → immediate response to P1/P2 issues | ||||




