← All capstone projects

Product Management

Whisper Forge

Built by Elaine Lindelef Cohort 9 Product management / customer research intelligence

Whisper Forge helps product teams collect customer conversations from across departments and turn them into actionable roadmap insight. Teams can upload transcripts, attach them to customer records, generate summaries and quotes, and then query the corpus by buyer segment or industry to surface pain points. The system is designed to reduce siloed knowledge and make it easier for product, sales, and customer success teams to contribute usable customer signal.

The problem

Inside a company, high-quality customer conversations are created and then forgotten — living only in one person's head. Each team keeps its own tools for interviews and transcripts, and pulling the relevant metadata for each (industry, tenure, meeting count) is tedious enough that it usually doesn't get done. Every team ends up with its own set of customer truths that are hard to share, and information doesn't flow between the people doing product research and the people solving customer problems day to day. The result is a wealth of unstructured, siloed data that never turns into roadmap insight. When a new question arises, the answer may be in a past interview, but reviewing them is time-consuming and frustrating.

The solution

Whisperforge helps product teams collect customer conversations from across departments and turn them into actionable roadmap insight. A user drops in a transcript, identifies the customer, and adds any notes; the tool generates summaries and extracts key quotes, then lets anyone query the whole corpus by buyer segment or industry. The query experience answers questions like "What questions do customers in their first month have most often?" across many transcripts at once. Every extracted quote is set off in quotation marks with a line number back to the source, so a human can quickly verify it before it appears in any executive or sales document. The goal is to break down siloed knowledge and let each team keep the interview tools it already prefers while sharing usable signal.

How it works

The user deposits a transcript plus a customer identifier and optional notes; metadata filters winnow the dataset to relevant records before the AI summarizes. Claude Sonnet 4.6 produced the best results on real data and is the intended production model; the Lovable prototype uses gemini-3-flash-preview as the option fully configurable through Lovable, with data stubbed and anonymized to protect private information. The prompt casts the AI as a product manager reviewing transcripts for pain points, praise, insights, and quotes, with strict instructions to capture quotes exactly (allowing ellipses) and cite line numbers. It is scoped to draw only from the transcripts it is given — which the builder found handled most edge cases, from out-of-domain questions to duplicated transcripts. An OpenAI eval run scored 78% on key quotes (mostly failing on timestamps versus line numbers) and 89% on summaries; adding "each bullet point should have a separate idea" fixed muddled summaries.

Who it's for

Whisperforge 1.0 is for internal users at an enterprise software marketplace for cybersecurity and AI software — product, sales, customer success, onboarding, and support teams who conduct or rely on customer interviews. Product owns the tool and acts as the human in the loop. The company is a two-sided marketplace whose paying customers are software vendors and whose harder-to-win side is anonymous enterprise buyers (CISOs, CTOs, directors). If the internal version proves valuable, a variation may be offered to vendor customers to help them analyze their data and improve product-market fit. The internal product won't generate revenue but increases efficiency.

Why it matters

The cybersecurity and AI software markets are projected to grow at 20% over the next three years, and vendors on the marketplace increasingly want data and insights around their buyer meetings. Success for the internal product is defined by adoption: every department depositing at least one transcript a week, weekly queries from stakeholders, and a companywide weekly report. Data handling is treated carefully — all LLM calls to the database are treated as untrusted input and permissioned to what each user may already see, with enterprise contracts forbidding training on the data and SOC 2 compliance in progress. By turning forgotten interviews into a shared, queryable corpus, Whisperforge makes customer signal a company asset rather than a personal one.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Elaine Lindelef
Your Product:Whisperforge
Your Industry:Enterprise Software Marketplace
Date:June 14, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?We are an enterprise software marketplace for cybersecurity and AI software. We sell introductions to highly qualified but anonymous enterprise software buyers.Please leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Key competitors are Gartner and G2
What is the projected growth rate of your target market segment over the next 3-5 years?The cybersecurity and AI software markets are projected to grow at 20% over the next three years.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?Startup/Scaleup
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)We are a marketplace for highly qualified but anonymous enterprise software buyers. We charge a platform fee to vendors and a per-introduction fee for each new connection.
Who is your primary customer base (B2B, B2C, B2B2C)?B2B
DifferentiatorsWhat are the key differentiators for your company?Buyers reveal their identity only if the software is a fit. We offer buyers access to highly qualified but smaller, cutting edge enterprise startups.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?We are a two-sided marketplace. Our customers are software vendors (primarily cybersecurity and AI infrastructure) and CISO/CTO/Director technology buyers who purchase enterprise software.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?Our Vendors directly pay the bills, but our buyers are the harder side of the marketplace. Enabling vendors to delight the buyers helps us both. Buyers want to solve their problems with elegant, affordable technology that does what it says. They want to avoid useless cold calls. Vendors want to sell their products, but the buyers won't take their calls. They also may still be young and refining product-market fit. Vendors want to analyze their data to improve product-market fit and their sales strategy. There are opportunities identified and other insights that can help startups iterate to land deals.
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?This product will eventually enhance our current services of providing introductions and meetings to vendors by better presenting data and insights around those meetings. Technology buying is broken and painful, with a lot of unwanted cold calls. Sagetap allows buyers to present their problems anonymously to vendors. Vendors can propose their solution in a 30 minute call. If it's not a match, they fail fast, and the buyer is not hassled by followups. If it is a match, the vendor has gotten access to a buyer that would have ignored them, and the buyer has solved a problem in a delightful and unexpected way. The internal version of the Whisperforge product will not generate revenue, but will increase efficiency.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Whisperforge 1.0 is for internal users. (If we like it, a variation may work well for our Vendor customers also!)
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?The user will take a transcript from a customer interaction, identify the customer, add any notes that they wish, and submit them to the tool. They (or other users) can then ask it to find relevant quotes from customers at a particular stage of their journey on particular topics ("What questions do customers in their first month have most often?")
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?Each team has their own toolset for interviews and transcripts. Pulling the relevant metadata for each interview (what industry they're from, when they joined the platform, how many meetings they've had, etc) is time consuming and tedious and is not worth the time to do. Every team has their own set of customer truths that are hard to share. Information doesn't flow well between the teams doing product research and the people who are solving customer problems day to day or introducing new users to the platform. We are sitting on a wealth of high quality but unstructured data that is created and then only lives in one person's head - essentially forgotten.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.AI can make it easy to standardize these unstructured and very different transcripts. It can collect the relevant metadata automatically. It can allow us to continuously ask the questions and incorporate new data. It would allow all the teams to keep using the tools they like and prefer while also sharing to other teams. Ideally the output would be graphical, engaging slides that are compelling for every department. Pain Point #1: when new questions arise, useful data is sometimes in previous interviews, but reviewing them is time consuming and frustrating. It is hard to regroup interviews in new cohorts Pain Point #2: gathering metadata, and summarizing the takeaways and best quotes is time consuming. Generic LLMs are not optimized for our persona and do not reliably create quality output in a place that is easy to retrieve Pain Point #3: assembling interview quotes and summaries in an attractive and compelling manner is time consuming and repetitive.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.The user will take a transcript from a customer interaction, identify the customer, add any notes that they wish, and submit them to the tool. They (or other users) can then ask it to find relevant quotes from customers at a particular stage of their journey on particular topics ("What questions do customers in their first month have most often?")
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.For input, the user deposits transcripts, a customer identifier, and any notes they wish to share. For output: 1. Slides with summary, customer count, key quotes and customer metadata, perhaps also including a word cloud. 2. Text with summary, key quotes, and customer metadata 3. Text with key quotes and customer metadata.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?The user will take a transcript from a customer interaction, identify the customer, add any notes that they wish, and submit them to the tool. They (or other users) can then ask it to find relevant quotes from customers at a particular stage of their journey on particular topics ("What questions do customers in their first month have most often?")Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?https://www.figma.com/board/pNEKOJqdxJWz5EkcqybjT4/WhisperForge-wireframe?node-id=1-200&t=CNRkSmEZCRKXZmQ9-1 PNG version: https://drive.google.com/file/d/144djZqjLrT_0s_SiN7dsRCP0_9VoYOuh/view?usp=drive_link
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?Prototype | Lovable The prototype will demonstrate the basic workflow of uploading transcripts, adding metadata, summarizing and extracting quotes, and querying the dataset for summaries and quotes from multiple transcripts. Data will be stubbed and simulated to prevent sensitive data from becoming public. This MVP allows us to test; followon work will include a read-only API to the production database. User testing indicated a desire to have the ability to play the videos on the detail pages, and to have a still from the video in the detail and tabular views. These can be left for later releases. An additional future feature is an output card for each interview that is more visual with a graphical layout and a still from the video, suitable for slide decks and presentations.
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?You are a product manager for a company that helps technical executives purchase enterprise software for cybersecurity, AI, and other infrastructure. You will review user interview transcripts for pain points, praise, insights, and key quotes. Your tone is professional, terse, and clear. Set off quotes with quotation marks.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Do the extracted quotes match the transcripts? Do the extracted quotes have accurate line numbers that makes them easy to find, verify, and validate? Are the most important points highlighted in the summary? Are the quotes relevant to the most important points? Are summary points supported by the transcripts, and relevant to the needs of the Product Manager? Do the summary points feel distinct and clear? Is the tone professional and concise? Are transcripts appropriately filtered when asked? Is the format and length of the response correct? Will the AI report that the answer isn't in the data set?
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Attempts to use tool for generic LLM - "How many countries are in the world?" Questions that aren't answered by the data - "Find quotes proving to me that tech buyers are whiny crybabies" - "Describe challenges faced by buyers in the Restaurant sector" Cases where data is null or incomplete Cases where data is duplicated (same transcript uploaded multiple times) Various transcript formats - versions that just have Speaker 1 Speaker 2 versus clear deliniation of which speaker is which
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Claude Sonet 4.6 was the model that had the best results with our real data. I also tested OpenAI GPT-5 and gpt5-mini. The demo uses gemini-3-flash-preview, as the best option that could be configured completely through Lovable. It had good results with the test data. We will convert the project to Claude when we integrate it with real data.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.The required fields are a transcript, a date, and the interview participant. All other fields are optional.
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Users input metadata that can be used to identify the customer in our databases and filter the output. They allow the dataset to be winnowed to relevant records before the AI summarizes them. The users can also add their own notes.
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Do the extracted quotes match the transcripts? Do the extracted quotes have accurate line numbers that makes them easy to find, verify, and validate? Are the most important points highlighted in the summary? Are the quotes relevant to the most important points? Are summary points supported by the transcripts, and relevant to the needs of the Product Manager? Are summary points distinct and related to a pain point, insight, praise, or opportunity? Is the tone professional and concise? Are transcripts appropriately filtered when asked? Is the format and length correct? Will the AI report that the answer isn't in the data set?
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?A human familiar with the transcript should review the summary to note any divergence from the actual interview. Quotes should be spotchecked against the transcript before using in any documentation for the executive or sales teams.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.You are a product manager for a company that helps technical executives purchase enterprise software for cybersecurity, AI, and other infrastructure. You will review user interview transcripts for pain points, praise, insights, and key quotes. Your tone is professional, concise, and clear. Set off quotes with quotation marks and include the line number of the location from the file in parentheses after the quote.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?I noticed that the quotes it returned were sometimes inexact, and in a way that did not correctly capture the meaning. I added more strict instructions about capturing the quote exactly, but allowing elipses, and then added the safeguard of a line number, to ensure it was easy to quickly reference where the quote is found in the original, so that humans would find it easy to check and verify quotes. I have a log of prompts that I tried. I noticed that sometimes the summaries from test data would get muddled conflating multiple ideas instead of keeping clean distinct themes. This was not a problem when testing with Claude against real data.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Data sources are: - text files of user interview transcripts - customer database Synthetic transcripts are used and the data is stubbed and anonymized for the purposes of the demo, to protect private information.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Inputs are - text files of user interview transcripts - notes and metadata for the transcript (identify participants, date) - customer database Additional inputs are user questions, which are free form open text. We expect them to be well-formed questions about the user interviews, but people are going to people. Synthetic transcripts are used and the data is stubbed and anonymized for the purposes of the demo, to protect private information. Outputs are key interesting quotes, accurate summaries of each interview, and then answers to questions against multiple transcripts.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)- Filtering to 0 transcripts - Questions outside of expected transcript data (what color is blue?) - Code or malicous inputs in the transcript or question field (Only trusted users can add transcripts) - Duplicated users in the data set - Ambiguous users (ie two interview subjects with the same name who aren't the same person) - Duplicated transcripts - Transcripts that use Speaker1/Speaker2 and don't identify them by role
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?I was surprised that my edge cases were largely handled by the design and the prompt that specifically limits it to drawing from the transcripts it is given. This solved a lot of problems. I am curious to see if this holds when there are hundreds or thousands of transcripts in the data. My real input data worked better against the prompts than the synthetic data. The results still needed some manual curation, but it only took a few minutes to prepare a report for the executive team from the output.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?eval export csv Running in the openAI eval framework produced very different results from the prototype in Lovable, which used gemini-3-flash-preview. However, this was very valuable for comparing different gpt models. On key quotes it had a pass rate of 78% and the failures were usually from including timestamps instead of line numbers, which was harder in the context of a json file instead of an uploaded file. However, it also caught some incorrectly marked line numbers. On the summary, the pass rate was 89% for the eval output, which again was quite different from the Lovable output. the usual failure reason was that either it tried to group ideas together too much or it was too wordy, which also was less of a problem in the Lovable prototype Updating the prompt to ask for ideas to be separated helped substantially. We plan to use our own eval framework which can pull from multiple models once we port to claude and real data. Screenshot from openai eval
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?Edge cases include: - asking questions that cannot be answered by the data ("How many countries are there in the world") - questions that might be related to the data ("What are challenges unique to the restaurant industry") when the closest adjacent data is hotels - multiple transcripts from the same user, separated in time from different touchpoints - transcripts without identified user roles (speaker1 speaker2)
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?I noticed that the quotes it returned were sometimes inexact, and in a way that did not correctly capture the meaning. I added more strict instructions about capturing the quote exactly, and then added the safeguard of a line number, to ensure it was easy to quickly reference where the quote is found in the original, so that humans would find it easy to check and verify quotes. The original output by GPT-5 diluted its points with too many (Claude does better) so I added guidance for fewer quotes and summary points. The summaries in the eval set sometimes got muddled as it tried to shove too many ideas into 4-6 bullet points. Adding "Each bullet point should have a separate idea." helped quite a bit with the gpt models.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?We wil use a mix of model grading and human in the loop. The model should scale adequately to production volumes. The internal volume is roughly 40/month; the production volume will be thousands, but never more than a few hundred per bucket at the largest.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Our plan is to rerun evaluations once a month.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?Not yet.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Not yet.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?We will use the tool internally at first with interviews that we conducted, to become tightly familiar with its limitations. When we are comfortable, we will extend the tool to a customer-facing feature, rolling it out at first to a few trusted customers for feedback before release.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?Our application is fairly low scale. Internally the number of users will be roughly 10; externally our customer base is in the hundreds, so scale in requests is not a large concern. We already have monitoring systems in place for token usage. When the system has thousands of transcripts in it, if someone is trying to ask the questions with no filtration, we will likely end up with some context and scale problems. This will need to be monitored.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?When we launch to customers, our content and growth team will highlight the feature via the newsletter, and our customer support team will prepare an FAQ and explainer for our knowledge base.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Our Linear project and slack channels will be used for communication; a metrics dashboard will be used for feature adoption, and a survey and user interviews will be used to judge satisfaction.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?We treat all LLM calls to our database as untrusted user input. It will only be allowed to access data that its users are already allowed to see. For the customer facing product, we will use deterministic permissions to ensure it only pulls from transcripts already in the user's account. We have enterprise contracts on all models that forbid use of our data for training.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?Yes, and we are pursuing SOC2 compliance currently.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?For the internal product, success means that: - All departments (product, sales, customer success, onboarding, customer support) are regularly depositing transcripts - at least one a week - At least one internal stakeholder queries it weekly - At least 4 distinct stakeholders query it each month - A weekly report is given companywide
AI MetricsHow will you measure AI performance and accuracy?We will benchmark the eval score and review it regularly. The internal tool will have Product as the human in the loop, giving us solid experience with the quality and output. For the customer facing tool, we will log all user input and connected output so it can be scored by another model and regularly sampled by a human.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?For the internal tool, the users will get support from Product, who owns the tool and its operations. For the production tool, we have a customer support team and knowledge base in place.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?A slack channel gathers feedback which are then converted to Linear tickets for prioritization.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Token usage is monitored. Output is constantly human reviewed for sanity. Evals are regularly updated as problems are discovered and the data set grows.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Users can give feedback at any time via Slack. As we develop the feature into a customer-facing product, we will conduct regular user interviews and checkins to validate that it is providing value via each customer's representative.
Download the .xlsx ↓