← All capstone projects

Public Library

Dwey

Built by Rushina Bhansali Cohort 9 Public library / reader discovery

Dwey is an AI library assistant that helps readers discover books through natural language and voice instead of rigid keyword search. The product is designed for library patrons who want recommendations that better match age, genre, and nuanced reading preferences. The demo shows personalized recommendations, title explanations, and downstream actions like adding items to a hold cart or wish list while learning from reading history over time.

The problem

Library patrons discover books through rigid, keyword-based catalog search that returns poor matches and offers no semantic understanding of theme, tone, or pacing. There's no recommendation tool to discover reads by what someone actually likes, no user ratings to judge a book's quality, and the family library account can't track an individual reader's history or maintain a member-specific list. General discovery tools like Goodreads, StoryGraph, and BookTok focus on popularity and live outside the library, disconnected from whether a title is actually available nearby — leaving patrons to bridge that gap manually.

The solution

Dwey (Doo-ee) is an AI library assistant that lets readers discover books through natural language and voice instead of keyword search. Patrons describe what they want — or ask for something like a book they loved — and receive age- and genre-appropriate recommendations filtered to titles available in their county's libraries. Each recommendation shows title, author, cover, genre, availability, nearest library and proximity, metadata, and a sourced user rating, with a short personalized explanation. Patrons can act by voice: add to wishlist, place a hold, or reserve a checked-out copy. Profile-based accounts track each family member's reading history to sharpen future recommendations.

How it works

Dwey uses an OpenAI GPT-5.5 model for recommendation reasoning and light agentic tool-calling (holds, wishlist, ratings), with OpenAI's Web Speech API for speech-to-text and gpt-4o-mini-tts for text-to-speech. A vector database of 50–75 books enables semantic retrieval; user and library data live in SQLite, and a FastAPI service pulls ratings and rating counts from the Google Books API. The system prompt casts the assistant as a librarian speaking in first person, adapting tone to the reader's age and prioritizing age, genre, availability, then proximity within a 10-mile radius. Strict guardrails require using the reviews tool for every recommendation and prohibit recommending books outside the inventory, outside the genre without asking, already read, or fabricated. A manual golden-set audit found an 80% baseline success rate; the main failure mode — ignoring a lower age floor and repeating titles on "more" requests — was fixed via a persona-first prompt structure and a check-before-guessing rule.

Who it's for

The business is B2B SaaS, sold to library systems on tiered annual licensing scaled by branches or cardholders, with a future affiliate revenue stream for titles the library doesn't carry. Buyers span public, academic, K-12 school, special, and national/state libraries. End users are the patrons those libraries serve: general community members (children via parents, youth, adults, seniors), students and faculty, and professionals in legal, medical, and government fields. The assistant adapts recommendations and tone across these age groups and reading needs.

Why it matters

Libraries are in a wave of modernization, with cities and counties committing multi-million-dollar capital packages to renovate branches and add technology, fueled by federal, state, and local bond funding. Dwey positions itself as "Librarian-as-a-Service" — an intent-based discovery layer that improves patron engagement and maximizes collection utilization across roughly 9,000 U.S. public library systems and tens of thousands of school libraries. Unlike popularity-driven competitors, Dwey combines semantic understanding with real library availability. Launch targets a limited pilot with a county library cohort to refine the recommendation logic before public release, with an evaluation roadmap moving toward assertion testing against the live catalog and LLM-as-a-judge grading against a 50+ persona golden dataset.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Rushina Bhansali
Your Product:Doo-ee: Personalized Context-Aware Book Recommendation Assistant
Your Industry:Education Tech
Date:June 15, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?Education TechPlease leave this area blank. This space is for the Instructor to provide you with feedback.
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds: 1. Strong competitors - Amazon Books, Goodreads, StoryGraph, BookTok Tailwinds: 1. Most recommendation systems focus on popularity. My tool focuses on context-aware discovery using semantic understanding of themes, tone, pacing, and reader preferences combined with book availability in their local library. 2. LLMs allow for semantic search and reasoning (LLMs can understand tone, pacing, theme; traditional recommendation systems rely on genres, ratings) 3. People are constantly looking for book recommendations. Key Competitors: - StoryGraph, Goodreads, Booktok - Springshare, BiblioCommons, Weblinx, Library Siteworks - Open source content management systems for their websites such as Drupal, WordPress, Joomla, Omeka. - Libraries also sometimes use custom-built portals or local contractors rather than a single off-the-shelf vendor.
What is the projected growth rate of your target market segment over the next 3-5 years?Cities and counties are committing large capital packages to build or replace central libraries and renovate branches. These projects are often multi‑million dollar investments to modernize space, add technology, and expand programming capacity. Federal and state funding opportunities and local bond elections are fueling a wave of modernization projects nationwide. Public library systems - 9K (17K outlets) School libraries (K-12) - 50K-60K Special libraries (corporate, law, medical, museum, government) - 5K-10K Academic libraries (college and university) - 3K-4K National/State Libraries - 1 National; 50+ State End users - The US population growth rate over the next 3-5 years is expected to grow by 0.2% to 0.5% per year (births and immigration). Age group - Approx. share of U.S. population Children (0–17) - 21.5% Youth / young adults (18–24) - 9% Adults (25–64) - 48.5% Seniors (65+) - 18%
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?start up
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)The business generates revenue by providing an intelligent, AI-driven discovery interface that improves library patron engagement and maximizes collection utilization. We monetize the efficiency and "discovery-as-a-service" value provided to library systems and their users. We sell a specialized "Librarian-as-a-Service" interface that enhances the existing library catalog experience, moving away from static, keyword-heavy search toward intent-based, conversational discovery. Revenue model: Our primary revenue model is B2B SaaS (Software-as-a-Service) Licensing. Tiered Licensing: Libraries and library consortia pay an annual recurring licensing fee for access to the AI interface. Pricing is scaled based on the institution's size (e.g., number of active branches or total registered cardholders). Ancillary Transactional Revenue (Future Phase): As a secondary stream, the platform can incorporate an Affiliate Model for titles requested by patrons that are unavailable in the library’s inventory. By providing seamless links to partner booksellers, we earn a commission on those external acquisitions, effectively serving as a "bridge" between library services and retail procurement.
Who is your primary customer base (B2B, B2C, B2B2C)?Primary customer base is B2B.
DifferentiatorsWhat are the key differentiators for your company?A personalized book recommendation assistant that allows the user to find their next read. Replaces the traditional search/advanced search functionality with an easier and intuitive way to discover books. Integrating book recommendations with library availability Maintaining profile based reading history to allow better future recommendations.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?Public Libraries (serving the general community) Academic Libraries (colleges and universities with curriculum-related collections, journals, research databases, course reserves, and study spaces (campus libraries, law/school-specific libraries) School (K-12) libraries (elementary, middle and high schools) Special libraries (corporate libraries, law libraries, medical/health libraries, government and legislative libraries) National / State libraries and archives (country’s or state’s published works, historical records, and official documents (e.g., Library of Congress, state archives). Institutional libraries - inside hospitals, museums, nonprofits, research institutes etc.
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?General community members - Children (Parents), Youth, Adults School, college or university's students, faculty and staff Professionals (legal, medical, government)
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?Core Features: 1. Personalized Content (Books/audio-books/Recommendation Assistant (User Need: Allows users to identify their next read based on what is immediately available) 2. Easier and intuitive content discovery (User Need: Current search and advanced search options are traditional/keyword based and have limited search criteria. Not good at semantic search or to search for books based on other book examples that the user may have read) 3. Profile based accounts to track each individual's reading history (if multiple members in family) 4. Using an individual's reading history to drive future recommendations. 5. Capture user feedback for the books that were checked out upon return to drive future recommendations.
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)General community members - Children (Parents), Youth, Adults and Seniors. School, college or university's students, faculty and staff. Professionals (legal, medical, government).
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?The user logs in to their library website. (general account for the family) The user navigates to the 'Catalog' section to look for books. The user uses the general search or advanced search to look for books. The user finds books based on the search results. The user puts a hold on the books they would like to read.
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?1. The current search funactionality does not return relevant books based on the criteria entered. 2. The current search functionality is keyword based and does not allow for semantic search. 3. The current website does not provide a recommendation tool that allows a user to discover content based what they like to read (theme, age etc.) 4. It is difficult to tell whether a book is good or not as there are no user ratings or rating count. 5. The library account is for the family and there is no profile based account setup to provide personalized content. 6. There is no way to track a specific user's reading history as the website only shows what books were checked out by the family. 7. The website allows to add books to a list which is a common family list and not member specific.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.Pain-points that can be addressed using generative AI. 1. Book recommendation tool to discover their next read using natural language processing and STT/TTS. Allows for semantic search. 2. Book availability in libraries within the county based on geographic proximity. 3. Capture Book rating and rating count using APIs 4. Identify user demographics up-front based on user profile during recommendations. 5. Consider past reading history as well as the user feedback about the books to drive future recommendations. 6. Support other tasks via voice commands - put books on hold, reserve a copy, add to wishlist. Pain-point Ranking: 1. The current website does not provide a recommendation tool that allows a user to discover content based what they like to read (theme, age etc.) 2. The current search funactionality does not return relevant books based on the criteria entered. 3. The current search functionality is keyword based and does not allow for semantic search. 4. It is difficult to tell whether a book is good or not as there are no user ratings or rating count. 5. The library account is for the family and there is no profile based account setup to provide personalized content. 6. There is no way to track a specific user's reading history as the website only shows what books were checked out by the family. 7. The website allows to add books to a list which is a common family list and not member specific.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.1. Build an AI assistant that will recommend books. 2. Add NLP and STT/TTS to easily allow users to talk to their assistant to find books. 3. Build profile based account setup to allow display of personalized content and recommend books based on user profile information without having to ask the user for a lot of details. 4. Integrate with library content inventory (knowledge base/RAG) to only recommend books that are part of the inventory also taking into account book availability and libraries based on geographic proximity. 5. Agentic workflows to put books on hold, reserve a copy. 6. Capture user feedback about the books when the user returns the books. (Prompt when the user logs in to find more books a1fter the previous books were returned) 7. If book is not part of the library inventory, provide information about where the user can buy the book (books tores/pricing information); add workflow for the book to be considered to be made part of the library inventory.
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.1. Build an AI assistant that will recommend books. 2. Add NLP and STT/TTS to easily allow users to talk to their assistant to find books. 3. Integrate with library content inventory (knowledge base/RAG) to only recommend books that are part of the inventory also taking into account book availability and libraries based on geographic proximity.
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?1. The family creates a library account and adds family members to the account. (Profile based sign-up) 2. After the account is setup, the user logs in to the library site. 2. The user is displayed the family dashboard page. The user selects the user profile of the person looking for books to read from the family dashboard. 3. Selecting the user profile will take the user to the 'Recommendations' page where the user can interact with the book assiatant Doo-ee to discover reads. 4. The user can ask for more recommendations or select a book to add to wishlist, place the book on hold or reserve a copy of a book that is currently unavailable. 5. The user can navigate to wishlist, reading history and other navigational items.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?Initial prototype in lovable: https://county-shelf-buddy.lovable.app Moved lovable codebase to Codex to link backend with front end and make other changes. Link to the application (Go to Sign up - Use Jetson family): https://book-recommendation-ai-capstone.vercel.app
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?1. Ask the AI book Assistant Doo-ee for recommendations using natural language processing and STT and TTS. 2. AI recommending books based on user's age/genre and availability at the primary library and other nearby libraries in the county. 3. Ask the assistant to add a book to the wishlist and place a hold. Book Recommendations, adding books to wishlist, placing holds on books and reserving a checked out copy are essential for launch. Later releases - Seamless links to partner booksellers to purchase books if not available in the library. Usage analytics dashboards that help library administrators make data-informed decisions about their collection development based on real-time patron intent.
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?a. You are a librarian with expert knowledge of all the books and resources available in all the libraries for the user's county. b. You must talk in first person as if you are having a conversation with the user. c. Your tone must be based on the user's age. (Children: 0–12 years, Teens/Adolescents: 13–19 years, Adults: 18–64 years, Seniors: 65+ years) d. You will recommend books to users only based on the book inventory of all the libraries in the county. e. You will recommend books based on the user's age and genre. f. You must always show 4 recommendations at a time. g. You must show the following details for each book you recommend - Title, Author, Cover Image, Genre, Availability, Library and proximity to the user's address, metadata describing the book and user rating h. When recommending books, you will prioritize based on following sequence: 1. User's age 2. Genre 3. Availability - books that are available now 4. Geographic proximity - libraries closest to the user's address i. You must limit the libraries within a 10 mile radius of the user's address. j. You must go to the web only to get the book rating. k. You must provide the source for the book rating. l. Do not recommend books that the user has already read. m. Do not recommend books outside of the genre without asking the user. n. Do not recommend books that are not part of the county's library inventory. o. Do not invent or fabricate data.
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Good output would be defined as a response that is accurate, relevant, complete and approved based on the source data (book inventory, availability, library information). Responses should answer the user's question and avoid unsupported or hallucinated information. The AI assistant must maintain a friendly and concise tone. Success will be measured based on the following benchmarks: Accuracy: >=95% Relevancy: >=95% Completeness: >=90% Hallucination Rate: <5% Source Attribution Accuracy: Average Response Time: <3 seconds User Satisfaction: >= 4.5/5
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?See embedded excel sheet of sample eval cases. https://docs.google.com/spreadsheets/d/1zEUTlEmnKf0mYeQ1ct47PBsAQhtRGIAM-tb3-lILIrs/edit?usp=sharing
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Used Open AI model GPT5.4 but then moved to GPT 5.5 as previous version was no longer supported. LLM is mainly used for generating book recommendations and tool calling to retrive book rating. LLM is used for small agentic tasks such as adding books to wishlist or place hold on books. Also used Open AI's webspeech API for STT and Open AI gpt-40-mini-tts for TTS. Capabilities: 1. Able to support basic STT and TTS functions. Limitations: 1. Latency when providing recommendations.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Age, Primary Library, Zipcode to find other libraries within the geographic proximity - Identified from user's profile Genre - User to provide and if not provided, AI assistant to ask Previously Read Books - Identified from user's reading history
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?Books that the user may have already read in the past - user input Author information - user input
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Tone Accuracy Relevance Hallucination Level Response Time
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?Ensure books are being recommended for the reader's age Ensure books are recommended for the genre requested Ensure books are recommended that are part of the library catalog Ensure books are accurately indicated as available, checked out based on library inventory
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.a. You are a librarian with expert knowledge of all the books and resources available in all the libraries for the user's county. b. You must talk in first person as if you are having a conversation with the user. c. Your tone must be based on the user's age. (Children: 0–12 years, Teens/Adolescents: 13–19 years, Adults: 18–64 years, Seniors: 65+ years) d. You will recommend books to users only based on the book inventory of all the libraries in the county. e. You will recommend books based on the user's age and genre. f. You must always show 4 recommendations at a time. g. Keep your conversational response brief. Do not restate full book details that are already shown in the UI cards. Use the cards for detailed metadata; use the assistant text for a short personalized explanation. h. When recommending books, you will prioritize based on following sequence: 1. User's age 2. Genre 3. Availability - books that are available 4. Geographic proximity - libraries closest to the user's address i. You must limit the libraries within a 10 mile radius of the user's address. j. You must go to the web only to get the book rating. k. You must provide the source for the book rating. l. Do not recommend books from the reader’s known reading history. m. Do not recommend any book the user says they have already read. n. Do not recommend books outside of the genre without asking the user. o. Do not recommend books that are not part of the county's library inventory. p. Do not invent or fabricate data. q. You must use the GET /api/books/{book_id}/reviews tool for every book recommendation. Call the tool to get average rating and ratings count. Do not answer without using the tool."
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?1. Updated SP to ask the LLM to keep the the conversation brief and give a short personalized explanation. It was also restating full book details besides what was being displayed on the UI within book cards so updated prompt to not list out all the book details. 2. Updated SP to to not recommend books if the user were to indicate they have already read the book in their user prompt and any books that were part of the reading history. 3. The AI assistant was filtering books above the user's age but did not have a lower age-fit floor so updated the system propmt to fix that.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Defined profile information (users/families) and library information in SQLite Created a vector db to create a data set of 50-75 books for semantic book retrieval Defined FastAPI to get book rating and rating count from Google Books API Data Preparation - 1. Users + Libraries - Created 3 mock families and defined 3 mock libraries in the county 2. Books - a. Created a data set of books for different age groups - early readers, young adults, seniors b. Defined books belonging to various genres - Mystery, Dystopian, Fantasy, Early Readers, Gardening, Health and Wellness c. Defined metadata for each book - title, author, genre, age group, themes, book summary d. Defined book inventory (For each library - whether the book is carried, defined total count, available count and checked out count for each book) 3. For book ratings and rating count, used the Google Books API.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Inputs: a. Can you recommend books like 'The Hunger Games'? b. Can you recommend books for a 6 year old? c. Can you recommend some mystery books? Output: a. Assistant recommends age appropriate books that have a similar theme/pace as 'The Hunger games' b. Assistant recommends age appropriate books c. Assistant recommends age appropriate books for the genre requested.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)a. Asked the assistant how was the weather today. b. Asked the assistant to recommend books without providing a genre.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Testing Scope: Conducted manual stress tests across diverse queries (e.g., specific genres, obscure authors, and complex demographic constraints). Successes: The model demonstrates high accuracy in standard genre-matching and tone alignment for general reader requests. Identified Failures: Testing revealed "Hallucination" in bibliographic data (guessing titles for obscure authors) and "Persona Mismatch" where constraints like reading age were ignored. Root Cause: These failures were traced to weak grounding in the RAG retrieval step and improper hierarchy within the system prompt instructions. Resolution: Updated the logic to mandate a "Check-Before-Guessing" rule and implemented a "Persona-First" instruction structure to ensure accurate demographic filtering.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Automated Evaluation: At this prototype stage, utilized a Manual Qualitative Audit as our primary evaluation methodology. Rather than abstract scores, we established a "Golden Set" of core user scenarios—ranging from "mood-based discovery" to "subject-specific constraints"—to build a performance baseline. Performance Audit: Manually reviewed outputs for three critical criteria: (1) Factuality/Grounding (Did the book exist?), (2) Relevance (Did it match user intent/age?), and (3) Tone (Did it mirror library standards?). Baseline Findings: We identified an 80% baseline success rate. The primary failure mode observed was not showing books based on reader's age (lower age limit), suggesting same books again when asked for more recommendations. Future Automated Roadmap: To transition from prototype to production, the roadmap includes: Assertion Testing: Automated scripts to verify that all AI-recommended titles are cross-referenced against the actual Library catalog and inventory. LLM-as-a-Judge: Implementing a secondary model to programmatically grade outputs against the Golden Set to allow for continuous integration and regression testing.
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?1. The "Hallucination" (False Positive) The Scenario: A user asks for a very obscure, fictional, or hyper-niche book that doesn't exist in the library catalog. The Problem: The LLM tried to be "helpful" by confidently fabricating a book title, author, and description that sounds plausible. The Mitigation: Defined a constraint: If the retrieved context from the library catalog does not contain a high-confidence match, the AI must admit it cannot find the book rather than inventing one. 2. "Context Drift" (Retrieval Mismatch) The Scenario: The assistant returns multiple results for a broad search (e.g., "books about time travel"), but none of them are actually highly rated or relevant to the specific user’s previous preferences. The Problem: The model picked the first (less relevant) result from the available books because it’s "stuck" on the top-ranking item. The Mitigation: Changes to handle irrelevant results by implementing a "relevance filter" to ignore low-quality matches. 3. The "Semantic Dead End" The Scenario: A user enters a query that is too restrictive (e.g., "Recommend a horror book written by a female author, set in the 1800s, under 200 pages, with a blue cover"). The Problem: The search returned zero results, leaving the user with a broken experience. The Mitigation: Changes to handle "empty states" by suggesting broadening the criteria, or by offering the "next closest" match. 4. "Prompt Injection" / Out-of-Scope Queries The Scenario: A user tries to get the AI to do something other than recommend books (e.g., "Write me a Python script" or "What is the capital of France?"). The Problem: The model deviated from its "Librarian" persona. The Mitigation: Tested for "persona drift" and logical changes to redirect the conversation back to books whenever a user goes off-topic.
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?1. Iterative Refinement: Transitioned from a basic "helpful assistant" prompt to a structured, constraint-heavy system persona to reduce ambiguity and hallucination. 2. Guardrail Implementation: I integrated negative constraints—explicitly listing topics and behaviors the model must refuse—to prevent off-topic discourse and improve safety. 3. Few-Shot Calibration: I introduced "Few-Shot" examples into the system context, providing the model with concrete "Ideal Interaction" patterns to guide its reasoning during complex queries.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?Goal: Scalable, consistent, and objective assessment. • Primary Method (LLM-as-a-Judge): Since manual human review doesn't scale, use a more powerful "Judge" model (e.g., a high-reasoning LLM) to grade the output of "Doo-ee" against a predefined rubric. • Grading Rubrics: The Judge evaluates against four dimensions: 1. Faithfulness: Is the recommendation supported by the source document? 2. Answer Relevance: Does it directly address the user's specific book request? 3. Safety: Does it adhere to the content moderation guidelines? 4. Diversity: Does the recommendation avoid "popularity bias" by offering a mix of genres? • Scalability: To test diverse scenarios, maintain a "Golden Dataset"-a collection of 50+ diverse user personas (e.g., "History buff," "Sci-fi novice," "Seeking diverse voices") and expected, high-quality outcomes. Use a test runner to execute these cases in parallel whenever I change the prompt or underlying logic.
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?Goal: Catch regressions before users do. • Pre-Deployment (Gatekeeping): Evaluations are automatically triggered on every GitHub Pull Request. No code is merged unless the test suite maintains or improves the "Golden Dataset" score (Regression Testing). • Post-Launch Monitoring (Continuous): I implement Continuous Evaluation (CE) on production traffic. By sampling 5–10% of real-world queries, the system automatically scores them using the LLM-as-a-Judge. This allows us to spot "drift"—cases where the model begins to perform poorly on types of queries it previously handled well. • Periodic Human Audits: Once a month, perform a "Human-in-the-Loop" (HITL) review. I manually inspect the transcripts of the lowest-scoring AI recommendations to calibrate my Judge model, ensuring the automated scoring remains aligned with real-world user satisfaction.
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?• Infrastructure Documentation: All core components—APIs (OpenAI), Vector Databases, and hosting environments (Vercel/Render)—are documented in the project README.md. This includes environment variables and dependency manifests. • Rate Limiting & Cost Controls: Implement a "soft limit" on API token usage to prevent runaway costs, along with simple client-side rate limiting to protect the backend from abuse. • Error Handling & Rollback: o Graceful Degradation: If the primary AI API fails, the application defaults to a "Service Temporarily Unavailable" message rather than breaking the UI. o Version Control: By utilizing GitHub, my code is versioned. If a new deployment introduces a regression, I can perform a "one-click" rollback to the last stable commit via my deployment platform (e.g., Vercel/Netlify). • Observability: Utilize basic logging to capture failed requests and latency issues, to identify exactly where a breakdown occurs.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Goal: Prove product is usable and "supportable." • Internal Knowledge Base: A "Living Doc" (e.g., Notion or GitHub Wiki) exists for the project. It serves as a "Single Source of Truth" containing: o The "One-Pager": High-level vision and business value. o Technical Runbook: Steps for troubleshooting the most common issues (e.g., "What to do if the API key expires"). o Support Playbook: A set of FAQs for how to handle potential user inquiries or "unhappy path" reports (e.g., "What to tell a user if the AI provides a bad recommendation"). • Training & Handoff: o Standardizing Inputs/Outputs: Clearly define what data enters the system and what the AI is permitted to output. o Stakeholder Alignment: Ensuring the "evaluators" (your professors/mentors) have access to the user-facing documentation and the "Support Playbook," demonstrating that the product is ready for external interaction.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?Target a limited Pilot with a cohort of [target user group, e.g., students or local library patrons] in a particular county library to gather qualitative feedback and refine the recommendation logic before an open public release.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?To scale, we would implement request-caching for common queries and move from a single-tenant to a multi-tenant database architecture to support increasing concurrent user volume.
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?Goal: Bridge the gap between the technology and the user's need for value. To ensure users understand how to use an AI assistant effectively, prepare a "Support & Adoption Kit": Interactive Demo/Tour: A short, guided "walkthrough" that highlights the "Aha!" moment—for example, how a user goes from a vague request ("I want something exciting to read") to a curated, perfect book recommendation. "How-to" Guides: Quick Start Guide: A one-page PDF or landing page section covering: How to phrase your request, How to provide feedback to the AI, and What to do if the recommendation misses the mark. Since the tool relies on user input, provide a "Cheat Sheet" of effective prompt templates to help users get the most accurate results. Living FAQ: A searchable database addressing common user hesitations, such as: "How does the AI choose my recommendations?" (Transparency regarding the logic). "Is my book preference data private?" (Addressing trust). "What if the AI suggests a book I've already read?" (Troubleshooting). Feature Explainer/Sell Sheet: A high-level document explaining the "Why now?" of the project, focusing on the problem (the overwhelming nature of book choices) and the unique AI solution.
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Goal: Maintain transparency and build confidence in the project’s progress. For internal communications, focus will be on rhythm and evidence. Launch Roadmap (The "What & When"): A clear visual timeline (e.g., a simple Gantt chart or milestone list) shared via email or a shared workspace, showing current status vs. upcoming milestones (e.g., Prototype, Beta Testing, Final Polish). Bi-Weekly "Show & Tell" Updates: A short 60-second video demo showing a specific feature that was just implemented to keep stakeholders engaged. The "Impact Dashboard": A simple, recurring summary of key metrics: Success Metrics: Number of recommendations provided, average time to get a recommendation, or user sentiment. Risk/Blocker Log: Be transparent about any technical hurdles (e.g., latency in the AI response) and how they will be mitigated. Post-Mortem/Feedback Loop:Share the "Lessons Learned" and "Future Iterations," (long-term lifecycle of the product)
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?All user queries are encrypted in transit via HTTPS. We do not store PII (Personally Identifiable Information). Future compliance will be handled by implementing data retention policies and user-controlled data deletion workflows.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?1. Content Moderation & Safety Safety Guardrails: Implemented System Prompt Guardrails that strictly constrain the AI to literary and reading-related topics, refusing to engage in sensitive, political, or non-literary discourse. Output Validation: I’ve designed the system with a "sanity check" layer (via prompt engineering or a lightweight classification model) that scans the AI's response for prohibited keywords or harmful sentiment before displaying the final recommendation to the user. Human-in-the-Loop: For the current prototype, I am the primary auditor. In a production environment, I would include a "Report a Recommendation" feature, allowing users to flag inappropriate suggestions for manual review. 2. Data Privacy & Compliance (The "Trust" Layer) Consent Management: Upon the first launch, the user is presented with a clear Privacy Disclosure stating exactly how their book preferences are used to train the local recommendation model, with a "Clear My History" button to exercise the right to be forgotten Transparency by Design: Included a disclosure label in the UI: "Recommendations are AI-generated and for entertainment/discovery purposes. While we strive for accuracy, please verify book availability and content via your local library or publisher." 3. Bias Mitigation & Fairness Algorithmic Transparency: Will incorporate a diversity factor in the recommendation prompt to ensure the model surfaces niche, long-tail, and diverse literary voices alongside mainstream titles. Regular Audits: Conduct periodic "Bias Audits" by testing the assistant with diverse user personas to ensure the recommendations aren't clustering around a single genre or demographic viewpoint.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User Engagement (The "Helpfulness" Indicator): Recommendation Acceptance Rate: The percentage of AI-generated suggestions that users click on, add to cart or add to a reading list. Session Conversion Rate: The percentage of user sessions that result in a successful book discovery (i.e., the user doesn't bounce immediately). Average Sessions per User: Measures "stickiness"— are users returning to Doo-ee for their next book? Business Metrics (The "Value" Indicator):User Retention (Day-30): The percentage of users who return to the platform after their first experience. Task Success Time: The average time taken from the initial "book discovery" prompt to the user saving a recommendation. Feedback/NPS Score: A simple post-recommendation "thumbs up/down" to capture direct user sentiment.
AI MetricsHow will you measure AI performance and accuracy?These metrics measure the technical health of the recommendations. For an LLM-based assistant, quality will be measured rather than just raw speed. • Recommendation Quality: o Click-Through Rate (CTR) on Recommendations: How often the top-ranked recommendation is acted upon. o Mean Reciprocal Rank (MRR): Measures where the "perfect" recommendation appears in the list. Ideally, the most relevant book should be the #1 or #2 suggestion. o Diversity Score: A check to ensure the model isn't just recommending the same 5 "bestseller" books to every user, ensuring the AI suggests a mix of genres and authors. • Model Reliability: o Hallucination/Invalid Reference Rate: Percentage of recommendations that refer to non-existent books or incorrect authors (crucial for a library assistant). o Answer Relevancy: Using an LLM-as-a-judge (or automated test cases) to score how well the AI's explanation for why a book was recommended matches the user's specific preferences.
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?For the pilot, support will be handled via a dedicated feedback link within the app. Escalations will be tracked in a issue tracker. Alignment will be achieved with the support team regarding the process to handle incoming requests and escalate when needed.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?1. Gather (The Input) • In-App Feedback: Integrate a "Thumbs Up/Down" widget directly alongside every book recommendation. • Explicit Reporting: A "Report an Issue" button for users to flag specific errors (e.g., "AI hallucinated this book," "link is broken"). • Usage Telemetry: Passive collection of session data (where users drop off, what queries generate no results) to identify friction points before a user even bothers to complain. 2. Triage & Act (The Processing) • Categorization: All feedback is sorted into three buckets: o Data Quality: (e.g., wrong author, broken link). o Model Performance: (e.g., "these suggestions are repetitive"). o Feature Request/UI: (e.g., "I wish I could filter by audiobook"). • The Iteration Loop: Bi-weekly review of feedback. High-frequency or "high-severity" issues are pushed to the next sprint; "noise" (individual preferences) is logged in a backlog for future UX analysis. 3. Prioritization (The Decision Matrix) Use a Severity x Impact matrix to rank items: • P0 (Critical/Blocker): Broken core functionality (e.g., the AI is down, or showing offensive content). Action: Immediate fix, pushed to production as an emergency hotfix. • P1 (High Impact): Inaccurate/Hallucinated recommendations for popular queries. Action: Refine system prompt/training data within the next sprint. • P2 (Low Impact/Enhancement): Feature requests, UI tweaks, or "nice-to-haves." Action: Add to the product backlog for the next release cycle.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?Goal: Move from "blind" deployment to "observable" operations. Since AI systems are non-deterministic, need to track both technical performance and model quality. • Operational Logging: Every user interaction is captured as a "trace" containing the full prompt, model response, metadata (latency, token usage), and session ID. • Key Monitoring Pillars: o Latency & Throughput: Monitor P50/P99 response times. Spikes in latency often signal underlying API issues or inefficient prompt/context bloat. o Cost/Token Tracking: Real-time budget monitoring. An unexpected spike in token usage is an early indicator of prompt engineering errors or "runaway" agent loops. o AI Guardrails: Real-time filtering of outputs. Any response flagged by safety/moderation layers triggers an immediate alert. • Automated Quality Checks: Use an "LLM-as-a-Judge" approach—a separate, smaller model that periodically scores a sample of the assistant's recommendations for relevance, tone, and accuracy against the defined rubrics.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?Goal: Close the loop between real-world performance and system updates. • Feedback Loops: Use the "Thumbs Up/Down" data to tag "Gold Standard" examples. These become the Evaluation Dataset for future testing. • Drift Detection: Periodically audit if the "type" of queries is changing (e.g., are users asking about non-book topics?). This helps determine when to update the system instructions to better handle new user behaviors. • Champion/Challenger Testing: Before deploying a major prompt update or model change, run it in parallel with your live version on a small slice of traffic. Only promote the "Challenger" if it outperforms the "Champion" in accuracy and user engagement. • Periodic Human-in-the-Loop Reviews: Monthly "Office Hours" with a small group of power users or subject matter experts to review low-scoring recommendations and manually tag corrections.
Download the .xlsx ↓