AI Tools
Norma AI
Norma AI is a compliance assistant for data protection and compliance professionals who need to understand AI regulations in plain English without a large legal budget. Users set a persona, industry, jurisdiction, and AI systems, then receive a personalized dashboard, searchable regulation library, and an assistant that answers compliance questions with citations, action items, penalties, and confidence levels. The demo focused on EU AI Act use cases for a fintech DPO.
The problem
Data protection officers must understand fast-moving AI regulation without a large legal budget. A DPO at a mid-market company faces 144 articles of the EU AI Act, GDPR overlap, and other jurisdictions — thousands of hours of dense legal language — while everyday questions ("which article applies to this new HR tool?") and law changes go unanswered without an expensive law firm at $300–$1,000 per hour. The pain is high-severity and daily: no quick reference to the applicable law, no plain-English translation, no alert when regulations change, and no way to judge the reliability of any answer they do find.
The solution
Norma AI is an AI governance copilot that explains privacy and AI regulations in plain English under one umbrella. Users set a persona, industry, jurisdiction, and AI systems, then get a personalized dashboard, a searchable regulation library documented down to sub-articles, and an assistant that answers compliance questions. Every answer is grounded exclusively in the official regulatory text — the demo focused on the EU AI Act for a fintech DPO — and follows a fixed format: a plain-English answer, the cited article, the recommended action, who it applies to, the applicable penalty, and a confidence level (🟢 High / 🟡 Medium / 🔴 Low). Penalties over €1M trigger an explicit recommendation to consult a lawyer. Pricing is freemium: 10 free assistant questions, then Pro and Team tiers.
How it works
Norma runs on a RAG pipeline over the EU AI Act PDF. Text is extracted with pdfplumber, then Claude Haiku structures each chunk into article, title, summary, obligations, penalty, and tags. Content is chunked (up to 1,500 words with overlap) and embedded into Pinecone across two namespaces — one for the assistant's raw chunks, one for the structured regulation library. Claude Haiku was chosen over Sonnet and GPT-5.5 because this is structured retrieval, not complex reasoning: it follows the format every time, cites articles consistently, and costs far less. It never answers from memory — an open-book exam where the book is the retrieved articles. Dynamic date injection ({{TODAY}}, {{DEADLINE}}, {{DAYS_LEFT}}) grounds deadline answers. Automated evals across accuracy, grounding, format, edge cases, and RAG quality rose from a 50% first pass to 85% after fixing retrieval (query enhancement mapping "chatbot disclosure" to Article 50, manual chunks for hard-to-retrieve articles) and the prompt.
Who it's for
Norma is aimed at DPOs and legal counsels, with CTOs, executives, and CFOs planned for future releases; the capstone is built around the DPO persona. The archetype is Sonia, a DPO at a fintech SaaS company using AI for recruitment, chatbots, and credit scoring, with no budget for a law firm and a looming EU AI Act deadline. The listed customer base is B2C, sold as individual Pro subscriptions and Team company subscriptions. Persona selection reframes answers — a DPO gets deployer obligations, a CTO gets technical requirements, a CEO gets a penalty summary.
Why it matters
The regulatory pressure is real and time-bound: the EU AI compliance market is projected to grow from $1.2B in 2024 to $6.8B by 2028, driven by the August 2026 enforcement deadline, an estimated 50,000+ companies required to comply, and parallel regulation across the US, UK, and Canada pointing to a $12B global addressable market by 2029. No dominant self-serve player exists. Because incorrect guidance carries liability, Norma keeps humans in the loop and refuses out-of-domain questions — asking clarifying questions for vague queries and declining to answer on unloaded regulations. Launch runs through a closed pilot with 10 friendly DPOs for accuracy validation, then public beta, then full self-serve. Critical issues — wrong penalty figures or hallucinated obligations — are defined as fix-within-24-hours.
The workflow
The PRD
| PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0 | ||||||
|---|---|---|---|---|---|---|
| Your Name: | Minal Phavde | |||||
| Your Product: | Norma - Privacy and Governace Laws- Explained in Plain English and Current | |||||
| Your Industry: | Legal | |||||
| Date: | 8th May 2026 | |||||
| 4D Method | AI PRD | Instructor Feedback | ||||
| Phase | Activity | Theme | Topic | Key Question(s) | Your Response Include external links to visuals/prototypes as required. | |
| DISCOVERY | Understand your market, business, product & user context | Business Value Map | Market Attractiveness | What industry is your business in? (ie Financial services, Healthcare, Education, etc)? | Legal - Ai Governance, other privacy laws explained in simple English. | Minal, the problem framing is strong. The current-state journey grounds the pain in a scenario any compliance professional recognizes, and the Workflow architecture shows real system-level thinking from onboarding through RAG through guardrails and audit logging. The competitive landscape will hurt most on Demo Day. Claiming nothing in the market serves this purpose invites pushback from any judge who knows OneTrust or Vanta. Name three to five competitors, state where each wins, and show the wedge they leave open for the mid-market professional. A tension runs through the business model. The PRD names individual consumers and a freemium tier, but the prototype is built for enterprise with SSO, RBAC, and team governance. Two products, two sales motions. Decide which launches first. The architecture earns credibility, but the operational AI layer is empty. Master prompt, evaluation criteria, and test cases are all blank. In legal, where bad guidance carries liability, these are load-bearing. Write the system prompt defining role, citations, and refusal behavior. Define three to five evaluation criteria with pass/fail thresholds. Prompt, then evals, then build. One forward question: the headwinds section names liability risk but does not describe mitigation. Guardrails, disclaimers, confidence indicators, and human review triggers belong in Design before development. |
| What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors? | Tailwinds: 1Ai Governance pressure is real Aug'2026 is the EU Ai act deadline. 2. Regulations are accelerating globally so many acts, a tool under one umbrella showing all regulations in simple english would help. 3. Data Privacy Office are overwhelmed , mid market companies have to pay high fees to Law firms. 4. No dominant self serve player exists. Headwinds: 1. Regulatory text keeps on changing - Maintenance is expensive.Continuous updates ongoing cost after launch 2. Liability risk : if incorrect guidance given, consequence could be significant, needs humal in the loop for reviews. | |||||
| What is the projected growth rate of your target market segment over the next 3-5 years? | "The EU AI compliance and governance market is projected to grow from $1.2B in 2024 to $6.8B by 2028 — driven by the EU AI Act August 2026 enforcement deadline, expanding scope to 27 member states, estimated 50,000+ companies required to comply, and parallel regulation emerging in the US, UK and Canada creating a global compliance addressable market of $12B by 2029." | |||||
| Business Model | What growth stage is your business currently in (e.g., startup, scale-up, mature)? | Start-up , Internal user base pilot. | ||||
| How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?) | Norma - Ai Governance Co-pilot A DPO who would spend hours drowning into Learning and understanding legal language for ever changing laws, or pay law firms $300-100$ per hour can now find all laws explained in plain simple english under one umbrella with a personal Ai Assistant to ask questions the tool will sell:: 1. Regulatory Intelligence- Eu Ai ACT, GDPR, other Global Laws explained in plain English. 2. Filters- PErsona based applicable law, Technology/domain based, Country based law filters. 3. Ai Assistant. 1. Free - 10 Ai assistant questions on full access to regulations library, no document export. 2. Pro (Individual Subscription) - DPOs 3. Team (Company subscription) | |||||
| Who is your primary customer base (B2B, B2C, B2B2C)? | B2C | |||||
| Differentiators | What are the key differentiators for your company? | Ai assistant- answers grounded exclusively in the official EU AI Act PDF with zero hallucination risk expandable to other regulations too, No Legal language., Rag Retrieval, formatted answers with confidence level and roadmap for action items. Articles and Sub articles Library briefly documented unlike any other tool in the market. Persona based dashboard, Change feed alerts based on personas Live regulations updates.Norma is built around , readiness score and personalised obligation checklist that turns regulatory anxiety into a clear action plan. | ||||
| Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.) | Customers | Who are the customers (ie buyers) of your product? | ||||
| End Users | Who are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context? | |||||
| Current Products / Services | If you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers? | |||||
| User Value Map | Target Persona | Who is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc) | DPOs, Legal Counsels. CTOS, Executives, CFOs for future releases. | |||
| Journey Map (current-state) | What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product? | Sonia a DPO with Fintech SAAS company which uses Ai tools for Recruitment, customer service Ai chatbot, Credit score calculator for customer, no budget for a law firm, her company needs to be ready for Eu Ai act. Sonia uses google finds 144 articles of EU AI Act, plus GDPR overlap, plus other jurisdiction. Tries to find the applicable article from EU AI act what needs to be done. Invests in a law firm which is very expensive. The most difficult part there are changes in LAw which Sonia is not aware of so no lice change feed. | ||||
| Pain-points | Where does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe? | 1. Privacy Laws- Difficult Legal Language and long documentation 1000 of hours. 2. No Quick reference and relate to the applicable law - dependency on Law firms $$$$. 3. Everyday questions like what are the new changes today, We are buying a new HR tools which has xyz capabilities which law DPIA, FRIA, which article of EU Ai act is applicable. 4. Alerts when Law changes 5. Risk scoring | ||||
| AI Opportunities | From your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first. | 1. Severity: High, Frequency: Daily: Cannot get quick Answers to regulatory questions Tight deadline- 2. Severity: HIGH, Frequency:Daily.Understanding Regulation legal language is very hard and documentation is too much. 2. Severity High: Frequency: Weekly: I miss the changes in regulations that apply to my domain. 3. Severity - Medium: Quarterly: Audit Trial - Audit report to monitor compliance 4. Severity: High: Which Ai tools fall under high-risk classify and apply a risk score | ||||
| Develop an AI Solution Hypothesis | AI Solution Hypothesis | Diverge | Ideate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage. | 1.A Regulation Library based on Regulations explained in simple Plain English. Covers various Ai Privacy laws like EU Ai act, US laws for different states, NIST etc. 2. Ai Assistant : Quick answers based on identified documentation. 3. Persona definition storage like DPO/CTO/CFO, Industry: Technology, CRM, Fintech, Ai tools Used. 4. Change Feed: A personalised feed that monitors authoritative regulatory sources continuously and delivers only changes based persona selected. 5. Ai Audit Trial: Showing every conversation , questions asked to depict regulations and investigated and followed. 5.Connection to Ai Inventory tools per processes and generate Risk score. | ||
| Converge | Rank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project. | Refer to Impact and Impact/Feasibility worksheet | ||||
| DESIGN | Define Target State Workflow | UX Flows & Wireframes Suggested Tool: Excalidraw | Workflow (future) | Assuming your product or feature works as desired, what is the target state workflow? | Please refer to the worksheet "Workflow" | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Build Wireframes | Wireframes | How will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features? | Please see the wireframes on this link https://normacomply.lovable.app | |||
| Develop Prototype to showcase AI interactions | Prototype Screens Suggested Tool: lovable.dev | What aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases? | Depicted in "Impact and Feasibility" worksheet. | |||
| Initial Prompt Design | Master Prompt [Initial Design] | Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency? | ou are Norma, an AI regulatory compliance assistant for Data Protection Officers and legal teams at mid-market companies. Today's date is {{TODAY}}. The EU AI Act high-risk compliance deadline is {{DEADLINE}} — {{DAYS_LEFT}} days away. Your job is to explain AI regulations in plain English to anyone without needing a lawyer for every question. Answer ONLY using retrieved regulatory content. STRICT RULES: - Always cite the specific article number e.g. Art. 27 - If the answer is not covered in the provided text say clearly: "This is not covered in the articles — please consult legal counsel" - Keep answers concise and practical — no jargon - Cpnvert Legal terms into simple english in laymans language - Always end with a confidence level: Confidence: High / Medium / Low - Add emojis based on confidence levels: 🟢 High — answer clearly supported by retrieved text 🟡 Medium — answer partially supported, some interpretation needed 🔴 Low — answer not well covered, legal advice recommended - If penalty exceeds €1M add: "⚠️ Penalty exceeds €1M — we recommend consulting a qualified lawyer" - If asked about deadlines or implementation timing, first ask: "Would you like me to generate a compliance roadmap based on today's date and the {{DEADLINE}} deadline?" Only generate the roadmap if they confirm yes. ANSWER FORMAT: ANSWER: [plain English explanation] **ARTICLE:** [article number and title] **ACTION:** [what the DPO should do next] **Whom it Applies To:** [Provider / Deployer / All] **PENALTIES Applicable:** [exact penalty from the regulation] CONFIDENCE: [🟢 High / 🟡 Medium / 🔴 Low] | |||
| Prepare for Testing & Iteration | Evaluation Criteria & Test Plan | Evaluation Criteria | What specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output? | Ran 5 categories of evals against Norma's AI assistant — Accuracy, Grounding, Format, Edge cases, and RAG quality. Our accuracy eval suite achieved an 80% pass rate on first prompt version,Can be improved more by fixing the Personas based answers and the penalties data uploading.However overall the Basic Accurate evals and grounding evals are passed. | ||
| Example Cases | What specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs? | Please see Evals worksheet | ||||
| DEVELOP | AI Model Selection & Justification | AI Model Selection & Justification | Which AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product? | Claude Haiku was chosen over Sonnet and GPT-5.5 because this task is structured retrieval — not complex reasoning. Haiku follows the format instructions every time, cites articles consistently, and costs 5x less than Sonnet with no quality difference for Norma's use case. Claude Haiku never answers from memory. It only reads what Pinecone retrieved and writes a plain English answer from that text. Think of it like an open-book exam — the book is the EU AI Act, and Claude can only write from what's on the pages in front of it | Please leave this area blank. This space is for the Instructor to provide you with feedback. | |
| Define Inputs | Input Specification Table | Required Fields | What are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement. | 1. Input Fields :question- 800 chars max-Source:Chat Box, Mandatory 2.Input Fields :Persona-String Enum:DPO,CTO..Source:User Persona settings, Not Mandatory Default set to DPO 3. Input Field: Today auto generated at runtime to calculate the days left for Eu Ai act. 4.deadline: hardcoded in .py file at the backend. 5. model - String · claude-haiku-4-5-20251001, hardcoded in the .py file at the backend | ||
| Optional Fields | Are there any optional or user-customizable fields? How do they impact the AI’s output? | 1. Persona set in the Persona settings. Changes answer framing — DPO gets deployer obligations, CTO gets technical requirements, CEO gets penalty summary. Currently the capstone project is built around DPO. 2. Question field in Ai assistant: Framing of the questions changed the answer retrieval. | ||||
| Define Good Output | Output Evaluation Checklist | Objective Criteria | What criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance) | 1.The structure should have a general Humas vibe/tone and as per format defined in the Prompt 2.The answers should show correct data like the penalties for different articles the correct articles cited 3.Currently the pinecone data has only Eu Ai act uploaded any other questions should not be answered, this confirms no hallucination.1. | ||
| Subjective Criteria | Are there any criteria that require human judgment or qualitative assessment? | Yes as the received text is converted into simple English from Legal language checking if the retrieved info is correct needs to be confirmed. | ||||
| Prompt Design Iteration | Master Prompt [Final Design] | Prompt Version 1 | What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints. | 1. the starting system prompt was very basic with direct questions like what does Article 50 say? 2. modified to add conversion to plain english, formatting adding confidence level, listing the deadline. 3. Added dynamic date injection — {{TODAY}} {{DEADLINE}} {{DAYS_LEFT}} 3. Later even added a question to give a roadmap and if confidence level was low reference to the Legal Counsel. 4. Also added instructions to ask further questions if question asked was not clear. | ||
| Prompt Iterations | If revised, what changes did you make and why? How do you track and record prompt evolution? | Stored prompt versions in Claude workbench. | ||||
| Data Preparation & RAG Implementation | Data Preparation & RAG Implementation | What data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information? | 1. Source: Eu Ai act PDF, document named by articles. 2.Data Prep: Extraction done by pdfplumber and Alude Haiku extracted articles , article · title · summary · obligations · penalty · tags per chunk 3. Rag pipeline:Chunking- chunk_size=500 words · overlap=80 words · ~216 chunks, embeddings. 4. Pinecone two namespaces, ---Ai assitant: default namespace stored 216 raw pdf chinks: OpenAI text-embedding ---Regulations Library: 87 structured article records- Hash-based vectors, served the purpose. | |||
| Create Evaluation Set | Example Input/Output Data for Testing | Typical Examples | What are the most common inputs and expected outputs? Use real data if possible. | Detailed in evals sheet. | ||
| Edge Cases & Negative Cases | What examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain) | Ambiguous input — asking "am I compliant?" with no context caused Claude to ask clarifying questions rather than answer — correct behaviour for a compliance tool. Out-of-domain — asking about CCPA or US state laws returned "not covered in retrieved articles" — Claude correctly refused to answer from general knowledge. Wrong article reference — asking "what is Art. 999?" returned "not found in provided text" with a lawyer recommendation — grounding rules held correctly. Missing Data: the penalties could not be retrieved right also Annexure 3, Fixed by loading dedicated manual chunks. | ||||
| Test Example Data & Review Results | Manual Review | Run your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why? | Detailed in the Eval sheet | |||
| Automated Evaluation | What pass/fail rate or scores did the AI achieve on core criteria? | The pass rate achieved with first test was only 50% which after fixing the data retrieval mechanism and prompt helped increase the test score to 85%. Still few failures if other personas are used. | ||||
| Handle Edge Cases & Iterate | Edge Case Identification | What edge cases did you identify in testing or real usage? | Asked questions like: 1. Am I compliant, it asked more questions for clarification 2. what can i Do: Asked clarification questions 2. Long Questions: Structured answer | |||
| Updates & Adjustments | What prompt or system adjustments have you made based on failures, feedback, or edge case observations? | Prompt improvements: 1. Answers too long and legal- Added plain English rule — explain jargon immediately after using it. 2. No deadline context in answers- Added dynamic date injection — {{TODAY}} {{DEADLINE}} {{DAYS_LEFT}} 3. Claude generating unsolicited roadmaps - more tokens used- Added conditional — ask user first before generating roadmap 4. No way to judge answer reliability- Added confidence indicator — 🟢 High / 🟡 Medium / 🔴 Low RAG retrievals: 1. Regulation Library: data was truncated also sub articles were not shown-Updated articles.py extraction prompt to detect when a chunk contains only part of a long article and create sub-article records — e.g. Art. 5(1)(a-c) and Art. 5(1)(d-f) as separate cards 2.Data truncation- Increased chunk_size from 800 to 1500 words — longer articles like Art. 5 (1,788 words) were being cut mid-article causing incomplete extraction 3. Ai assistant: Art. 14 and Art. 27 not retrieved- Loaded dedicated manual chunks with keywords prepended for better matching 4. Ai assistant: Art. 50 not retrieved by Pinecone - Added enhance_query() — maps "chatbot disclosure" to "Article 50 transparency" 5. Ai assistant: Hash vectors returning wrong chunks- Hash vectors returning wrong chunks | ||||
| Automate Evaluation Approach | Evaluation Method | What is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets? | Hybrid approach is better,Manual review for edge cases. Model grader for complex retrieval comparing with other models using the model graders.However if 100 evals need to be tested with every change across the same data script way is better also in case for fixed output checks like citations/penalty comparisons. | |||
| Evaluation Frequency | How often will you re-run evaluations for new data, new prompts, or post-launch monitoring? | With the data changes i have tested new data was not added then too i had to test everytime if the model is not hallucinating. definitely everytime testing needs to be performed , with addition of new data or changes with the existing data. | ||||
| DEPLOY | Finalize Launch & Rollout Plan | Operational Readiness Checklist | Technical Readiness | Is infra (APIs, databases, rate limits, monitoring, rollback) tested and documented? | All three API endpoints were manually tested pre and post deployment · Pinecone namespaces verified via stats · token usage tracked per response · rollback handled via git revert and Railway auto-redeploy · all scripts and architecture documented in GitHub." | Please leave this area blank. This space is for the Instructor to provide you with feedback. |
| Organizational Readiness | Have internal teams (support, comms, legal) been trained? Is documentation complete? | yes communicated through newsletter. | ||||
| Launch & Rollout Strategy | Launch Approach | What is your launch approach? Pilot, AB test, or all users—who gets access and when? | My approach would be "Norma launches via closed pilot with 10 friendly DPOs for accuracy validation · opens to public beta in month 6 with weekly invite batches · goes fully self-serve in month 9. | |||
| Scale Readiness | How will you ensure readiness for scale? How will you monitor initial volume and scale up? | monitor the token usage and adding the Subscriptions for Pinecone , Railway, Anthropic tiers accordingly. Not very sure. | ||||
| Go-to-Market Plan | Marketing / Training Assets | What assets (FAQ, demo, guides) will you prepare for external communication/marketing? | I will prepare short demos which are more impactful with FAQs for end users. Technical documentation with api details etc. | |||
| Stakeholder / Internal Comms | How will you communicate launch plans, progress, and outcomes internally? | "Launch progress communicated via GitHub commits for technical changes · Internal chats for daily updates and incidents · -weekly pilot feedback summaries for stakeholders - newsletters or Formal announcements for upcoming launch dates. · Scheduling demos for DPOs , legal team to build engagement. | ||||
| Confirm Legal, Privacy & Risk Protocols | Data & Privacy | How do you handle and protect user data, including storage, privacy, and compliance? | "Norma stores no user questions or PII — all inputs are processed in memory and discarded · persona settings stored in browser local storage only · Pinecone holds regulation text exclusively · production will add authentication · EU data residency and signed DPAs with all third party processors. User data is collected minimally · processed in memory where possible · stored encrypted with access controls · never used to train third party models · governed by signed DPAs with all vendors · and users retain the right to view export and delete their data at any time in compliance with applicable privacy law." | |||
| Policy & Compliance | Are content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain? | Audit trail tracked. Architecture, PII data usage and Ai usage details shared with the architecture board and LEgal team for review. | ||||
| Define Success Metrics | Success Metrics | User/Business Metrics | What user metrics will indicate success? What business metrics will demonstrate value? | "User success is measured by answer accuracy above 85% · session depth above 3 questions · and return rate above 40% · cost per question, token usage. | ||
| AI Metrics | How will you measure AI performance and accuracy? | Ae performance can be measured by setting benchmark against automated scripts pass fails. Metrics tracked after each eval run.Every change in the data, automated evals script to be run,Run test_pinecone.py · verify top 5 retrieval queries return correct chunks. User feedback - Collect user thumbs up/down · calculate human accuracy rate | ||||
| Monitor, Iterate & Improve | User Support & Feedback Plan | Support Channels | Where can users get support? Is escalation and ownership clear? | Added a contact us page with support email on the app. Can also ask the question on Ai assistant and configure the answer accordingly. | ||
| Feedback Workflow | How do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated? | User feedback collected via pilot DPO interviews and thumbs up/down ratings · bugs triaged by severity in GitHub Issues · critical issues defined as wrong penalty figures or hallucinated obligations · fixed within 24 hours · communicated via weekly release emails, teams/slack channels and tracked through Git commit referencing the issue number. | ||||
| Monitoring & Continuous Improvement | Monitoring Approach | What monitoring/logging is in place to spot operational/AI issues post-launch? | AI issues flagged when confidence score distribution drops below 70% green or error rate exceeds 5% · reviewed daily during pilot and weekly post-launch. | |||
| Ongoing Improvement | How will you collect learnings, review performance, and update your system continuously post-launch? | AI eprformance can be measured by setting benchmark against automated scripts pass fails. Metrcs tracked after each eval run.Every change in the data, automated evals script to be run,Run test_pinecone.py · verify top 5 retrieval queries return correct chunks. User feedback - Collect user thumbs up/down · calculate human accuracy rate | ||||




