← All capstone projects

AI Tools

ProcuLens

Built by Mifuyu Fujii Cohort 9 Manufacturing procurement / supply chain

ProcuLens is an LLM-powered tool for manufacturing procurement teams making indirect-material supplier decisions after collecting quotes. It extracts and normalizes quote data, combines that with quantitative and qualitative purchase history plus market price references, and recommends a supplier based on evidence beyond lowest price. The workflow also lets users review extracted data, inspect supporting evidence, and see suggested next actions.

The problem

Manufacturing procurement teams spend significant time comparing supplier quotes for indirect materials. After collecting quotes, staff manually enter each one into Excel and pick a supplier on price, delivery, and past relationship. Market-price research is done when possible, but most people skip it. Indirect materials are often deprioritized versus direct materials, so each site buys as needed and cost optimization suffers. For teams managing multiple factories and 1,000-5,000 SKUs, comparison work is a heavy, know-how-dependent burden — and negotiation, which should follow from history and market prices, usually doesn't happen at all.

The solution

ProcuLens is an LLM-powered supplier recommendation assistant that goes beyond lowest price. It extracts and normalizes quote data, combines it with quantitative and qualitative purchase history plus market-price references, and recommends a supplier grounded in evidence. The explicit objective is not simply the cheapest supplier. It separates quote information, transaction history, qualitative feedback, and market prices, then explains in practical terms why a supplier is recommended. Users can review and edit extracted data, inspect the supporting evidence, and see suggested next actions such as price negotiation or delivery confirmation — treating market prices as reference for validation, never the primary basis.

How it works

The user uploads quote files (PDF, Excel, image) and enters the target material. The system extracts quote data via OCR and LLM parsing, normalizes units, dates, currencies, and model numbers, and validates that fields are complete and totals match. After the user reviews the results, LLM function calling retrieves context three ways: SQL search on Supabase for quantitative history, semantic search (RAG) for qualitative notes like quality issues and delivery delays, and a Tavily web search for market prices. Retrieval uses OpenAI text-embedding-3-small with pgvector. The model runs at low temperature and returns strict JSON, keeping each source clearly separated and writing "Not retrieved" or "No relevant information found" when data is missing. Across 10 evaluated runs it scored 10/10 on recommendation reasoning and next-action clarity, 9/10 on missing-data handling, 8/10 on non-hallucination, and 7/10 on source citation; Claude Sonnet 4.6 and Haiku 4.5 both produced accurate, easy-to-read decisions.

Who it's for

The users are procurement staff at manufacturing companies who collect indirect-material quotes from multiple suppliers — especially people managing many factories and SKUs who spend significant time on comparison work. Customers are B2B small and mid-sized manufacturers handling roughly 1,000-5,000 SKUs of indirect materials and MRO supplies. The delivering business provides FDE-style AI and DX implementation support, differentiating by helping SMEs design and implement targeted AI workflows quickly and at low cost rather than deploying a large core system.

Why it matters

Rising raw-material, logistics, and labor costs have made procurement cost optimization increasingly important, while finance teams are asked to do more with the same headcount. The procurement software market is forecast to grow around 10-12.5% annually, and procurement analytics in Japan at a 25.7% CAGR through 2030. ProcuLens targets operational efficiency, less dependence on individual know-how, and standardized processes across sites — the gap left by tools like Coupa and SAP Ariba. It launches as a limited pilot with a small team of procurement staff to validate the workflow from upload through extraction, comparison, recommendation, and email drafting, with success measured by time saved per comparison and estimated cost reduction.

The workflow

The PRD

PRODUCT FACULTY — AI PRODUCT REQUIREMENTS DOCUMENT (PRD) TEMPLATE Version 1.0
Your Name:Mifuyu Fujii
Your Product:Indirect Procurement Optimization AI Agent for Manufacturing
Your Industry:Manufacturing industry
Date:April 29, 2026
4D MethodAI PRDInstructor Feedback
PhaseActivityThemeTopicKey Question(s)Your Response Include external links to visuals/prototypes as required.
DISCOVERYUnderstand your market, business, product & user contextBusiness Value MapMarket AttractivenessWhat industry is your business in? (ie Financial services, Healthcare, Education, etc)?ManufacturingMifuyu, the journey map is carrying this section. Nine steps grounded in real procurement behavior, naming specific platforms like MonotaRO and ASKUL, showing exactly where time disappears. That level of operational specificity is rare and it gives you a genuine foundation to build on. Three places to sharpen before moving forward. First, the ICP. "Mid-sized to large manufacturing companies with multiple factories" covers an enormous range. A 500-person automotive parts maker with 2,000 indirect SKUs and a 5,000-person diversified manufacturer are different buyers with different procurement team structures, different ERP maturity, and different willingness to pay for consulting-led AI. Define the revenue band, headcount range, and approximate number of indirect material categories. That is what makes the rest of your PRD decisions (prompt design, pricing, rollout) coherent. Second, the competitive landscape names only Big4 consulting firms. Procurement software tools like Coupa, SAP Ariba, and Japan-market platforms such as infomart are closer competitors to your proposed workflow than McKinsey is. Your differentiation claim ("hands-on, frontline consulting") needs to hold up against software that automates price benchmarking without consultants. Spell out why your approach wins against those tools specifically. Third, your AI Opportunities section labels pain points with GenAI patterns (Scaling, Translation, Consistency), but not all of them require an LLM. Comparing a quoted price against past purchase prices and flagging a gap is a database query with basic logic. Where exactly does the LLM become necessary versus merely convenient? Answering that honestly will also resolve your scope question about chaining three solutions. If one of them does not need generative AI at all, it should not be in this project. What does Design look like once you have picked the tightest possible ICP from the range you are currently sitting on?
What are the key challenges (headwinds) and opportunities (tailwinds) impacting growth in your industry? Who are the key competitors?Headwinds - Rising raw material, logistics, and labor costs have made procurement cost optimization increasingly important across manufacturing. - In small and mid-sized manufacturing companies in particular, frontline teams are not fully keeping up with the work required to reduce costs. - Direct materials tend to be managed relatively closely, while indirect materials are often deprioritized. - As a result, each site tends to purchase items as needed, making procurement cost optimization difficult. Opportunities - Procurement teams increasingly need not only cost reduction, but also operational efficiency, less dependence on individual know-how, and standardized procurement processes across locations. - Generative AI could automate a connected workflow that includes quote collection, price comparison, reference to past purchasing data, and drafting negotiation emails to suppliers. Key competitors - Coupa - SAP Ariba - Other procurement and purchasing management software
What is the projected growth rate of your target market segment over the next 3-5 years?- Multiple market research reports forecast the procurement software market to grow at around 10% annually. - For example, Grand View Research forecasts a 10.0% CAGR for 2026-2033, while Strategic Market Research forecasts a 12.5% CAGR for 2024-2030. - In procurement analytics, Grand View Research forecasts a 25.7% CAGR for the Japanese market from 2023-2030, suggesting strong growth potential for AI- and analytics-enabled procurement support.
Business ModelWhat growth stage is your business currently in (e.g., startup, scale-up, mature)?The business is a growing company transitioning from the startup stage to the scale-up stage.
How does your business make money? What do they sell? What is your primary revenue model (e.g., subscription, freemium, licensing, marketplace, transactional, etc?)- The business provides FDE-style AI / DX implementation support for manufacturing companies. - It generates revenue by working closely with client companies' frontline operations and providing end-to-end support, including identifying business issues, discovering AI use cases, building prototypes, implementing solutions into operations, and continuously improving them. - The main revenue source is project-based consulting fees. The fee model is designed for each project based on deliverables and required effort.
Who is your primary customer base (B2B, B2C, B2B2C)?- B2B - Main customers are SME, manufacturing companies. - Handle roughly 1,000-5,000 SKUs for indirect materials and MRO supplies.
DifferentiatorsWhat are the key differentiators for your company?Differentiate by helping SMEs design and implement AI workflows for specific operational pain points in a short period of time and at low cost, rather than requiring the implementation of a large core system.
Feature Value Map (IMPORTANT NOTE: This section is only relevant if you are working on enhancing an existing product. It does not apply if you are developing a new product from 0 to 1.)CustomersWho are the customers (ie buyers) of your product?
End UsersWho are the end-users of your product? Which users are the most revenue-generating / revenue-impacting for your company? What are their goals, roles, and context?
Current Products / ServicesIf you are a Product-led business: What are the core features of your product, and how do they address user needs? If you are a Service-led business: What are the key services you offer to customers?
User Value MapTarget PersonaWho is your AI product / feature for? (Internal users, external users, an influencer, a buyer, etc)Procurement staff at manufacturing companies who collect quotes for indirect materials from multiple suppliers. I n particular, this targets people who manage multiple factories and many SKUs, and who spend significant time on comparison work.
Journey Map (current-state)What is the typical journey for your target persona when they are using your product / service, focusing on their ideal experience (happy path) as they interact with your product?1. Review requests from frontline teams 2. Identify and select supplier candidates 3. Prepare RFQ 4. Collect quotes and compare/analyze them 5. Negotiate 6. Obtain purchase order approval 7. Delivery and inspection 8. Invoice and payment processing 9. Review performance
Pain-pointsWhere does the user experience friction, obstacles, or unmet needs throughout the journey? Identify which pain-points are most frequent and severe?3. Preparing quote requests It is time-consuming to manually prepare requests in each supplier's format. 4. Collecting quotes and comparing/analyzing them After collecting quotes, users have to enter each supplier's quote into Excel and choose a supplier based on price, delivery date, and past relationship. Market price research is done when possible, but most people do not do it. 5. Negotiation Users need to negotiate prices with suppliers by email based on past procurement history and market prices, but this is cumbersome, so they usually do not negotiate for indirect materials. 8. Invoice and payment processing Invoice and payment processing is burdensome because indirect materials are procured in many varieties and at high frequency.
AI OpportunitiesFrom your list of pain points, identify those that can effectively be addressed using Generative AI. Remember, this project focuses on leveraging LLM-powered AI to solve your target persona's pain points. Rank these pain points starting with the most severe and frequently occurring first.3. Preparing quote requests -> No. This can be handled with simple logic. 4. Collecting quotes and comparing/analyzing them -> Yes. AI can extract information from unstructured data such as supplier emails, summarize market research, and recommend a supplier based on many variables, including price, market research results, and past relationships. 5. Negotiation -> Yes. An LLM can draft negotiation emails. 8. Invoice and payment processing -> No. A workflow can be built without an LLM.
Develop an AI Solution HypothesisAI Solution HypothesisDivergeIdeate potential solutions to address your AI-solvable pain points. Focus on generating a high quantity of ideas rather than evaluating their quality at this stage.4. Collecting quotes and comparing/analyzing them - Extract information from unstructured data such as supplier emails - Summarize market research - Recommend a supplier based on multiple variables, including price, market research results, and past relationships 5. Negotiation - Draft negotiation emails and send them to suppliers
ConvergeRank your ideated solutions based on impact and feasibility. Identify your top three AI solutions, and clearly select the one you'll focus on for your project.- Extract information from unstructured data such as supplier emails business impact: 3 feasibility: 10 - Summarize market research business impact: 6 feasibility: 8 - Recommend a supplier based on many variables, including price, market research results, and past relationships - selected theme for this project business impact: 9 feasibility: 7 - Draft negotiation emails and send them to suppliers business impact: 6 feasibility: 9
DESIGNDefine Target State WorkflowUX Flows & Wireframes Suggested Tool: ExcalidrawWorkflow (future)Assuming your product or feature works as desired, what is the target state workflow?The user selects "Start quote comparison" from the procurement dashboard. ↓ The user uploads quote files and enters the target material. PDF / Excel / image, etc. ↓ The system automatically extracts quote data. OCR (reads PDFs) LLM (understands the meaning of the extracted text and breaks it down into quote items) Normalization (standardizes variations in units, dates, currencies, company names, model numbers, etc.) Validation (checks that quantity and unit price are not empty, total amounts match, and currencies are not mixed) ↓ The user reviews and edits the extracted results. Quantity, unit price, delivery date, terms, etc. ↓ Using LLM function calling, the system searches past procurement history for the target material through (1) SQL search and (2) semantic search, and also performs (3) web search for the target material. 1. SQL search on Supabase data: retrieve the following information - firstPurchaseDate: - latestPurchaseDate: - totalPurchaseCount: - purchaseCount2025: - sameCategoryPurchaseCount: - sameProductPurchaseCount: 2. Semantic search (RAG) on Supabase data: check qualitative information such as past supplier quality issues, delivery delays, complaints, returns, and special notes. 3. Tavily API: web search for the target material; retrieve prices and site links for the target material. ↓ The LLM generates: - Recommended supplier selection, reasons for the selection, and future to-dos based on past procurement history and market prices. ↓ The user views comparisons by supplier. The user checks price, delivery date, terms, past performance, and market prices, and confirms that the information is reflected correctly. ↓ The user approves or proceeds to recheck. ↓ The user reviews the draft email to the supplier.Mifuyu, the future-state workflow is the strongest new piece here. Eleven steps with clear frontend and backend ownership, an honest flag on product equivalency as a hard problem, and a logical sequence from material input through to a saved Gmail draft. The most urgent fix is sequencing. A prototype exists and a workflow tool is being configured, but the master prompt and evaluation criteria are both empty. You are building before defining what good output looks like. Flip the order. Write evaluation criteria first: what makes a market price comparison accurate enough to trust, what makes a negotiation email effective enough to send, what counts as a correct product match. Those definitions become the rubric every prompt iteration gets tested against. Without them, Demo Day judges will ask how you know the output is good, and you will not have an answer. Second, the workflow assumes a clean path with no failures. What happens when the market price search returns nothing? What happens when the model matches the wrong product? What happens when the generated email contains a hallucinated price? A procurement professional sending a fabricated number to a supplier is a relationship-destroying event. Add confidence indicators on price matches, a manual override for uncertain equivalency, and source citations on every price figure. Third, draw a clean line in the workflow between deterministic steps and LLM steps. Comparing a quoted price against a stored purchase price is a lookup and a subtraction. That does not need a language model. The steps that genuinely need an LLM are judging whether two differently named products are functionally equivalent, synthesizing data into a negotiation rationale, and generating the email itself. That boundary shapes your architecture, your cost structure, and which steps need eval criteria versus simple unit tests. Direct answers to your questions On whether coding tools like Codex or Claude Code are mandatory: no. Lovable for the UI and n8n for the workflow is a valid path. Use whatever lets you move fastest. Coding tools become useful only if you hit a step neither tool can handle, like a custom eval script or a specific API call. On how to improve the flow: split your workflow diagram into two explicit categories. Label each step as either deterministic (rules, lookups, arithmetic) or LLM-required (reasoning, generation, synthesis). For deterministic steps, n8n handles them with simple logic nodes and no model call. For LLM steps, write a separate prompt for each, define the expected input and output format, and create three to five test cases with clear pass or fail thresholds. That separation is the single highest-leverage improvement available right now.
Build WireframesWireframesHow will users navigate through your AI solution? What are the key steps and decision points? What information will be displayed at each stage? What specific UI elements are needed on each screen? How will the layout accommodate AI features?https://excalidraw.com/#json=FV2kKFqw3hFo3kbADDH2D,9vK2zqSvbHXwEGwU2estFg
Develop Prototype to showcase AI interactionsPrototype Screens Suggested Tool: lovable.devWhat aspects of the AI solution will you demonstrate in your prototype? How will the AI inputs, processing, and outputs be presented visually to users? Which features are essential for launch? What can be left for later releases?The prototype is implemented in Next.js, so I cannot show the running prototype, but I created the same prototype screens in Lovable. https://lovable.dev/projects/4f7c21d4-82a0-4aee-b906-380998d2d2e4?magic_link=mc_9c82a303-249b-4d95-9b99-ca2366d5ff90
Initial Prompt DesignMaster Prompt [Initial Design]Create an initial master prompt. Consider the following: What tone or personality should the AI use? How should user input and system instructions be structured for maximum clarity? What system instruction will govern the AI's behavior? What examples might improve performance? How will you format outputs for consistency?"You are a Supplier Recommendation Assistant for procurement teams in the manufacturing industry. ## Role and Objective Based on user-reviewed quote JSON data, review past procurement history, internal notes, and market price information, then create a concise summary that helps procurement users make the next decision. The goal is not simply to select the cheapest supplier. Instead, organize price information, transaction history, qualitative feedback, and market price references separately, and explain in practical terms why a supplier is recommended. ## General Principles - Write concisely and practically. - Do not make overly definitive statements; only write information supported by evidence. - Do not create nonexistent suppliers, prices, history IDs, URLs, or facts. - Keep quote information, SQL-derived quantitative information, RAG-derived qualitative information, and web-derived market price information clearly separated. - If information cannot be retrieved, state ""Unable to retrieve"". - Do not provide generic advice. Include only information relevant to the current decision. ## Input Format The input includes: - comparisonTarget - Product information being evaluated - items - Reviewed quote line items - Only items with status = ""included"" should be considered for recommendations ## Information Sources ### Quote Summary - Use values from the input JSON items exactly as provided. - Do not rename, infer, or create missing quote fields. - If a field is missing or null, write ""Not provided"". ### Quantitative Information (SQL-derived) - Use values from supplierRelationshipSummary as-is. ### Qualitative Information (RAG-derived) - Summarize comments from qualityNotes only. - Capture: - Positive feedback - Areas of concern - Delivery performance trends - Quality and issue trends - Historical purchases completed at lower prices ### Market Price Information - Include only web search links and relevant notes. - Do not use market prices as the primary recommendation factor. - Treat market prices as reference information for price validation and negotiation. ## Supplier Recommendation Logic ### 1. Exclusion Rules - Do not recommend suppliers whose status is not ""included"". - If RAG-derived qualitative information contains recent major quality issues or delivery issues, treat the supplier cautiously even if it has the lowest price. - Normal negotiation history or routine confirmation items should not be treated as major concerns. ### 2. Core Evaluation - Compare prices primarily using pricePerPiece. - If pricePerPiece is unavailable, use unitPrice. - Review SQL-derived transaction history. - Review RAG-derived strengths and concerns. - As a guideline, if delivery is more than 10 days slower than other suppliers, include it in the reasoning or next actions. ### 3. Recommendation Rules - If there are no major concerns, recommend the lowest-priced supplier. - If recommending a supplier that is not the cheapest, clearly explain why: - Significantly faster delivery than competitors - Stronger historical transaction record - Positive quality or delivery history from RAG-derived records - Major concerns identified for the lowest-priced supplier - Do not use market price information as the primary basis for recommendation. ## Output Format ### 1. Supplier Summary For each supplier, present information in the following order: Supplier Name Quote Summary - supplier: - quoteDate: - productName: - modelSpecification: - unitPrice: - quantity: - totalAmountTaxIncluded: - deliveryDate: - pricePerPack: - piecesPerPack: - pricePerPiece: Past Procurement History Quantitative Information (SQL-derived) - firstPurchaseDate: - latestPurchaseDate: - totalPurchaseCount: - purchaseCount2025: - sameCategoryPurchaseCount: - sameProductPurchaseCount: Qualitative Information (RAG-derived) - Positive Feedback: - Areas of Concern: - Delivery Trend: - Quality / Issue Trend: - History of Lower-Priced Transactions: ### 2. Market Price Information - Provide reference information from web search results. - Include source links. - Sales units, package quantities, tax treatment, and shipping conditions may differ, so direct comparison with the current quotes may not be valid. - Always review the source links before making a decision. ### 3. Recommended Supplier ### 4. Reasoning ### 5. Next Actions - Include only actions relevant to the current decision, such as: - Price negotiation - Delivery confirmation - Quality verification - Do not include the same generic recommendations every time."
Prepare for Testing & IterationEvaluation Criteria & Test PlanEvaluation CriteriaWhat specific quality benchmarks (e.g., clarity, relevance, tone, accuracy, SEO, hallucination avoidance) will define “good” output?Evaluation Item: Non-hallucination What to Check: - No information is invented beyond the quote data, SQL results, RAG retrievals, or web search results. - Supplier names, prices, dates, IDs, URLs, and comments are supported by source data. Evaluation Item: Source citation What to Check: - Key facts, qualitative insights, and market price information include a clear source reference. - Users can trace information back to the relevant history record, date, or source. Evaluation Item: Missing-data handling What to Check: - Missing or unavailable information is explicitly stated. - The output does not infer or estimate information that was not retrieved. Evaluation Item: Recommendation reasoning What to Check: - The recommendation is supported by price, procurement history, qualitative insights, and delivery information. - 推奨理由の決め手が一瞬で分かるか。理由が重複せず、「なぜこの仕入れ先なのか」が明確か
Example CasesWhat specific example use cases, edge cases, and negative cases should be covered by test prompts and outputs?Use Case - The quote contains complete and valid information. - Relevant past procurement history is available. - The item exists on the indirect-materials e-commerce site used for market price research. Edge Cases - No past procurement history is available. - The quoted price or delivery date contains unrealistic values. - Past procurement history includes serious errors or issues. - The item cannot be found on the indirect-materials e-commerce site. Negative Cases - User prompt: Not applicable because the UI controls user input. - Past procurement history :cannot be retrieved through the API. - Web search:cannot be performed through the API.
DEVELOPAI Model Selection & JustificationAI Model Selection & JustificationWhich AI model is best suited for your solution and why? What capabilities and limitations does it have? How will it integrate with your product?Because the structure is simple: quote information is provided as input, past procurement history and web search information are retrieved, and the model judges the appropriate supplier. For that reason, I started by testing legacy models. Since the model retrieves fact-based information and then decides what to do based on it, I set the temperature low. The important evaluation points are whether the model can choose the appropriate supplier based on evidence, whether the response time is not too long, and whether token usage is not excessive. gpt5.4 nano (temperature: 0.2, reasoning: medium) Response time: 39 seconds (I expected this to be the fastest, but for some reason 5.4-mini was faster), tokens: 7406, judgment: not wrong, but it is not very clear what the decision was based on. gpt 5.4-mini (temperature: 0.2, reasoning: medium) Response time: 24.5 seconds, tokens: 8130, judgment: able to make a reasonable decision based on past procurement history, but the writing is hard to read. got-5.4 (reasoning: medium) Response time: 83 seconds, tokens: 8130, judgment: able to make an accurate decision based on past procurement history. Haiku 4.5 (thinking adaptive, medium) Response time: 43 seconds, tokens: 11712, judgment: able to make an accurate decision, and the writing is easy to read. Claude sonnet 4.6 (thinking: adaptive, medium) Response time: 78 seconds, tokens: 10379, judgment: able to make an accurate decision, and the writing is easy to read. Claude sonnet 4.6 (thinking: adaptive, low) Response time: 54 seconds, tokens: 10379, judgment: able to make an accurate decision, and the writing is easy to read.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Define InputsInput Specification TableRequired FieldsWhat are the required input fields for the AI (e.g., title, description, keywords, tone)? Indicate format, source, requirement.Information extracted and normalized from the target material and quote documents Quote information includes: - supplier: - quoteDate: - productName: - modelSpecification: - unitPrice: - quantity: - totalAmountTaxIncluded: - deliveryDate: - pricePerPack: - piecesPerPack: - pricePerPiece:
Optional FieldsAre there any optional or user-customizable fields? How do they impact the AI’s output?[Insert your response here]
Define Good OutputOutput Evaluation ChecklistObjective CriteriaWhat criteria will you use to judge output as “good”? (e.g., structure, use of keywords, tone, factuality, relevance)Evaluation Item: Non-hallucination What to Check: - No information is invented beyond the quote data, SQL results, RAG retrievals, or web search results. - Supplier names, prices, dates, IDs, URLs, and comments are supported by source data. Evaluation Item: Source citation What to Check: - Key facts, qualitative insights, and market price information include a clear source reference. - Users can trace information back to the relevant history record, date, or source. Evaluation Item: Missing-data handling What to Check: - Missing or unavailable information is explicitly stated. - The output does not infer or estimate information that was not retrieved. Evaluation Item: Recommendation reasoning What to Check: - The recommendation is supported by price, procurement history, qualitative insights, and delivery information. - The reasoning is consistent with the defined evaluation criteria. Evaluation Item: next action clarity What to Check: - The recommendation is supported by price, procurement history, qualitative insights, and delivery information. - The reasoning is consistent with the defined evaluation criteria.
Subjective CriteriaAre there any criteria that require human judgment or qualitative assessment?- Whether the qualitative information from semantic search, such as quality issues, is accurate. - Whether the target material can be searched successfully.
Prompt Design IterationMaster Prompt [Final Design]Prompt Version 1What is your starting system prompt for the model? What variations will you test? What techniques will you use to optimize performance? List initial instructions, persona, inputs, and constraints.You are a supplier recommendation assistant for procurement teams in manufacturing companies. ## Role and Objective Use the reviewed quote JSON to check past procurement history, internal notes, and market price references, then create a summary that helps the procurement user decide what to do next. The objective is not simply to choose the cheapest supplier. Your role is to separate quote information, past transaction data, related notes, and market price references, then explain in practical terms why a supplier is recommended. ## Basic Principles - Write concisely, practically, and based only on available evidence. - Do not overstate conclusions. - Do not invent suppliers, prices, history IDs, URLs, dates, or comments. - Keep quote JSON, SQL-derived past procurement data, RAG-derived related notes, and web-search-derived market price information clearly separated. - If information could not be retrieved, write "Not retrieved." - Do not repeat generic advice. Write only what is useful for this specific decision. ## Output Tone - Write in clear, practical business English for procurement users. - Keep the tone confident but not overstated. - Prefer natural decision-support language over audit-style wording. - Make the main recommendation easy to understand at a glance. - Use short paragraphs and focused bullets instead of long explanations. - When evidence is available, state the implication clearly. For example: "This makes Precision Edge Tools the strongest option." - When evidence is limited, be transparent without sounding alarmist. For example: "This should be confirmed before ordering." ## Length Guidelines - Keep each sentence between 10 and 20 words. - Avoid sentences longer than 30 words unless necessary for clarity. - Keep each bullet to one sentence. - In the Reason section, use 3 to 4 bullets only. - In the Next action section, use 2 to 3 bullets only. - Keep each major section concise enough to scan in a UI. ## Handling Missing Information - If a tool or search could not be executed, write "Not retrieved." - If a tool or search was executed but no relevant information was found, write "No relevant information found." - Even when no relevant information is found, do not write "no issue," "no risk," or "no record" as a definitive statement. - If missing information affects the evaluation, write "This could not be confirmed from the retrieved information" or "Additional confirmation is needed." - If a RAG-derived related note has the comment `No note`, do not treat it as quality, delivery, or trouble information. Write "No relevant comment." - In the reference history table, if the original comment is `No note`, write "No relevant comment" instead of copying `No note`. - If a search was executed and returned no matching records, use "No relevant information found," not "Not retrieved." - Avoid phrases such as "records do not show an issue." Prefer "This could not be confirmed from the retrieved information." ## Input Format The input contains: - comparisonTarget - Product information for the comparison target - items - Reviewed quote line items - Only items with status `included` should be considered for recommendation ## Information to Use - Quote summary - Use the values from the input JSON items as-is. - Do not rename, infer, or create missing quote fields. - If a field is missing or null, write "Not provided." - Supplier past procurement data (SQL-derived) - Use the values in `supplierRelationshipSummary` as-is. - Do not summarize, infer, or recalculate numbers or dates. - Past related notes (RAG-derived) - Summarize only the comments in `qualityNotes`. - Extract positive points, caution points, delivery tendencies, quality or trouble tendencies, and lower past transaction references. - Do not mix price history into "positive points." Put it under "lower past transaction references." - Under related notes, include a reference table that makes the source history ID, date, and content traceable. - The reference table should include `sourceRecordId`, `purchaseDate`, and key point. - Market price - Write only web-search-derived links and caution notes. - Limit reference links to 3 to 5 representative links. - Do not use market price as the main reason for the recommendation. Treat it as reference information for price checking or negotiation. ## How to Choose the Recommended Supplier 1. Exclusion Rules - Do not recommend suppliers whose status is not `included`. - If RAG-derived related notes show a recent serious quality or delivery issue, treat the supplier carefully even if its price is low. - Do not treat normal negotiation history or ordinary confirmation items as serious concerns. 2. Basic Evaluation - First compare prices using `pricePerPiece`. - If `pricePerPiece` is unavailable, use `unitPrice`. - Check SQL-derived past procurement data. - Check RAG-derived related notes. - Use delivery date as a decision factor only when it is clearly later than other suppliers. - As a guideline, if delivery is more than 10 days later than other suppliers, include it in the reasoning or next action. 3. Recommendation Rules - If there are no serious concerns, generally recommend the supplier with the lowest price. - If recommending a supplier that is not the cheapest, clearly explain why. - Delivery is significantly faster than others - Past procurement history is clearly stronger - RAG history provides quality or delivery reassurance - The cheapest supplier has a serious concern - Do not use market price as the main reason for the recommendation. ## Output Format 1. Supplier-by-supplier summary For each supplier, write in the following order. Supplier name Quote summary - supplier: - quoteDate: - productName: - modelSpecification: - unitPrice: - quantity: - totalAmountTaxIncluded: - deliveryDate: - pricePerPack: - piecesPerPack: - pricePerPiece: Past procurement history Past transaction data - firstPurchaseDate: - latestPurchaseDate: - totalPurchaseCount: - purchaseCount2025: - sameCategoryPurchaseCount: - sameProductPurchaseCount: Related notes - Positive points: - Caution points: - Delivery tendency: - Quality or trouble tendency: - Lower past transaction references: Reference history | sourceRecordId | purchaseDate | Key point | |---|---|---| 2. Market price - Write web-search-derived reference information. - Include 3 to 5 reference links. - Selling unit, package quantity, tax treatment, and shipping conditions may differ, so the current quote cannot be directly compared with the web prices. - The user must check the linked pages before using the information. 3. Recommended supplier - State exactly one recommended supplier from the current target suppliers. - If no supplier can be recommended, write "No recommendable supplier." 4. Reason - Start with a short, direct decision sentence. - Then explain the 2 to 4 strongest reasons only. - Avoid listing every possible factor. Focus on what actually changes the decision. - Use price, past transaction data, related notes, and delivery differences only when they affect the decision. - If the recommended supplier is not the cheapest, explain why non-price factors were prioritized. - Mention why major non-recommended suppliers were not selected only when necessary. 5. Next action - As a rule, write actions only for the recommended supplier. - Make clear what should be checked or negotiated if proceeding with this supplier. - Include actions for non-recommended suppliers only if they could change the recommendation. - Do not create actions for every supplier just because they were candidates. - Write only actions needed for this specific decision, such as price negotiation, delivery confirmation, or specification/quality confirmation. - Do not repeat generic advice. ## JSON Response Requirement Return the output above as valid JSON only. Do not wrap the JSON in Markdown. Do not include explanatory text before or after the JSON. Use exactly the following top-level keys: { "supplierSummaries": [ { "supplierName": "string", "quoteSummary": { "supplier": "string", "quoteDate": "string", "productName": "string", "modelSpecification": "string", "unitPrice": "string", "quantity": "string", "totalAmountTaxIncluded": "string", "deliveryDate": "string", "pricePerPack": "string", "piecesPerPack": "string", "pricePerPiece": "string" }, "pastTransactionData": { "firstPurchaseDate": "string", "latestPurchaseDate": "string", "totalPurchaseCount": "string", "purchaseCount2025": "string", "sameCategoryPurchaseCount": "string", "sameProductPurchaseCount": "string" }, "relatedNotes": { "positivePoints": "string", "cautionPoints": "string", "deliveryTendency": "string", "qualityOrTroubleTendency": "string", "lowerPastTransactionReferences": "string" }, "referenceHistory": [ { "sourceRecordId": "string", "purchaseDate": "string", "keyPoint": "string" } ] } ], "marketPrice": { "summary": "string", "observedMarketRange": "string", "cautionNotes": ["string"], "referenceLinks": [ { "title": "string", "url": "string" } ] }, "recommendedSupplier": "string", "reason": { "decisionSentence": "string", "bullets": ["string"] }, "nextAction": ["string"] } Map each original output section to JSON as follows: - Supplier-by-supplier summary -> supplierSummaries - Market price -> marketPrice - Recommended supplier -> recommendedSupplier - Reason -> reason - Next action -> nextAction Use "Not provided", "Not retrieved", "No relevant information found", or "No relevant comment" inside the JSON values when required by the rules above.
Prompt IterationsIf revised, what changes did you make and why? How do you track and record prompt evolution?What changes were made to the prompt - The reasons for the recommended supplier and the checklist content included too much information, making it unclear what the actual point was. -> First state the reason briefly in one line, then add supporting details. - No tone was specified, so the output varied depending on the input. -> Added tone guidance. - Some cases produced long responses. -> Added word-count constraints. - It was unclear whether information could not be retrieved at all or whether a search was performed and no results were found. -> Added instructions to clearly state whether information was retrieved or not. How do I track and record prompts: I record prompts and outputs in the Cursor directory, then have Codex evaluate the outputs against the evaluation criteria on a 1-3 scale.
Data Preparation & RAG ImplementationData Preparation & RAG ImplementationWhat data sources will you use? How will you prepare data for model training or evaluation? (e.g., cleaning, structuring). For RAG: How will you chunk, embed, retrieve relevant information?Data preparation: Past procurement history data: supplier, purchase date, purchased product, delivery date, quote deadline, and notes such as quality issues are organized into one table and stored in Supabase (using mock data). RAG: Vectorized past procurement history data is stored in Supabase. OpenAI text-embedding-3-small is used for embeddings. pgvector is used for vector search.
Create Evaluation SetExample Input/Output Data for TestingTypical ExamplesWhat are the most common inputs and expected outputs? Use real data if possible.Quote documents with the necessary information, suppliers with past procurement history, and target materials that can be searched on MonotaRO as indirect materials.
Edge Cases & Negative CasesWhat examples test the AI’s limits? (e.g., missing data, ambiguous input, out-of-domain)Edge Cases - No past procurement history is available. - The quoted price or delivery date contains unrealistic values. - Past procurement history includes serious errors or issues. - The item cannot be found on the indirect-materials e-commerce site. Negative Cases - User prompt: Not applicable because the UI controls user input. - Past procurement history: cannot be retrieved through the API. - Web search: cannot be performed through the API.
Test Example Data & Review ResultsManual ReviewRun your input data with the prompt. How did your output perform in manual review? Which examples failed which criteria, and why?Failed the evaluation criterion "non-hallucination": when the item cannot be found on the indirect-materials e-commerce site, completely unrelated material links were retrieved.
Automated EvaluationWhat pass/fail rate or scores did the AI achieve on core criteria?Implemented 10 runs: Evaluation Item: Non-hallucination -> 8/10 passed Evaluation Item: Source citation -> 7/10 passed Evaluation Item: Missing-data handling -> 9/10 passed Evaluation Item: Recommendation reasoning -> 10/10 passed Evaluation Item: next action clarity -> 10/10 passed
Handle Edge Cases & IterateEdge Case IdentificationWhat edge cases did you identify in testing or real usage?
Updates & AdjustmentsWhat prompt or system adjustments have you made based on failures, feedback, or edge case observations?Improved the system prompt so that when a web page for the target material cannot be found, the output says that it was not found.
Automate Evaluation ApproachEvaluation MethodWhat is your chosen approach for evaluation (human, model grader, script)? How will you scale testing to diverse/large test sets?LLM as judge
Evaluation FrequencyHow often will you re-run evaluations for new data, new prompts, or post-launch monitoring?weekly
DEPLOYFinalize Launch & Rollout PlanOperational Readiness ChecklistTechnical ReadinessIs infra (APIs, databases, rate limits, monitoring, rollback) tested and documented?APIs, databases, and rate limits have been tested and documented. However, monitoring and rollback have not yet been tested or documented, so they need to be addressed going forward.Please leave this area blank. This space is for the Instructor to provide you with feedback.
Organizational ReadinessHave internal teams (support, comms, legal) been trained? Is documentation complete?Documentation and formal training for internal teams are planned for the future.
Launch & Rollout StrategyLaunch ApproachWhat is your launch approach? Pilot, AB test, or all users—who gets access and when?At this point, the product will start with a limited pilot rather than a full release to all users. First, access will be provided to a small team close to the target users, such as procurement staff in manufacturing companies, to validate the practical workflow from quote upload, AI extraction, comparison, recommendation reasoning, and email drafting. A/B testing is not planned for the initial launch. After using the pilot to assess accuracy, usability, risks, and support load, the target user base will be expanded gradually. The official access start date and scope will be decided after pilot preparation is complete.
Scale ReadinessHow will you ensure readiness for scale? How will you monitor initial volume and scale up?
Go-to-Market PlanMarketing / Training AssetsWhat assets (FAQ, demo, guides) will you prepare for external communication/marketing?Before launch, the following external-facing materials will be prepared: - Product overview deck - Demo video - First-use guide - FAQ
Stakeholder / Internal CommsHow will you communicate launch plans, progress, and outcomes internally?Launch plans, progress, and results will be managed through regular internal updates and documentation. Specifically, the launch plan will be summarized in a shared document that clearly states target users, release scope, schedule, risks, and decision criteria. Progress will be shared through weekly meetings and Teams, where development status, test results, and unresolved issues will be reviewed. After launch, usage status, errors, user feedback, support inquiries, and improvement points will be summarized and reported to stakeholders. If needed, a retrospective will also be held to decide the next release scope and improvement priorities.
Confirm Legal, Privacy & Risk ProtocolsData & PrivacyHow do you handle and protect user data, including storage, privacy, and compliance?User data will be handled only to the minimum extent necessary and managed securely on the server side. ProcureLens handles quote files, extracted quote data, comparison results, recommendation reasons, and past procurement history. For the MVP, uploaded original files are expected not to be stored permanently and will be used only during comparison processing. Extracted comparison data and history data will be stored in Supabase. For compliance, data retention periods, the scope of data sent to external AI/APIs, deletion policies, terms of use, and the privacy policy will be organized before launch. At this point, a formal legal and compliance review still needs to be completed.
Policy & ComplianceAre content moderation, legal, and audit processes in place? Are you compliant with regulations needed for your domain?At this point, this has not yet been formally established. It will be addressed before launch. Because ProcureLens is a procurement support product, it handles quote data, supplier information, past purchase history, and AI-generated recommendation reasons. Therefore, legal review, data handling rules, scope of external AI/API usage, logs and audit trails, terms of use, and privacy policy need to be organized.
Define Success MetricsSuccess MetricsUser/Business MetricsWhat user metrics will indicate success? What business metrics will demonstrate value?User Metrics - Time saved per quote comparison - User satisfaction from pilot users Business Metrics - Estimated cost reduction per comparison
AI MetricsHow will you measure AI performance and accuracy?- Hallucination rate - Source citation rate - Missing-data handling accuracy - Response latency
Monitor, Iterate & ImproveUser Support & Feedback PlanSupport ChannelsWhere can users get support? Is escalation and ownership clear?At this point, a formal support channel and escalation process have not yet been established. They will be clarified before the pilot starts. First, an inquiry form will be prepared for pilot users, and users will be able to send inquiries from the AI recommended supplier results screen.
Feedback WorkflowHow do you gather, triage, and act on feedback and bugs? How are critical issues prioritized and communicated?Before the pilot starts, the feedback and bug management process will be established. Feedback will be collected through an on-screen inquiry form. Inquiry content will be shared into a tool such as Asana, and Ops will handle items in order. Priority will be determined based on impact scope and severity. Issues such as data leakage, business risk caused by incorrect recommendations, service outages, or problems that prevent completion of the main flow will be handled with the highest priority. If needed, affected users will also be informed of the situation, workaround, and planned fix. After resolution, the cause and recurrence prevention measures will be recorded.
Monitoring & Continuous ImprovementMonitoring ApproachWhat monitoring/logging is in place to spot operational/AI issues post-launch?At this point, minimal logs exist, but a formal monitoring system has not yet been established. However, the full monitoring needed after launch, such as continuously tracking error rates, response time, external API failures, database connection errors, and cost spikes, will be developed going forward.
Ongoing ImprovementHow will you collect learnings, review performance, and update your system continuously post-launch?After launch, usage logs, user feedback, support inquiries, and bug reports will be collected continuously.
Download the .xlsx ↓