Claude Certified Architect Professional Prep Course
← सभी पाठ
पाठ 02Claude Certified Architect Professional Prep Course

एंटरप्राइज इंटीग्रेशन और प्रोडक्शन

सारांश ऑडियो

इस पाठ के लिए कोई ऑडियो सारांश नहीं है।

अध्ययन नोट्स

स्क्रीन 1: ओरिएंटेशन: अंत तक आप क्या कर सकेंगे

मॉड्यूल शेल 2 मिनट ओरिएंटेशन: अंत तक आप क्या कर सकेंगे

मॉड्यूल 1 ने आर्किटेक्चरल अवधारणाओं का परिचय दिया। यह मॉड्यूल विशिष्टताओं में गहराई से जाता है।

यह मॉड्यूल आपको वे विशिष्ट उपकरण देता है जो एक प्रोटोटाइप को प्रोडक्शन सिस्टम से अलग करते हैं: बिल्ड करने से पहले एक क्वालिटी गेट, डिप्लॉय करने से पहले एक कॉस्ट और रिलायबिलिटी मॉडल, कमिट करने से पहले एक फीजिबिलिटी फ्रेमवर्क, एक इंटीग्रेशन आर्किटेक्चर जो सिक्योरिटी रिव्यू से बचता है, और एक एक्सपेरिमेंटेशन मेथड जो आपको बताता है कि क्या परिवर्तन वास्तव में काम करते हैं।

इस मॉड्यूल के अंत तक, आप सक्षम होंगे:

1 सफलता के मानदंड को परिभाषित करें और प्रोडक्शन कोड की पहली लाइन लिखने से पहले एक eval सूट बनाएं, मॉडल-आधारित evals को कोड-आधारित evals से अलग करें, eval वर्कफ़्लो चरणों का चयन करें, और किसी भी परिवर्तन के लिए gating मैकेनिज्म के रूप में evals का उपयोग करें एक प्रोडक्शन सिस्टम में। 2 POC-से-प्रोडक्शन चेकलिस्ट के माध्यम से काम करें, कॉस्ट और लेटेंसी को एक बजट में मैप करें, रिलायबिलिटी पैटर्न निर्दिष्ट करें (retries, fallbacks, circuit breakers), चुने गए आर्किटेक्चर के लिए विफलता मोड का नाम दें, और प्रत्येक के लिए शमन को स्पष्ट करें, जिसमें agents (सिस्टम जो tools का उपयोग करते हैं, turns में कारण देते हैं, और बहु-चरणीय कार्य करते हैं) को प्रोडक्शन-विश्वसनीय बनाना शामिल है। 3 कॉल वॉल्यूम, token खपत, और कॉस्ट का अनुमान लगाकर एक use case बनाएं, AI Capabilities and Limitations से चार AI गुणों के विरुद्ध तकनीकी व्यवहार्यता का आकलन करें, और एक व्यावसायिक समस्या को स्पष्ट सीमा शर्तों के साथ एक scoped समाधान आर्किटेक्चर में अनुवाद करें। 4 एक Claude डिप्लॉयमेंट को आर्किटेक्ट करें जो एंटरप्राइज के लिए तैयार है compliance के लिए इंटीग्रेशन पैटर्न निर्दिष्ट करके (regulated-industry constraints, BAA coverage, data-residency), identity (SSO/OAuth), authorization, data handling, और observability instrumentation। प्रत्येक इंटीग्रेशन पॉइंट पर सही इंटीग्रेशन पॉइंट (API, SDK, MCP, Claude Code) रखें। 5 एक live Claude सिस्टम पर एक A/B टेस्ट या structured experiment की योजना बनाएं और interpret करें, hypothesis सेट करें, metrics का चयन करें, आवश्यक sample size का अनुमान लगाएं, और overclaiming के बिना एक परिणाम को पढ़ें।

यह मॉड्यूल एक मजबूत systems background मानता है और सीधे मॉड्यूल 1 पर बनता है। यह foundational concepts को छोड़ता है और उन निर्णयों पर गहराई से जाता है जो एक प्रोटोटाइप को प्रोडक्शन सिस्टम से अलग करते हैं।

शैक्षणिक सामग्री के लिए अस्वीकरण / नोटिस

हमने यह Architect कोर्स मॉड्यूल 2: Enterprise Integration & Production Claude के साथ वास्तविक काम करने में आपकी मदद करने के लिए बनाया। इसे शैक्षणिक सामग्री के रूप में मानें। यह कानूनी, वित्तीय, या अन्य व्यावसायिक सलाह का गठन नहीं करता है, इसलिए आप जो सीखते हैं उसे अपनी स्थिति के अनुसार अनुकूलित करें। हमारे उत्पाद और सेवाएं तेजी से विकसित होती हैं, इसलिए कुछ सामग्री में त्रुटियां हो सकती हैं या पुरानी हो सकती हैं; Anthropic की वेबसाइट या docs पर सत्यापित करना याद रखें। कोर्स में उपयोग किए गए उदाहरण और परिदृश्य illustrative हैं और अक्सर काल्पनिक हैं। यदि कोर्स सामग्री किसी कंपनी या उत्पाद का उल्लेख करती है, तो इसका मतलब यह नहीं है कि Anthropic उनका समर्थन करता है, वे Anthropic का समर्थन करते हैं, या कि हम संबद्ध हैं। यह भी ध्यान दें कि Anthropic उत्पादों और सेवाओं का आपका उपयोग हमारी शर्तों, नीतियों और documentation द्वारा कवर किया गया है; यदि इस कोर्स में कुछ उनके साथ विरोध करता है, तो वे नियंत्रण करते हैं।

स्क्रीन 2: Evals को स्वीकृति मानदंड के रूप में: बिल्ड प्रक्रिया में गुणवत्ता बनाना

शिक्षण Evals17 मिनट Evals को स्वीकृति मानदंड के रूप में: बिल्ड प्रक्रिया में गुणवत्ता बनाना पहले मॉड्यूल में, आपने आर्किटेक्चर निर्णय लिए: पैटर्न, इंटीग्रेशन पॉइंट, और आपका सिस्टम विभिन्न inputs पर कैसे प्रतिक्रिया करना चाहिए। आपके पास अभी तक यह आत्मविश्वास नहीं है कि जब inputs वास्तविक हों और users अप्रत्याशित हों तो वे निर्णय कायम रहते हैं।

यह वह जगह है जहां evals चित्र में आते हैं। Evals evaluations के लिए short हैं, और वे आपको प्रोडक्शन में जाने से पहले या मॉडल अपडेट के बाद आपके सिस्टम के व्यवहार को परीक्षण करने देते हैं, इसलिए आप समस्याओं की खोज करते हैं इससे पहले कि वे आपके users को मिलें। यह सेक्शन समझाता है कि evals क्या हैं, वे क्यों महत्वपूर्ण हैं, और समस्याओं से आगे रहने के लिए उनका उपयोग कैसे करें।

कोड से पहले Evals: क्यों क्रम महत्वपूर्ण है एक eval, या evaluation, एक structured test है जो जांचता है कि क्या आपका सिस्टम expected और accurate output देता है। हालांकि यह सीधा लगता है, timing अत्यंत महत्वपूर्ण है। मानक दृष्टिकोण सिस्टम बनाना है, देखना है कि यह सही दिखता है या नहीं, और बाद में परीक्षण करना है। यह हमेशा सर्वोत्तम दृष्टिकोण नहीं है।

प्रोडक्शन कोड लिखने से पहले अपना eval suite लिखना तीन चीजों को होने के लिए मजबूर करता है जो अन्यथा defer करना आसान है:

पहला, सफलता का अर्थ measurable terms में बताएं दूसरा, design assumptions को जल्दी expose करें, जब उन्हें बदलना अभी भी सस्ता है तीसरा, अपने आप को एक gate दें जो यह निर्धारित कर सकता है कि क्या एक model swap, एक prompt change, या एक नई retrieval strategy ने measurably सिस्टम में सुधार किया है

एक eval suite बिल्ड के अंत में QA step के रूप में नहीं, बल्कि शुरुआत में, प्रोडक्शन कोड लिखने से पहले, परिभाषित होना चाहिए। वास्तव में, यदि आप किसी व्यवहार के लिए एक eval नहीं लिख सकते, तो आपके पास यह verify करने का कोई विश्वसनीय तरीका नहीं है कि वह व्यवहार मौजूद है। इसका मतलब है कि आप सिस्टम में जो भी परिवर्तन करते हैं वह verifiable नहीं है। शुरुआत में एक eval suite जोड़ने से आप पूरे बिल्ड में verify कर सकते हैं।

Eval वर्कफ़्लो कैसे चलता है: task definition से result तक एक well-constructed eval वर्कफ़्लो नीचे दिए गए चरणों के माध्यम से sequentially चलता है। प्रत्येक चरण एक artifact produce करता है जो अगले चरण में feeds करता है:

चरण | क्या होता है | आउटपुट

  • Task को परिभाषित करें | आप जिस व्यवहार का मूल्यांकन कर रहे हैं उसे specific, measurable terms में बताएं, और वह prompt लिखें जिसका उपयोग आप इसे परीक्षण करने के लिए करेंगे। एक vague definition एक vague eval produce करता है। behavioral specification और prompt की concreteness का स्तर ही है जो परिणाम को meaningful बनाता है। | Task specification with prompt to test and pass criteria
  • Golden dataset बनाएं | उन inputs को assemble करें जो आपका सिस्टम encounter करेगा, जिसमें edge cases और counterexamples शामिल हैं। यह dataset वह है जिसके विरुद्ध आपका eval चलता है। यदि dataset representative नहीं है, तो scores meaningful नहीं हैं। | Labeled dataset with expected outputs
  • Automated checks चलाएं | प्रत्येक prompt को सिस्टम के माध्यम से pass करें और output को अपने expected result के विरुद्ध compare करें। Automated checks fast और cheap हैं। उन्हें उन व्यवहारों के लिए उपयोग करें जो unambiguous हैं: format compliance, schema validation, और authoritative data के विरुद्ध factual lookups। | Pass/fail record per item
  • Judge के साथ Score करें | उन व्यवहारों के लिए जिनके लिए interpretation की आवश्यकता है, जैसे tone, reasoning की accuracy, और edge-case responses की appropriateness, एक model-based judge outputs को scale पर assess कर सकता है। | Score per item with reasoning
  • Interpret और Act करें | Aggregate scores आपको बताते हैं कि सिस्टम कहां है और क्या एक परिवर्तन इसे सही दिशा में ले गया है। एक परिवर्तन जो mean score को बढ़ाता है जबकि quietly edge cases या adversarial inputs पर performance को degrade करता है, सिस्टम को बेहतर नहीं बनाता है। | Overall score, per-category breakdown

Model-based बनाम code-based evals: प्रत्येक का उपयोग कब करें आपको जिस प्रत्येक व्यवहार का मूल्यांकन करने की आवश्यकता है उसे same तरीके से check नहीं किया जा सकता। कुछ व्यवहारों का एक single correct answer है: output या तो एक valid JSON है, या नहीं है। दूसरी category यह है कि क्या output language के expected tone या style से मेल खाता है। ये दोनों categories के व्यवहार को different evaluation tools की आवश्यकता है, और किसी दिए गए व्यवहार के लिए सही चुनना accuracy सुनिश्चित करने और costs बचाने के लिए महत्वपूर्ण है।

तीन प्रकार के evals में different speed-versus-flexibility tradeoffs हैं:

Code-based evals milliseconds में deterministic checks चलाते हैं और लगभग कुछ भी cost नहीं करते। Model-based evals एक judge model का उपयोग करते हैं उन outputs को assess करने के लिए जिनके लिए interpretation की आवश्यकता है और लगभग model call जितना cost करते हैं। Human-review evals high-stakes या novel व्यवहारों के लिए human judgment पर निर्भर करते हैं जहां न तो code और न ही एक model judge को reliably evaluate करने के लिए विश्वास किया जा सकता है। Human-review evals सबसे slow और most expensive विकल्प हैं।

Eval type | कैसे काम करता है | कब उपयोग करें | Cost | सीमा

Code-based eval | एक function programmatically output को check करता है: schema validation, regex match, JSON parse, length check, authoritative data के विरुद्ध assertion। | कोई भी व्यवहार जो unambiguous है। Format compliance, schema correctness, lookup accuracy, length constraints। | बहुत कम: प्रति check milliseconds, कोई API call नहीं। | उन व्यवहारों को assess नहीं कर सकता जिनके लिए interpretation की आवश्यकता है। Tone, helpfulness, reasoning quality, और edge-case appropriateness सभी को judgment की आवश्यकता है जो एक function supply नहीं कर सकता। Model-based eval | एक judge model को original prompt, system output, और एक scoring rubric मिलता है। Judge एक score और reasoning return करता है। Judge prompt स्वयं एक prompt है जिसे engineer और test करने की आवश्यकता है। | कोई भी व्यवहार जो interpretation की आवश्यकता है: response quality, instruction following, reasoning accuracy, safety, और ambiguous inputs की handling। | मध्यम से उच्च: प्रति evaluated item एक API call, judge model के per-token rate पर scale पर, यह जोड़ता है। | Judge models borderline cases में inconsistent हो सकते हैं। Judge को score के साथ reasoning produce करने के लिए force किए बिना, वह inconsistency detect करना मुश्किल है। Human-review | एक human evaluator output को पढ़ता है और इसे एक rubric या criteria के set के विरुद्ध score करता है। यह structured (एक scoring sheet) या unstructured (open annotations और feedback) हो सकता है। | High-stakes या novel व्यवहार जहां न तो एक function और न ही एक judge model को विश्वास किया जा सकता है: safety-critical edge cases, नए capability areas बिना established rubrics के, या कोई भी output जहां एक गलत evaluation significant risk carry करता है। Calibrating और validating model-based evals के लिए भी उपयोगी। | उच्च: human time सबसे expensive resource है, और throughput limited है। Sampled subsets के बिना scale पर viable नहीं है। | Slow, expensive, और sampled subsets से परे scalable नहीं। Human evaluators भी अपना inconsistency introduce करते हैं।

Grading ladder: कैसे grade करें चुनना प्रत्येक व्यवहार को same तरीके से grade नहीं किया जाना चाहिए, और grading method की choice एक deliberate ladder को follow करती है। सबसे सस्ता reliable method के लिए पहुंचें और केवल तभी climb करें जब व्यवहार इसकी मांग करे।

Code-based grading, जहां भी व्यवहार इसे allow करता है। Deterministic checks, जिसमें schema validation, exact match, length, और presence शामिल हैं, milliseconds में चलते हैं, लगभग कुछ भी cost नहीं करते, और कभी drift नहीं करते। यदि एक व्यवहार को code में check किया जा सकता है, तो इसे होना चाहिए। LLM-as-judge, जब व्यवहार को interpretation की आवश्यकता है। उन outputs के लिए एक judge model का उपयोग करें जिनके लिए judgment की आवश्यकता है। Detailed rubrics, constrained verdicts (एक small fixed set of labels rather than free-form scores), calibration against human-labeled examples, और एक different model के साथ grading का उपयोग करके judging को rigorous बनाएं जिसके outputs आप evaluate कर रहे हैं, self-preference से बचने के लिए। Human grading, last resort के रूप में। High-stakes या novel व्यवहारों के लिए human review reserve करें जहां न तो code और न ही एक calibrated judge अभी तक trustworthy है। यह सबसे expensive और least scalable विकल्प है।

Judge calibration: वह चरण जो कई teams छोड़ते हैं एक LLM judge स्वयं एक system है जो गलत हो सकता है। इसके verdicts पर विश्वास करने से पहले, सुनिश्चित करें कि इसे calibrate करें। ऐसा करने के लिए, इसे human-labeled outputs के एक set के विरुद्ध चलाएं और confirm करें कि human judgment के साथ इसकी similarity rely करने के लिए काफी high है। एक uncalibrated judge confident scores produce करता है जो high quality नहीं हो सकते। यह कोई automated grade नहीं होने से बदतर है, क्योंकि यह trustworthy लगता है।

Volume को perfection पर favor करें। कई automatically gradable cases manually-graded ones की एक handful को beat करते हैं: broad, cheap coverage अधिक regressions को catch करता है than एक small, painstaking set, और यह हर change पर चल सकता है।

Success criteria को परिभाषित करना: एक business requirement को एक measurable threshold में बदलना एक business requirement जैसे "summarize claims accurately" वास्तव में आपको नहीं बताता कि क्या measure करना है। इसे एक eval criterion में बदलने की प्रक्रिया के निम्नलिखित चरण हैं:

व्यवहार को specifically identify करें: "Summarize claims accurately" को "extract the filer's name, claim number, incident date, और claimed amount from each document" में update किया जाना चाहिए। Threshold सेट करें: तय करें कि क्या passing count करता है। यदि, उदाहरण के लिए, आपके thresholds structured fields पर 100% accuracy, 2% से कम hallucination rate, और 99. 5% of the time response within schema हैं, तो वे numbers business requirement से आने चाहिए। बस अपने first prototype को achieve करने के लिए choose न करें; platform. claude. com/docs/en/test-and-evaluate/develop-tests पर eval thresholds सेट करने के लिए guidance खोजें। Failure modes को identify करें: उसी उदाहरण को जारी रखते हुए, एक output जो acceptable नहीं हो सकता है में एक fake claim number, एक missing incident date, या एक value from the wrong claim शामिल हो सकता है। प्रत्येक failure mode आपके eval dataset में एक category है। Adversarial inputs शामिल करें: इसमें missing fields, handwritten sections, unusual formatting, और non-standard layouts वाले documents शामिल होने चाहिए। यदि आपका golden dataset केवल clean inputs contain करता है, तो आपके eval scores production performance को predict नहीं करेंगे।

Evals को change के लिए gating mechanism के रूप में एक production Claude system में हर परिवर्तन, चाहे वह एक model swap हो, एक prompt revision हो, एक context strategy change हो, या एक retrieval configuration update हो, development के दौरान eval suite के माध्यम से चलना चाहिए इससे पहले कि यह production में जाए। यह जानने का एकमात्र विश्वसनीय तरीका है कि क्या एक परिवर्तन सिस्टम में सुधार किया है।

एक single-turn eval set आपको नहीं बताएगा कि सिस्टम एक conversation में कैसे hold up करता है। Multi-turn evals एक separate category हैं जो एक single prompt और response के बजाय exchanges के एक sequence पर सिस्टम को score करते हैं। एक multi-turn eval कुछ criteria को check करता है: क्या सिस्टम prior context को turns में straight रखता है, क्या यह एक follow-up prompt का जवाब देता है बिना details को invent किए जो conversation में पहले कभी नहीं कहे गए थे, और क्या output quality conversation के रूप में चलता है जब conversation लंबा चलता है। क्योंकि scored किया जा रहा unit एक whole conversation है, इस category को अपना golden dataset की आवश्यकता है। यह full conversation transcripts से consist करता है known high quality responses के साथ प्रत्येक turn पर, covering the follow-ups, topic shifts, और conversation lengths जो सिस्टम production में देखेगा।

एक team पर विचार करें जो एक document summarization workflow बना रहा था जिसने अपने summarization prompt को revise किया लेकिन अपने eval suite को match करने के लिए update नहीं किया। Eval suite सभी required checks को pass किया। Production में swap के दो दिन बाद, field reports दिखाते हैं कि multi-clause legal sentences को summaries में truncate किया जा रहा था। Root cause यह था कि eval set prompt change से पहले का था और behavior से match नहीं करता था जो changed था।

संक्षेप में, आप हर change से पहले evals चलाना चाहिए eval set को current रखने के लिए system के साथ जो measure कर रहा है।

Cost · Complexity · Risk Cost: हर model-based eval एक API call है। हालांकि cost को mind में रखना worth है, इसे आपको एक smaller eval set की ओर drive न करने दें than आपके use case को warrant करता है। बड़ा risk under-evaluating है: एक production-breaking change जो एक undersized eval suite के माध्यम से slip करता है एक few extra API calls से far अधिक cost करता है। अपने dataset को size करें जो आपको अपने results में confidence देता है और cost को एक secondary constraint के रूप में treat करें। Complexity: Eval infrastructure एक parallel system को maintain करने के लिए जोड़ता है। Golden dataset को current रखा जाना चाहिए, judge prompts को engineer और test किया जाना चाहिए, और pass thresholds को revisit किया जाना चाहिए जब system की requirements change करती हैं। अगले कहां जाएं: Working implementation patterns के लिए, जिसमें grading designs और golden-answer comparison का एक walkthrough शामिल है, github. com/anthropics/claude-cookbooks/blob/main/misc/building_evals. ipynb पर Claude Cookbooks देखें। Risk: एक out-of-date eval suite false confidence provide करता है। यह एक misleading impression create करता है कि एक change safe है जब checks उस behavior को measure कर रहे हैं जो अब system में exist नहीं करता है। एक regression के लिए highest risk moment तब है जब evals present हैं और out of date हैं।

स्क्रीन 3: Eval suite जिसने गलत चीज़ को measure किया

Watch Out Evals5 मिनट Eval suite जिसने गलत चीज़ को measure किया

यह गलती क्यों होती है जब एक demo काम कर रहा है, इसे "good enough" declare करना एक reasonable call लगता है। Manual spot-checks को time लगता है, और system हर tested input पर correctly respond करता लगता है। जो इसे एक issue बनाता है वह यह है कि teams केवल उन inputs को test कर सकते हैं जिनके बारे में उन्होंने सोचा था, जबकि production बाकी को surface करता है।

एक partner postmortem जिसे reconstruct करने की आवश्यकता है निम्नलिखित एक composite postmortem है जो field deployments में repeatedly एक pattern को represent करता है। Team ने एक professional services firm के लिए एक contract review assistant बनाया। उन्होंने assistant को दस contracts के विरुद्ध manually test किया जो उनकी team को अच्छी तरह से जानते थे, system को ready declare किया, और फिर production में चले गए। Regression दो हफ्ते बाद आया।

Postmortem ने क्या पाया System एक class of contracts पर fail कर रहा था जिसके विरुद्ध इसे कभी test नहीं किया गया था, अर्थात् non-standard obligation structures वाले। Model wrong section से obligations को extract कर रहा था। Eval suite exist करता था। यह project की शुरुआत में बनाया गया था और development के दौरान team द्वारा उपयोग किए जाने वाले दस contracts से बनाया गया था। यह full contract population का एक representative sample नहीं था। जब prompt changed, eval suite passing रहा, लेकिन केवल क्योंकि golden dataset अभी भी old prompt के expected outputs को reflect करता था। इसे कभी नए behavior को account करने के लिए update नहीं किया गया था।

वे निर्णय जो हमें यहां ले आए

Eval dataset को convenient inputs से rather than एक representative sample से बनाया गया था। एक eval suite जो production में face करने वाले input distribution को cover नहीं करता है एक different system को measure कर रहा है than आप ship कर रहे हैं। Eval input set को prompt change के बाद update नहीं किया गया था। Eval हर बार passing रहा क्योंकि यह उस behavior को assess कर रहा था जो prompt अब produce नहीं करता था। Scores stable दिख रहे थे क्योंकि कुछ भी नहीं test कर रहा था जो changed था। Manual spot-checks को एक eval substitute के रूप में treat किया गया था। Spot-checks confirm कर सकते हैं कि एक specific input एक specific output produce करता है, लेकिन वे आपको नहीं बता सकते कि क्या system unknown inputs पर correctly behave करता है।

क्या Watch Out करें Eval suite present था लेकिन दो तरीकों से misconfigured था: dataset representative नहीं था, और यह current नहीं रखा गया था। दोनों problems invisible हैं जब तक production gap को expose नहीं करता। System development में healthy दिख रहा था क्योंकि यह केवल उन inputs पर test किया गया था जिसके लिए यह पहले से optimize किया गया था।

स्क्रीन 4: Eval types को sort करें

Checkpoint Evals5 मिनट Eval types को sort करें नीचे आठ evaluation tasks listed हैं। प्रत्येक को उस bucket में drag करें जहां यह belongs: model-based eval या code-based eval। सही placement इस बात को follow करता है कि क्या checked किया जा रहा behavior को interpretation की आवश्यकता है या straightforward है।

  • जांचें कि हर response valid JSON है जो defined output schema को match करता है।
  • एक complex analysis में model की reasoning को assess करें कि क्या यह sound और complete है।
  • Verify करें कि response length 500 tokens के अंतर्गत है।
  • Score करें कि system एक emotionally charged customer complaint को कितनी appropriately handle करता है।
  • Confirm करें कि extracted claim number source document में quantitative value को match करता है।
  • Evaluate करें कि क्या summary एक long briefing document से most important points को capture करता है।
  • Verify करें कि सभी required sections (executive summary, methodology, findings, recommendations) output में present हैं।
  • Assess करें कि क्या एक customer-facing message का tone brand के लिए appropriately professional है।

Code-based eval

Model-based eval

उत्तर जांचें अभी के लिए छोड़ें

स्क्रीन 5: POC से production तक: cost, latency, और reliability

शिक्षण POC to Prod16 मिनट POC से production तक: cost, latency, और reliability आपका eval suite आपको बताता है कि क्या system correctly behave करता है, लेकिन कुछ भी नहीं कि क्या यह उस volume पर correctly behave करने को afford कर सकता है जो आपके partner को expect करता है। वह gap Proof of Concept (POC)-to-production gap है, और इसके चार dimensions हैं: cost, latency, reliability, और failure modes। सभी चार एक demo में invisible हैं।

एक POC आपको भी पहला signal देता है कि क्या system उस business metric को move करता है जिसे improve करने के लिए यह design किया गया था। वह signal deliberately capture करने के लिए worth है: production cost को model करने से पहले भी, अपने POC sample पर partner को care करने वाले outcome को measure करें। एक cost profile जो एक system पर budget में fit करता है जो उस metric में measurable improvement produce नहीं करता है जो matters है अभी भी एक failed deployment है।

जहां एक Proof of Concept (POC) और एक production system differ करते हैं एक POC capability को demonstrate करने के लिए design किया गया है। यह low volume पर चलता है, clean inputs पर, और एक patient user के साथ। एक production system उस volume पर चलता है जो आपके partner का business generate करता है, real inputs पर, और users के साथ जिनके पास slow या incorrect responses के लिए कोई tolerance नहीं है। चार dimensions जहां एक POC आपको mislead करता है वह cost, latency, reliability, और failure modes हैं।

Dimension | क्यों यह एक demo में invisible है | जब यह fail करता है तो यह कैसा दिखता है

Cost | एक POC जो 10–50 requests per day चलता है एक negligible bill produce करता है। Production volume पर monthly cost projections एक different calculation हैं। | Billing dashboard एक cost दिखाता है जो project approval पर signed off budget को exceed करता है। Architecture को deployment के बाद renegotiate किया जाना चाहिए। Latency | एक demo typically एक बार में एक request चलाता है। Concurrent load के अंतर्गत latency p95 एक different number है than एक single request के अंतर्गत median latency। | SLA breaches और user abandonment। Latency जो एक demo के लिए acceptable है एक real-time user-facing workflow के लिए unacceptable हो सकता है। Reliability | एक POC के पास कोई retry logic नहीं है, कोई fallback नहीं है, और कोई circuit breaker नहीं है। जब यह fail करता है, developer refresh करता है और फिर से try करता है। कोई users waiting नहीं हैं। | Retry logic या fallback handling के बिना, कोई भी transient API failure पूरे user-facing workflow को down ले जाता है rather than gracefully degrade करने के। Failure modes | एक demo को उन inputs पर test किया जाता है जो developer expect करता है। Production उन inputs को raise करता है जो developer expect नहीं करता है। Failure modes architecture type के लिए specific हैं। | Silent degradation, made-up outputs on edge-case inputs, या एक input class पर complete failure जो कभी test नहीं किया गया था।

Cost और latency modeling: Build करने से पहले numbers को जानें Cost और latency model को architecture को finalize करने से पहले build किया जाता है। तीन inputs जो आपको चाहिए वह हैं call volume (requests per day या per month), token budget per request (input tokens plus expected output tokens), और model tier। उन तीन से, आप monthly cost का अनुमान लगा सकते हैं और इसे अपने budget ceiling के विरुद्ध check कर सकते हैं कोई code लिखने से पहले।

Token budget per request वह जगह है जहां most cost models गलत जाते हैं। Teams average token count को inputs पर calculate करते हैं और assume करते हैं कि यह distribution है। Practice में, token distributions अक्सर skewed होते हैं: most requests short हैं, लेकिन long requests की एक tail total cost का एक disproportionate share consume करती है। एक cost model जो average usage पर based है इन longer requests के cost impact को significantly underestimate कर सकता है, अक्सर एक factor of two या three से।

Latency एक similar pattern को follow करता है। Median latency pack के middle को reflect करता है, लेकिन SLA breaches आमतौर पर high end पर slower requests द्वारा caused होते हैं। यही कारण है कि p95 median से अधिक एक useful design target है। P95 वह latency value है जिसके नीचे 95% of requests complete होते हैं। केवल slowest 5% इसके ऊपर fall करते हैं।

Caching सबसे effective cost और latency lever है जब system prompt long और stable है। Prompt caching processed prompt prefix को cached tokens के लिए preserve करता है, इसलिए API subsequent requests पर उन्हें reprocess नहीं करता है। Savings cached prefix की length और कितनी बार यह reused है दोनों के साथ scale करते हैं। यदि, उदाहरण के लिए, cache reads को standard input token rate के 10% पर charge किया जाता है, तो एक long prefix जो many requests में reused है सबसे बड़ी effective savings produce करता है। Platform. claude. com/docs/en/about-claude/pricing पर most current cache read rate खोजें। Risk consistency है: यदि cached content को live state को reflect करने की आवश्यकता है, तो caching एक consistency window create करता है जो use case की requirements को violate कर सकता है। जो stored है वह prompt prefix है।

Reliability controls: जो हर model call में build करना है Reliability controls जो production Claude system में belong करते हैं different failure scenarios को address करते हैं और call stack के different layers पर sit करते हैं।

Transient error recovery with exponential backoff। जब एक model एक transient error return करता है, जैसे एक rate limit 429, timeout, या 5xx, system को progressively longer delays के साथ retry करना चाहिए attempts के बीच। यह एक brief hiccup को एक prolonged outage में turn करने से एक flood of retries को prevent करता है। Maximum number of attempts और total wait time को set करें based on कितना delay आपके use case को tolerate कर सकता है। Fallback chains। यदि primary model या endpoint unavailable है, system को automatically request को एक alternative जैसे एक different model tier या एक cached response में route करना चाहिए। इसे user को एक error raise नहीं करना चाहिए। Fallback behavior को आपके eval suite के part के रूप में test किया जाना चाहिए। Circuit breakers। एक circuit breaker एक downstream dependency पर error rate को measure करता है और trips करता है जब errors एक established threshold को exceed करते हैं। एक बार trip होने के बाद, requests immediately fail करते हैं rather than एक timeout के लिए wait करने के। यह एक degraded dependency को broader system को down ले जाने से prevent करता है।

Reliability controls को सही stage पर sit करना चाहिए effective होने के लिए: new attempts API call के close होने चाहिए, circuit breakers service boundary पर, और fallback chains orchestration layer में। उन्हें गलत layer में place करने का मतलब है system के गलत part को protect करना और सही part को exposed रखना।

Architecture type के आधार पर failure modes Agents उन tasks को handle करते हैं जो एक single model call में complete नहीं किए जा सकते: वे tools का उपयोग कर सकते हैं, results को observe कर सकते हैं, execution के mid में plans को adjust कर सकते हैं, और multi-step processes को complete कर सकते हैं जिनके लिए प्रत्येक step पर dynamic reasoning की आवश्यकता है। Claude Code एक production example है, और यह codebases को navigate करता है, tests चलाता है, fixes apply करता है, और एक workflow में iterate करता है जो एक single-turn architecture में impossible है। नीचे दिए गए controls govern करते हैं कि इन systems को reliably कैसे build करें।

Architecture | क्या पहले break करता है | Mitigation

Agent | Unbounded tool use और growing context। एक agent जो tools को call कर सकता है बिना budget constraints या turn limits के बिना cost और latency को उन तरीकों में run up करेगा जो invisible हैं जब तक एक single request budget ceiling को exceed नहीं करता। | Per-turn token budgets, maximum tool call counts, और explicit stopping criteria set करें। Tool set को minimum required में constrain करें। Agent की stopping behavior को eval करें, केवल output quality नहीं। RAG (retrieval-augmented generation) | Retrieval quality drift। Retrieval layer degrade होता है जब documents को index से add या remove किया जाता है बिना reindexing के, जब query और document representation fall out of alignment करते हैं, या जब index को एक schedule पर refresh किया जाता है जो live-state queries के लिए staleness create करता है। | Retrieval quality को eval loop में रखें। Retrieval precision और recall को system metrics के रूप में monitor करें, केवल output quality नहीं। Live-state queries को static knowledge queries से separate करें। Document processing pipeline (Evaluator-optimizer) | Low-confidence extractions के लिए कोई exception path नहीं। एक pipeline जो सभी documents को same flow के माध्यम से route करता है regardless of extraction confidence सभी documents पर edge cases पर wrong outputs produce करेगा same rate पर जो clean documents पर correct outputs produce करता है। | Extraction step में confidence scoring add करें। Low-confidence extractions को एक human review queue में route करें rather than downstream processing। Edge-cases और difficult documents को अपने eval set में include करें। Orchestrator-workers | Orchestrators और subagents के बीच failure boundaries blur होते हैं, traces fragment होते हैं, और एक dropped subagent silently synthesis पर fail कर सकता है। | Recoverable (subagent: retry या flag) बनाम unrecoverable (orchestrator) boundaries को define करें। सभी agents में एक shared trace ID create करें। Coverage को synthesis पर reconcile करें ताकि results submitted units के बराबर हों।

Cost · Complexity · Risk Cost: अपने proof-of-concept costs को production से match करने के लिए assume न करें। Architecture को commit करने से पहले costs को model करें, rather than first billing cycle के बाद। Complexity: Retries, fallback chains, और circuit breakers को एक system में add करना जो उनके लिए design नहीं किया गया था बहुत harder है। Reliability को start से build करें rather than first production incident के बाद scrambling करें। Risk: एक system जिसके पास कोई fallback नहीं है और कोई circuit breaker नहीं है एक point of failure है: primary model endpoint। जब वह endpoint peak load पर down जाता है, कोई recovery path नहीं है, पूरा user-facing workflow fail करता है instead of gracefully degrade करने के।

Model version pinning नोट: Model version pinning ऊपर दी गई table में हर architecture पर equally apply होता है। यह एक architecture choice नहीं है, एक operational discipline है। अपने configuration में model versions को pin करें, Anthropic model deprecation page को platform. claude. com/docs/en/about-claude/model-deprecations पर monitor करें, और एक version-update runbook maintain करें।

स्क्रीन 6: Demo cost profile जो production bill बन गया

Watch Out POC to Prod5 मिनट Demo cost profile जो production bill बन गया

यह गलती क्यों करना आसान है एक POC capability को prove करने के लिए build किया गया है। यह cost को model करने के लिए design नहीं किया गया है। Team एक small dataset के विरुद्ध build करता है, कुछ सौ requests चलाता है, और bill negligible है। एक POC bill low volume पर एक sample है, और यह production spend को reflect नहीं करता है।

Post-launch review से तीन quotes नीचे दिए गए quotes एक single team के post-deployment review से हैं, production में एक document triage system को move करने के 60 दिन बाद। प्रत्येक quote same गलती के एक different aspect को identify करता है।

Quote 1 "एक POC जो 10–50 requests per day चलता है अभी भी एक production cost estimate को inform कर सकता है, लेकिन केवल यदि numbers को expected production volume में appropriate error bounds के साथ scale किया जाता है। Raw demo costs को एक client को present करना बिना उस extrapolation के वह जगह है जहां risk lives।"

Quote 2 "हमने assume किया कि token distribution uniform होगा। यह नहीं था। Tail में long documents 80% of the total token spend को consume कर रहे थे।"

Quote 3 "जब endpoint ने peak पर एक 529 error return किया, पूरा workflow down चला गया। हमारे पास कोई fallback नहीं था क्योंकि हमने कभी test नहीं किया कि क्या होता है जब call fail करता है।"

क्या break हुआ और क्यों प्रत्येक quote एक distinct failure को name करता है, लेकिन वे order में compound करते हैं। Cost model गलत था क्योंकि यह incorrect volume पर build किया गया था। Token distribution assumption गलत था क्योंकि यह incorrect inputs पर build किया गया था। Reliability failure development में invisible था क्योंकि failure cases को कभी test नहीं किया गया था।

तीनों failures एक ही root cause को share करते हैं: POC को एक cost और reliability model के रूप में treat किया गया था, केवल एक capability demonstration नहीं। एक POC सवाल का जवाब देता है "क्या system यह कर सकता है", लेकिन यह "क्या यह scale पर करने के लिए cost करता है" या "क्या होता है जब एक dependency fail करता है" का जवाब नहीं देता है।

क्या Watch Out करें POC को तीन dimensions में simultaneously एक production model के रूप में treat किया गया था: cost, input distribution, और reliability। तीनों production-system properties हैं जिन्हें separately design किया जाना चाहिए, और एक POC उनमें से कोई भी establish नहीं करता है।

स्क्रीन 7: Cost & reliability calculator

Checkpoint POC to Prod5 मिनट Cost & reliability calculator Calculator का उपयोग करके configurations को explore करें। Model tier, prompt caching, max_tokens cap, और call-volume multiplier को adjust करें; readouts हर setting के लिए monthly cost और p95 latency दिखाते हैं। आपका goal: एक configuration खोजें जो cost ceiling और latency target दोनों को same time पर meet करता है। Calculator केवल values display करता है; यह आपकी exploration को score नहीं करता है। (Figures modeling के लिए representative हैं; publish time पर current rates को confirm करें।)

Scenario एक customer service agent 50,000 requests per month को process करता है। System prompt 5,000 tokens है और requests में stable है। Average user input 300 tokens है, average output 400 tokens है। Cost ceiling $800/month है। P95 latency target 3 seconds है।

Model tier

Haiku Sonnet Opus

Prompt caching

On

Max_tokens cap: 512

Call-volume multiplier: 1.

Est. monthly cost $720 Cost ceiling: $800/mo

p95 latency 2. 5s Latency target: ≤3s

Decision: एक बार जब आपके पास एक configuration है जो दोनों targets को meet करता है, कौन सा lever वहां पहुंचने के लिए सबसे अधिक काम करता है?

A. Opus में switch करना, क्योंकि सबसे capable model हमेशा safest है। B. Prompt caching को turn on करना, क्योंकि 5,000-token system prompt सभी 50,000 requests में stable है, इसलिए इसे cache करना dominant input-cost driver को cut करता है। C. Max_tokens cap को raise करना, क्योंकि अधिक headroom quality को improve करता है।

Submit करें अभी के लिए छोड़ें

स्क्रीन 8: Use-case sizing और feasibility

शिक्षण Sizing18 मिनट Use-case sizing और feasibility Production readiness checklist आपको बताता है कि viable होने के लिए एक system को क्या achieve करना चाहिए। यह model के outputs की quality और इसके चारों ओर system की reliability दोनों को cover करता है। Output quality को evals के माध्यम से validate किया जाता है, और system reliability को architecture controls जैसे retries, fallbacks, और circuit breakers के माध्यम से validate किया जाता है। दोनों bars को meet करना वह है जो production readiness का मतलब है।

Sizing आपको बताता है कि क्या एक specific business problem उस bar को meet कर सकता है, और कौन सी constraints design को govern करती हैं। Feasibility तीन states में से एक में fit करता है: feasible as scoped, feasible with constraints, और not feasible। State को correctly identify करना वह है जो एक scoping document को useful बनाता है।

एक use case को कैसे size करें एक use case को size करने का मतलब है कोई code लिखने से पहले एक cost model produce करना। Model को precise होने की आवश्यकता नहीं है, लेकिन यह accurate होना चाहिए enough ताकि architecture को budget के विरुद्ध validate किया जा सके और token distribution assumptions को formalize किए जाने से पहले surface किया जा सके।

चार inputs model को drive करते हैं: call volume, token budget per request, model tier, और sensitivity parameters।

Step 1: Call volume का अनुमान लगाएं। कितने requests per day या per month बनाए जाते हैं? यह number business requirement से आता है, developer की intuition से नहीं। एक customer service agent जो 1,000 conversations per day को handle करता है 1,000 Claude calls per day produce करता है, plus कोई भी multi-turn continuation calls। यह number business owner से get करें। एक sample dataset आपको एक accurate figure नहीं देगा। Step 2: Token budget per request को set करें। Token budget के दो components हैं: input tokens (system prompt, retrieved context, और user message) और output tokens (expected response length)। Distribution को model करें rather than केवल average। यदि document lengths widely vary करती हैं, cost model typical cases को account करना चाहिए साथ ही extremes। यदि system prompt long और stable है, prompt caching meaningfully input costs को reduce कर सकता है। Caching को explicit cache_control markers की आवश्यकता है request में। Cache writes को standard input से higher per-token cost incur करते हैं, इसलिए cost model को first use पर write cost को account करना चाहिए। Default cache TTL 5 minutes है; workloads जिनकी request frequency TTL से lower है consistent caching savings को realize नहीं करेंगे। Step 3: Monthly cost को project करें। Call volume को input token count से multiply करें input token rate पर। Separately, output token count को output token rate से multiply करें। फिर दोनों figures को add करें। Input और output tokens सभी model tiers पर different rates पर priced हैं। यदि prompt caching apply करता है, cached input tokens के लिए cache read rate का उपयोग करें, standard input rate नहीं। Finalize करने से पहले platform. claude. com/docs/en/about-claude/pricing पर current rates को verify करें। यदि applicable है तो caching savings को add करें। Result को production readiness checklist से cost ceiling के विरुद्ध compare करें। यदि projection ceiling को exceed करता है, architecture को code की एक line लिखने से पहले change करने की आवश्यकता है। यदि, उदाहरण के लिए, Batch API standard API pricing के relative में 50% price reduction provide करता है और 100,000 requests per batch तक support करता है, इसे किसी भी workload के लिए एक cost alternative के रूप में model करें जहां SLA asynchronous processing को permit करता है। Regulated workloads के लिए, verify करें कि क्या batch processing partner के BAA और compliance configuration के अंतर्गत covered है PHI या similarly governed data को इसके माध्यम से route करने से पहले। Platform. claude. com/docs/en/about-claude/pricing पर most current Batch API discount rate और batch size limit खोजें। Step 4: Sensitivity analysis चलाएं। क्या होता है cost को यदि call volume double होता है? क्या होता है यदि token distribution tail की ओर shift करता है? Sensitivity analysis आपको बताता है कि cost model कितना fragile है और कहां assumptions को business owner के साथ verify करने की आवश्यकता है design को commit करने से पहले।

एक use case को कैसे scope करें एक business requirement को एक scoped architecture में turn करने के लिए discovery sequence चार steps में चलता है। किसी भी step को skip करना एक commitment produce करता है जो business owner के साथ अगली conversation को survive नहीं करेगा।

Step 1: Business requirement को capability list में। System को क्या करने की आवश्यकता है? प्रत्येक capability को separately name करें। "Process insurance claims" एक goal है। Capabilities में claim document से structured fields को extract करना, policy database से policy coverage को look up करना, claim को appropriate adjuster queue में route करना based on claim type और value, और adjuster notification को draft करना शामिल हो सकता है। उन्हें separately identify करें ताकि आप प्रत्येक को appropriate owner को assign कर सकें। Step 2: Capability list को architecture sketch में। प्रत्येक capability के लिए, decide करें कि यह कहां belong करता है। कौन सी capabilities Claude own करता है? कौन सी existing systems को belong करती हैं? कौन सी एक human in the loop को require करती हैं? यह decomposition step है Module 1 से, एक specific use case पर applied। Step 3: Architecture sketch को boundary conditions में। उन conditions को state करें जिनके अंतर्गत architecture काम करता है और conditions जिनके अंतर्गत यह नहीं करता है। Feasibility एक verdict plus constraints है जो verdict को true बनाते हैं। एक architecture जो 20 pages तक के documents के लिए काम करता है लेकिन longer documents के लिए fail करता है एक boundary condition है जिसे document किया जाना चाहिए। Step 4: Boundary conditions को SOW में scope में। Statement of work boundary conditions को contain करता है। यह ensure करता है कि development team और business owner दोनों समझते हैं कि system को क्या handle करने के लिए design किया गया है और क्या explicitly out of scope है।

एक technical feasibility assessment को कैसे conduct करें एक feasibility assessment जो केवल "क्या Claude यह कर सकता है" पूछता है एक capability check है। चार AI properties आपको एक structured तरीका देते हैं identify करने के लिए कि design को कहां compensating controls की आवश्यकता होगी, और वे controls क्या होने चाहिए।

प्रत्येक property को select करें feasibility question को देखने के लिए जो यह raise करता है और जहां design compensate करता है।

Next Token Prediction Knowledge Working Memory Steerability

Feasibility question to ask: क्या यह task probabilistic generation को require करता है, या क्या यह specific values पर precision को require करता है? Classification, summarization, और drafting probabilistic tasks हैं जहां model excel करता है। Specific authoritative values (account numbers, policy dates, claim amounts) का extraction source of truth के विरुद्ध verification को require करता है। जहां design compensate करता है: Generator-verifier loops; extracted values पर code-based evals; quantitative data को retrieve करने के लिए tool calls।

Feasibility question to ask: क्या यह task उस information पर depend करता है जो rare, contested, recent, या domain-specific है उन तरीकों में जो training data में represent नहीं हो सकते हैं? यदि हां, design को knowledge को context window में bring करना चाहिए। Model को इसे supply करने पर rely न करें। जहां design compensate करता है: Stable knowledge के लिए retrieval-augmented generation; live-state data के लिए tool calls; contested claims पर uncertainty को flag करना।

Feasibility question to ask: क्या inputs comfortably context window में fit करते हैं, या क्या task को inputs को process करने की आवश्यकता है जो aggregate में window को exceed करते हैं? Long documents, multi-document tasks, और extended conversations सभी इस constraint को hit करते हैं। जहां design compensate करता है: Chunking strategies; progressive context loading; turns में summarization; context limit को exceed करने वाले inputs के लिए pipeline architecture।

Feasibility question to ask: क्या instructions specific, concrete, और verifiable हैं? Abstract या ambiguous instructions, long reasoning chains, और tasks जो precise numerical या logical computation को require करते हैं सभी वह जगहें हैं जहां model intent से drift कर सकता है। जहां design compensate करता है: Explicit output schemas के साथ system prompts; structured outputs; numerical precision के लिए code execution; evaluator-optimizer loops।

Feasibility verdicts एक बार scoping sequence और technical assessment complete होने के बाद, architecture एक feasibility assessment के लिए ready है। तीन possible outcomes हैं।

Verdict | इसका मतलब क्या है | क्या document करें

Feasible as scoped | चार AI properties में arguments Claude को प्रत्येक capability के लिए favor करते हैं। Cost model ceiling के अंदर है। Latency p95 SLA के अंदर है। कोई भी capability एक compensating control को require नहीं करता है जो architecture को change करता है। | Assumptions को clearly state करें। Feasible-as-scoped verdicts infeasible-with-constraints बन जाते हैं जब assumptions change करते हैं। Feasible with constraints | Design specific conditions के अंतर्गत काम करता है जिन्हें enforce किया जाना चाहिए। Document length एक threshold के अंतर्गत रहना चाहिए। Retrieval index को एक defined schedule पर refresh किया जाना चाहिए। एक human review gate outputs के लिए एक confidence threshold के ऊपर exist करना चाहिए। Constraints architecture का part हैं। | प्रत्येक constraint को explicitly document करें। प्रत्येक violated constraint के लिए, failure mode को identify करें। Development team को जानने की आवश्यकता है कि वे क्या build कर रहे हैं, केवल यह नहीं कि वे क्या build कर रहे हैं। Not feasible | कम से कम एक capability एक AI property limitation का सामना करता है जिसे scope और budget के अंदर compensate नहीं किया जा सकता है। Cost model ceiling को एक margin से exceed करता है जिसे model tier, caching, या architecture changes से close नहीं किया जा सकता है। एक not-feasible verdict एक correct assessment है जो engagement को एक अधिक expensive failure से बचाता है। | कौन सी constraint disqualifying है और क्यों state करें। जहां एक scope reduction verdict को change करेगा, इसे name करें और business owner को एक choice के साथ present करें।

Business value और ROI mapping: एक feasible design को एक justified investment में turn करना एक feasibility verdict business owner को बताता है कि system को budget के अंदर और constraints के अंतर्गत build किया जा सकता है। यह उन्हें नहीं बताता कि क्या इसे build करना worth है। Business value और ROI mapping को determine करना वह steps हैं जो दूसरे सवाल का जवाब देते हैं। वे scoped architecture को financial और operational outcomes से connect करते हैं जो business expect करता है, business owner पहले से ही use करने वाली terms में expressed: hours saved, error rates reduced, cycle time shortened, या revenue protected। Mapping एक technical design को एक decision में turn करता है जो एक budget holder defend कर सकता है।

Business case पांच main pillars पर rests करता है, और उन्हें identify करना ROI conversation को blueprint और business owner दोनों use करने वाली language में रखता है: efficiency (same work faster या cheaper में done), transformation (work जो पहले feasible नहीं था possible बन जाता है), productivity (same people से अधिक output), solution cost (system को run करने की cost), और performance SLAs (deployment को hold करने वाली service levels)। प्रत्येक ROI claim को उस pillar में map करें जिसे यह advance करता है ताकि value statement number और value का kind दोनों को capture करे।

Mechanism दो states के बीच एक comparison है। Baseline state वह है कि work आज कैसे किया जाता है, business को care करने वाली unit में measured। Projected state वह है कि work एक बार Claude को workflow में होने के बाद कैसे किया जाता है, same unit में measured। Value दोनों states के बीच difference है, minus system को run करने की cost। Cost figure directly sizing के दौरान produce किए गए cost model से आता है, इसलिए ROI calculation work को reuse करता है जो आप पहले से ही किए हैं rather than starting over।

Mapping चार steps में build किया जाता है, और प्रत्येक step को एक number में ground किया जाना चाहिए जो business owner को recognize करेगा:

Step 1: Baseline को एक business unit में name करें। आज task कैसे perform किया जाता है से start करें और इसे उस unit में measure करें जो business पहले से ही track करता है। एक claims review workflow के लिए वह analyst hours per claim या average days to resolution है। Baseline को business owner के own operational data से आना चाहिए, क्योंकि हर बाद का number इसके विरुद्ध compared है। Intuition से pulled एक baseline एक ROI figure produce करता है जो कोई भी finance team accept नहीं करेगा। Step 2: Post-deployment state को same unit में predict करें। Estimate करें कि same task कितनी अच्छी तरह perform करता है एक बार Claude workflow में है, baseline के identical unit में measured। जहां feasibility verdict human review को require करता है, projection को उस cost को include करना चाहिए। Low-confidence output को एक reviewer में route करना labor को reduce करता है, लेकिन यह पूरी तरह से eliminate नहीं करता है। जब design human-in-the-loop review को specify करता है, full automation को project करना value को overstate करता है और एक number produce करता है जो operations reject करेगा। Step 3: Sizing model से run cost को subtract करें। Sizing के दौरान produce किए गए projected monthly cost को लें और इसे new state की recurring cost के रूप में treat करें। Deployment की value operational gain से Step 2 minus यह run cost है। यह step केवल recurring run cost को isolate करता है। Build cost को separately treat किया जाता है और Step 4 में payback period calculation में accounted है। Sizing output को ROI calculation में include करना दोनों analyses को consistent रखता है। Token budget या model tier में एक change फिर cost ceiling और value case दोनों को update करता है together। Step 4: Payback period और sensitivity को state करें। Result को एक payback period के रूप में express करें, जो वह time है जो accumulated operational gain को build cost और run cost को cover करने के लिए लेता है। फिर state करें कि वह period कैसे move करता है यदि volume assumptions या gain-per-task assumptions गलत हैं। एक payback period को quote करना बिना sensitivity analysis के एक single optimistic scenario पर based decisions को invite करता है, एक जो अक्सर fail करता है जब business case launch के बाद real volumes को meet करता है।

Mapping का output एक short value statement है जो Architect business owner को feasibility verdict के साथ hand करता है। यह "क्या हम यह build कर सकते हैं? " और "क्या यह build करना worth है? " के जवाबों को numbers में tie करता है जो business पहले से ही own करता है। दोनों artifacts statement of work में together travel करते हैं।

ये सबसे common ROI map errors हैं, क्या उन्हें cause करता है, और जहां वे appear करते हैं।

Risk | क्यों यह होता है और जहां यह दिखता है

Baseline को estimated rather than measured किया जाता है। | जब business owner के पास clean operational data नहीं है, baseline को intuition से fill किया जाता है। यह apparent gain को misleading बनाता है। Error तब तक hidden रहता है जब तक finance team business case review के दौरान baseline number के source के लिए नहीं पूछता। इस point पर पूरे case को rebuild किया जाना चाहिए। Projection assume करता है full automation जब design human review को require करता है। | एक feasibility verdict जो एक human review gate को require करता है labor को reduced करता है, fully eliminated नहीं। Value case अक्सर misleadingly इसे eliminated के रूप में model करता है। Gap launch के बाद first operational period में surface करता है, जब actual analyst hours उतनी दूर fall नहीं करते जितनी business case ने promise किया था। Run cost को average से rather than sizing distribution से लिया जाता है। | Average token cost को reuse करना instead of sizing model से distribution से recurring cost को understate करता है, जो net value को overstate करता है। यह heavy-tailed inputs वाले workflows पर concentrate करता है, जहां small fraction of large requests most of the cost को drive करता है।

Cost · Complexity · Risk Cost: Average token counts पर based sizing जब कुछ requests दूसरों से बहुत बड़े हों तो cost को underestimate करेगा। यह गलत करना architecture को commit करने के बाद renegotiate करने का मतलब है। Complexity: एक feasibility assessment जो चार AI properties में से किसी को skip करता है एक constraint को miss करने का risk रखता है जो design को change करता है। Working memory सबसे overlooked है, क्योंकि यह development पर small, clean inputs के दौरान rarely show up करता है, लेकिन यह production में surface करेगा। Risk: एक feasible-with-constraints verdict जिसे document नहीं किया गया है एक infeasible system बन जाता है जब constraints production में violated होते हैं। Constraints design का part हैं और architecture को carry करते हैं same weight।

स्क्रीन 9: Scoping call जिसने constraints को skip किया

Watch Out Sizing5 मिनट Scoping call जिसने constraints को skip किया

Demo trap जब एक business owner एक use case के बारे में excited है और initial demo काम करता है, feasibility को confirm करना volume और SLA constraints को gather करने से पहले efficient path लगता है। Capability वहां है, prototype काम कर रहा है, और constraints के बारे में पूछने के लिए slow down करना reasons को look करने जैसा लग सकता है कि नहीं कहने के लिए। Problem यह है कि "technically feasible" meaningless है यदि constraints expected scale पर apply नहीं किए जाते हैं।

एक scoping conversation: Commitment बहुत जल्दी किया गया निम्नलिखित एक pattern का extract है जो discovery calls में surface करता है जब capability question को constraint questions से पहले answer किया जाता है।

Partner: "हमें एक document review assistant की आवश्यकता है जो हमारे legal contracts को process कर सकता है और non-standard clauses को flag कर सकता है।" Architect: "हम यह कर सकते हैं। Model contracts को read करने और clause patterns को identify करने में अच्छा है। मुझे एक feasibility write-up put together करने दें।" Partner: "बहुत अच्छा, build को कितना समय लगेगा? " Architect: "Initial version के लिए छह हफ्ते।" [Build में दो हफ्ते] Partner: "मुझे mention करना चाहिए: हम एक दिन में लगभग 800 contracts को process करते हैं। कुछ framework agreements हैं जो 300 pages तक चलते हैं। और हमें 30 seconds में results की आवश्यकता है।"

क्या गलत हुआ Feasibility verdict को तीन essential constraints को gather करने से पहले issue किया गया था: call volume (800 per day), input size (up to 300 pages), और latency requirement (30 seconds)।

क्या एक lengthy contract context window में fit करता है model tier के चयन पर depend करता है। 1 million token context window वाले models एक 300-page contract को chunking के बिना handle कर सकते हैं; एक 200k token window वाले model को longest documents के लिए एक chunking strategy को require कर सकता है। Context window capacity इसलिए model tier decision का part है, एक settled assumption नहीं।

800 requests per day पर, 30-second latency requirement उतना limiting नहीं है जितना लगता है। यह एक request हर 108 seconds में average करता है। Sequential processing उस volume पर viable है बिना requests को parallel में run किए। उस request rate पर, volume latency को pressure नहीं करता है, cost को pressure करता है। Latency task complexity, model size, और output length द्वारा driven है। वे variables हैं जिनके चारों ओर model tier decision को build करने की आवश्यकता है।

Architect ने capability को एक necessary first step के रूप में confirm किया, लेकिन यह भी last step था, जिसका मतलब है commitment को design possible होने से पहले किया गया था।

क्या Watch Out करें Capability question को constraint questions को भी ask किए जाने से पहले answer किया गया था। Volume, latency, और input-size constraints feasibility verdict के लिए inputs हैं। Verdict केवल उतना ही sound है जितना constraints को gather किया गया है।

स्क्रीन 10: Feasibility call को justify करें

Checkpoint Sizing5 मिनट Feasibility call को justify करें

प्रत्येक scenario के लिए, वह option चुनें जो correct feasibility verdict और इसके पीछे single load-bearing constraint दोनों को name करता है। Verdict अकेला काफी नहीं है, constraint वह है जो verdict को defensible बनाता है।

SCENARIO 1 OF 3 एक professional-services firm एक research assistant चाहता है जो 10–40 page industry reports को summarize करता है और client briefings को draft करता है। 50 reports per week, briefings within 24 hours, $500/month ceiling।

A. Not feasible, reports context window के लिए बहुत लंबे हैं। B. Feasible as scoped, input size context window में fit करता है, और volume, latency, और cost सभी Sonnet tier पर range में हैं; कोई भी AI property एक disqualifying constraint को present नहीं करता है। C. Feasible with constraints, हर briefing पर एक human review gate की आवश्यकता है।

SCENARIO 2 OF 3 एक logistics company एक delay predictor चाहता है जो unstructured carrier emails को read करता है, delay reasons और new ETAs को extract करता है, और उन्हें order-management system में write करता है। 5,000 emails/day, under 10 seconds, $1,000/month।

A. Feasible as scoped, Haiku with caching volume और latency को handle करता है। B. Not feasible, volume बहुत high है। C. Feasible with constraints, load-bearing constraint एक transactional write पर extraction accuracy है, इसलिए यह extraction accuracy पर एक code-based eval plus low-confidence extractions पर एक human review gate को require करता है system of record में write करने से पहले।

SCENARIO 3 OF 3 एक financial-services firm real-time trading recommendations चाहता है current market conditions plus अपने proprietary model से, 2 seconds में delivered।

A. Not feasible as described, load-bearing constraint live-state knowledge gap है: real-time market data को एक live feed में एक tool call को require करता है, और क्या वह round trip 2-second budget के अंदर fit करता है feasibility verdict को issue किए जाने से पहले validate किया जाना चाहिए। B. Feasible as scoped, Claude पहले से ही current market conditions को जानता है। C. Feasible with constraints, बस एक human gate add करें।

अभी के लिए छोड़ें

स्क्रीन 11: Enterprise integration patterns: identity, auth, data, और observability

शिक्षण Integration14 मिनट Enterprise integration patterns: identity, auth, data, और observability

Compliance constraints किसी भी अन्य निर्णय से पहले entry point options को eliminate करते हैं, इसलिए entry point selection पहले आता है। शेष पांच layers govern करते हैं कि integration वहां से कैसे build किया जाता है।

Sizing आपको बताता है कि system को क्या करने की आवश्यकता है और क्या यह constraints के अंदर कर सकता है। Integration patterns आपको बताते हैं कि यह enterprise stack से कैसे connect करता है। वह connection पांच layers है, और प्रत्येक layer के पास एक architectural decision है जो Architect को belong करता है, implementation team को नहीं।

Entry point selection: कौन सा integration use करें और कब किसी भी integration में पहला decision कौन सा entry point system connect करेगा। Compliance constraints इस first stage पर options को eliminate करते हैं, किसी भी अन्य architectural decisions से पहले। Table पांच available entry points को cover करता है, कब प्रत्येक apply करता है, और प्रत्येक flexibility या maintenance में क्या cost करता है।

Entry points: कौन सा integration use करें और कब Entry Point | इसे use करें जब | आप क्या trade off करते हैं

Direct API | आपको requests को कैसे build किया जाता है, responses को कैसे handle किया जाता है, और errors को कैसे manage किया जाता है पर full control की आवश्यकता है। System का हर part आपका own code है। | आप retry logic, streaming, tool orchestration, और error handling को build और maintain करने के लिए responsible हैं। Implementation effort SDK path से higher है। SDK (Python/TypeScript) | आप एक convenience layer चाहते हैं जो basics को handle करता है बिना design को कैसे करें पर control को give up किए। SDK HTTP layer को handle करता है और typed interfaces देता है जबकि orchestration को आपके code में छोड़ता है। | Raw API से कम granular control। SDK version upgrades कभी-कभी behavior को उन तरीकों में change कर सकते हैं जिन्हें deploy करने से पहले code review की आवश्यकता है। Claude Code | Primary user एक developer है, और task में code को write, review, या navigate करना involve करता है। Multiple users या customer-facing deployments वाले products के लिए appropriate नहीं है। | Multiple users या customer-facing deployments के लिए design नहीं किया गया। Claude Code developer workflows के लिए design किया गया है, multi-tenant products के लिए backend नहीं; product integrations को Claude API, client SDK, या Agent SDK का उपयोग करना चाहिए। Agent SDK | Claude को आपके own product के अंदर multiple turns में act करने की आवश्यकता है, आपके application को surrounding workflow को control करने के साथ। Agent SDK एक managed loop चलाता है, iteration, tool execution, और termination को handle करता है, इसलिए आपकी team को उस infrastructure को build करने की आवश्यकता नहीं है। Python और TypeScript में available। जब एक single request और response sufficient है, या जब task को tools पर multi-turn reasoning को require नहीं करता है तब सही choice नहीं है। | Managed loop fine-grained control को give up करता है प्रत्येक iteration step पर। यदि आपके use case को turns के बीच custom logic को require करता है, एक raw API loop आपको उस control को देता है managed loop को build और maintain करने की cost पर। MCP (Model Context Protocol) | आप Claude को existing tools या internal services से connect कर रहे हैं और एक standard तरीका चाहते हैं उन connections को manage करने के लिए बिना उन्हें अपने orchestration logic में mix किए। MCP एक protocol provide करता है tool integration के लिए जो integration layer को orchestration layer से separate रखता है। | MCP एक protocol layer add करता है Claude और आपके tools के बीच, जो tool calls को debug करना direct function calls से अधिक complex बनाता है।

Security और compliance constraints integration layer पर Module 1 ने establish किया कि regulatory और policy constraints, HIPAA, GDPR, और FedRAMP जैसे laws, attorney-client privilege, और data-residency requirements entry point options को eliminate करते हैं किसी भी अन्य decisions से पहले। यह section same logic को एक level deeper apply करता है: कौन सा integration pattern उस entry point पर, appropriate architectural choices, internal policies, contractual terms, और अन्य items के साथ, help satisfy या otherwise mitigate concerns को constraint के बारे में।

नीचे दी गई table entry point selection से full integration design तक उस analysis को extend करता है।

नोट: यह legal guidance के रूप में नहीं देखा जाना चाहिए - अपने own legal और compliance team के साथ काम करें controls को implement करने के लिए जो आपके organization की needs को meet करते हैं।

Constraint-to-integration matrix Constraint | जहां यह चलता है (Route) | यह कैसे connect करता है (Integration pattern) | यह किसे trust करता है (Identity) | यह क्या handle करता है (Data handling) | यह क्या logs करता है (Observability)

Attorney-client privilege* | API firm के own application के पीछे चलता है, एक firm-approved gateway के माध्यम से जो हर request को logs करता है। | Gateway user और Claude के बीच बैठता है। यह API key को hold करता है, enforce करता है कि कौन क्या access कर सकता है, और एक audit log produce करता है जो firm own करता है। | Users SSO के माध्यम से log in करते हैं। User identity और permissions server द्वारा assign किए जाते हैं और user द्वारा claim नहीं किए जा सकते। | सभी privileged content gateway के माध्यम से flow करता है, जो official record serve करता है। | हर request और response gateway पर logged है और firm policy के अनुसार retained है। HIPAA (PHI handling) | API एक configuration पर चलता है जो एक Business Associate Agreement (BAA) द्वारा covered है। HIPAA-eligible cloud paths में Claude API directly (Anthropic से एक signed BAA के साथ), AWS Bedrock, और Google Vertex AI शामिल हैं। BAA को specific configuration को cover करना चाहिए use में, केवल provider को generally नहीं। | Cloud provider integration को mediate करता है। BAA को specific configuration को cover करना चाहिए use में, केवल provider को generally नहीं। | Users partner के authentication system के माध्यम से verified होते हैं। PHI को access करना minimum necessary तक limited है HIPAA के अंतर्गत। | PHI को केवल task को require करने वाले में strip down किया जाता है। जहां possible है, reference IDs को full data fields की जगह use किया जाता है। | Logs request, model version, user identity, और data scope को capture करते हैं, और HIPAA requirements के अनुसार retained हैं। GDPR और data residency | Model execution को एक approved region में confined किया जाता है। यह एक cloud route (AWS Bedrock या Google Vertex AI) के माध्यम से या Claude API directly को inference_geo parameter का उपयोग करके achieve किया जा सकता है, जो currently "us" और "global" को values के रूप में support करता है। नोट करें कि inference_geo EU pinning को support नहीं करता है – जबकि GDPR compliance EU residency को require नहीं करता है, cross-border transfers lawful हैं एक valid transfer mechanism के साथ, जहां एक deployment को EU data residency requirement है, एक cloud route rather than direct API को use किया जाना चाहिए। जहां Claude API use किया जाता है, DPA Anthropic के साथ directly है। जहां एक cloud route use किया जाता है, DPA terms cloud contract से inherited हैं। नोट: Microsoft Foundry EU data residency को Coming 2026 के रूप में list किया गया है confirmed timeline के बिना publish पर। यदि deployment Foundry को reference करता है और EU data pinning को require करता है, उस route को commit करने से पहले current availability को verify करें। | Execution region को integration layer पर locked किया जाता है और हर request पर checked किया जाता है। Data approved region को leave नहीं करता है। | Users approved data region के अंदर verified होते हैं और GDPR requirements के अनुसार handled होते हैं। | Personal data को केवल pinned region के अंदर process किया जाता है। Borders में data को move करने को एक documented legal basis को require करता है और केवल explicitly justified होने पर build किया जाता है। | Logs record करते हैं कि किसने क्या data access किया, processing के लिए legal basis, और जब यह delete किया जाएगा। FedRAMP और government | Integration एक cloud configuration पर चलता है जो correct impact level पर required FedRAMP authorization को hold करता है। FedRAMP-eligible paths Claude for Government, AWS Bedrock GovCloud, और Google Vertex Assured Workloads हैं। Claude Enterprise direct API पर FedRAMP authorized नहीं है और इन paths के लिए substitute नहीं किया जा सकता है। | Architecture को specific cloud और configuration तक limited है जो authorization को carry करता है। उस boundary के बाहर कुछ भी allowed नहीं है। | Users agency के approved identity provider के माध्यम से authenticate करते हैं। Access agency के role और policy layer द्वारा controlled है। | Data agency के classification rules के अनुसार handled है। Controlled unclassified information authorized boundary के अंदर stay करता है। | Logs agency की continuous monitoring requirements को meet करते हैं। Internal data-residency policy | Integration उस cloud provider पर चलता है जिसे partner के organization ने पहले से approve किया है। सही route वह है जिसे CIO ने clear किया है, convenience की परवाह किए बिना। | Architecture को constrain किया जाता है जो procurement ने approve किया है, engineering preference की परवाह किए बिना। | Users partner के standard SSO के माध्यम से log in करते हैं। Roles partner के existing policy के अनुसार assign किए जाते हैं। | Data handling partner के existing classification scheme के अनुसार। Claude layer existing controls को inherit करता है और नए ones को introduce नहीं करता है। | Logs partner के existing logging infrastructure में feed करते हैं rather than एक separate system।

  • Privilege preservation appropriate contractual terms, retention settings, internal policies, और अन्य items पर depend करता है। इस chart में information privilege waiver concerns को mitigate करने में help करने के लिए intended है।

किसी भी integration design को शुरू करने से पहले, constraints के माध्यम से काम करें order में: governing regulation या policy को identify करें, determine करें कि कौन से entry point और route अभी भी available हैं, integration pattern को choose करें जो fit करता है, और identity, data handling, और observability requirements को document करें जो follow करते हैं। किसी भी step को skip करना कुछ ऐसा build करने का risk रखता है जो technically काम करता है लेकिन एक legal या security review को fail करता है।

हर enterprise Claude integration को नीचे defined layers में decisions की आवश्यकता है। Compliance पहले आता है, regulations जैसे HIPAA, GDPR, और FedRAMP से constraints, और policies जैसे data residency और attorney-client privilege, routes और entry points को eliminate करते हैं किसी भी अन्य decisions से पहले। शेष चार layers उन options के अंदर operate करते हैं जो उस filter को survive करते हैं। नोट करें कि ये sections ZDR use cases को address नहीं करते हैं।

Layer | Architectural decision | क्या break होता है जब यह गलत है

Compliance और regulated-industry constraints | कौन से delivery routes और entry points governing constraint को survive करते हैं? BAA coverage, FedRAMP authorization, data-residency pinning, और approved-vendor lists प्रत्येक किसी भी अन्य decisions से पहले options को eliminate करते हैं। | Integration एक route पर build किया जाता है जो next legal या security review को fail करता है। उस point पर redesign की cost पहले से invested time plus एक scratch से एक नया architecture है। Identity और SSO | User identity boundary कहां sit करता है Claude integration point के relative में? Context में एक Claude call में user कौन है, और वह identity safely prompt में कैसे pass होता है? | जब user identity correctly prompt में pass नहीं होता है, Claude अपने responses को scope नहीं कर सकता है जो user को authorized है देखने के लिए। Identity raw field के रूप में user message में pass किया गया है manipulable है। Server-side injection वह risk को remove करता है। Authorization और policy | कौन सी capabilities इस user या role के पास हैं? कौन सा data वे access कर सकते हैं? Authorization model जो आपके existing systems को govern करता है Claude layer को भी govern करना चाहिए। | एक Claude integration जो underlying system के authorization model को bypass करता है users को data को access देता है जिसे वे authorized नहीं हैं, एक path के माध्यम से जो access policy को enforce करने के लिए design नहीं किया गया था। Data handling और PII | कौन सा data context window में जाता है? Sensitive fields जो directly user message या system prompt में pass किए जाते हैं API request का part बन जाते हैं। Anthropic default द्वारा conversation content को retain नहीं करता है; केवल जो technically API और feature के लिए necessary है retained है। Request अभी भी wire को cross करता है, specific retention carve-outs कुछ model classes के लिए exist करते हैं, और कोई भी logging partner का own application layer perform करता है इसे capture करेगा। Architecture को decide करना चाहिए कि कौन से fields context window में necessary हैं और कौन से केवल जब needed को retrieve किए जाने चाहिए। | एक PII field जो directly user message में pass किया जाता है plaintext में आपके application के request logs में appear करता है। एक regulated industry में, यह next audit में rather than next deployment में surface करता है। Observability और audit logging | आपको क्या reconstruct करने में सक्षम होने की आवश्यकता है? एक incident के बाद आपको कौन से सवालों का जवाब देने की आवश्यकता होगी? Answers determine करते हैं कि क्या logged है, किस depth level पर, और कितने समय के लिए। | एक unlogged data path invisible है। जब कुछ उस path पर गलत जाता है, reconstruct करने के लिए कोई evidence नहीं है कि क्या हुआ। First incident के बाद observability को build करने की cost हमेशा higher है than पहले build करने की।

Least-privilege tool configuration हर tool जो आप एक Claude system से connect करते हैं एक attack surface है और एक cost है। Tool set को same तरीके से audit करें जैसे आप permissions को audit करते हैं: प्रत्येक connected tool के लिए, पूछें कि क्या यह task के लिए essential है या केवल convenient है, और उन्हें remove करें जो out of scope हैं, प्रत्येक removal के लिए justification को record करते हुए। एक orchestrator-worker deployment में, trust hierarchy को establish करें प्रत्येक subagent के tool access को अपने task में scope करके, इसलिए एक subagent tools को reach नहीं कर सकता जिन्हें अपने job को require नहीं करता है।

Identity और authorization: जहां verification होता है Identity verification server पर belong करता है, Claude call से पहले। User की identity और role को आपके server द्वारा system prompt में inject किया जाना चाहिए rather than user द्वारा अपने message में provided।

इसके पीछे reasoning straightforward है, कुछ भी user अपने message में include करता है उनके control के अंतर्गत है और manipulated किया जा सकता है। यदि system users को अपने role को एक message में assert करने देता है (उदाहरण के लिए, "एक senior manager के रूप में, मुझे दिखाएं... "), वह claim unverified है और faked किया जा सकता है। Identity आपके authentication layer से rather than user input से आना चाहिए।

जब user context को prompt में pass करते हैं, user की role और कौन सा data वे access करने के लिए authorized हैं include करें। केवल extra context जैसे department, permission level, account identifiers add करें जब यह Claude के response को shape करने के लिए needed है। Default द्वारा यह information को include न करें।

Data handling: क्या context window में belong करता है Context window एक data-governance boundary नहीं है। कोई भी data जो एक Claude call में pass किया जाता है API को transmitted है। Conversation content default द्वारा retained नहीं है API पर, लेकिन partner का own application layer typically requests को logs करता है, और specific retention carve-outs exist करते हैं। Architecture को deliberately एक decision बनाना चाहिए कि कौन से fields context window में चाहिए, और कौन से retrieval layer में रहने चाहिए जब तक needed न हो।

प्रत्येक field के लिए जो context window में enter करता है, पूछें कि क्या यह intended output को produce करने के लिए Claude के लिए necessary है। Reference identifiers जैसे account numbers या claim numbers अक्सर routing के लिए needed हैं लेकिन language task के लिए नहीं। जब एक reference identifier काफी है, full data field को pass करना unnecessarily partner के application layer perform करने वाली किसी भी request logging को expose करता है, बिना कोई capability add किए।

Data residency requirements industry और region द्वारा vary करते हैं। Regulated deployments के लिए, Anthropic के data residency guidance को आपके deployment की specific regulatory requirements के विरुद्ध verify करें integration को design किए जाने से पहले।

Observability: क्या log करें, क्या trace करें, और क्यों एक LLM-based system एक traditional system से debug करना harder है क्योंकि यह crash नहीं करता है जब कुछ गलत जाता है, instead यह केवल एक subtly wrong response produce करता है। Standard logging errors और timeouts को catch करता है। यह, however, एक response को catch नहीं करता है जो quietly incorrect है एक तरीके में जो real business consequences रखता है।

एक production Claude system को चार चीजों को log करना चाहिए:

Request: model version, input token count, prompt identifier Response: output token count, latency, stop reason Context: user role, session ID, क्या caching apply किया गया था Outcome: क्या downstream system ने output को accept किया और कोई भी rejection signals

Security organizations increasingly observability को precondition के रूप में treat करते हैं agents को enable करने के लिए, क्योंकि एक trustworthy audit trail के बिना, एक autonomous system को act करने के लिए approve नहीं किया जाता है। उस standard के लिए design करें। एक design-review checklist item के रूप में, verify करें कि कौन से agentic actions आपके chosen surfaces में audit logs में recorded हैं। Coverage surfaces द्वारा vary करता है, और एक action जो taken है लेकिन logged नहीं है, एक security reviewer को, एक action है जिसे allow नहीं किया जा सकता है।

Cost · Complexity · Risk Cost: एक integration बिना PII redaction के API call से पहले sensitive fields को expose करता है partner के application-layer request logging को हर request पर। Retroactive redaction की cost एक log history में जो कभी इसे support करने के लिए design नहीं किया गया था एक production Claude system में सबसे expensive data handling fix है। Complexity: एक observability layer जो first production incident के बाद add किया जाता है root cause को एक system से reconstruct किया जाना चाहिए जो question को answer करने के लिए set up नहीं किया गया था incident को surface किया। Build logging को answer करने के लिए questions को आपको answer करने की आवश्यकता होगी, पहले आपको उन्हें पूछने की आवश्यकता है। Risk: एक multi-tenant system जो एक shared API key पर चल रहा है एक rate limit breach को tenant को attribute करने का कोई तरीका नहीं है जिसने इसे cause किया। जब org-level limit peak load पर trip करता है, spike visible है लेकिन source नहीं है, और हर tenant impact को absorb करता है। Separate API keys per tenant production multi-tenant deployment में attribution और isolation के लिए required हैं।

स्क्रीन 12: PII field जो सीधे prompt में चला गया

Watch Out Integration5 मिनट PII field जो सीधे prompt में चला गया

Demo को skip करने के लिए security के लिए trap जब goal एक demo को quickly काम करना है, एक working Claude integration के लिए fastest path वह data को pass करना है जो आपके पास directly prompt में है। एक PII redaction layer को add करना, एक server-side identity injection को build करना, और observability stack को instrument करना सभी time add करते हैं। वे भी कोई visible capability add नहीं करते, जैसा कि system उनके बिना काम करता है। Skip करने की cost पहले audit तक appear नहीं करता है।

एक regulated-industry deployment से एक trace excerpt नीचे दिया गया trace एक healthcare-adjacent deployment में एक field pattern का composite है। Team ने एक patient intake summarization tool बनाया। Summarization correctly काम किया। Data handling नहीं।

API request: application layer के request logs में captured Model: claude-sonnet-4-6 System: "आप एक clinical intake summarizer हैं। Intake form से key presenting concerns, medications, और allergies को extract करें।" User: "Patient: Jane Doe, DOB: 1978-04-12, SSN: 123-45-6789, Insurance ID: BCB-88712. Chief complaint: chest pain, onset 3 days ago... "

SSN और Insurance ID user message में हैं और इसलिए API request के साथ travel करते हैं, किसी भी application-layer request logging को expose किए गए, despite summary को produce करने के लिए needed नहीं होने के। System prompt presenting concerns, medications, और allergies के लिए पूछता है। उनमें से कोई भी patient के SSN या Insurance ID को context window में होने की आवश्यकता नहीं है।

क्या break हुआ Tool design के अनुसार काम किया, लेकिन data handling गलत था। कई fields जो HIPAA के अंतर्गत Protected Health Information (PHI) को qualify करते हैं, patient का name, date of birth, SSN, Insurance ID, chief complaint, medications, और allergies, API call में pass किए गए थे और application layer के request logs में plaintext में captured किए गए थे। ये language task के लिए necessary नहीं थे। जब deployment को production certification से पहले review किया गया, application layer के request logs में patient SSNs के साथ thousands of entries user message field में contained थे।

Fix एक data architecture change था: एक server-side redaction step जो non-essential PII fields को Claude call से पहले strip करता है, और एक retrieval function जो केवल fields को supply करता है जो language task को Claude perform कर रहा है। दोनों को original design में होना चाहिए था।

क्या Watch Out करें Data handling architecture को convenient के चारों ओर design किया गया था pass करने के लिए, necessary के चारों ओर नहीं। Necessity सही filter है: यदि field language task के लिए required नहीं है Claude perform कर रहा है, यह context window में नहीं होना चाहिए।

स्क्रीन 13: Integration diagram को critique करें

Checkpoint Integration5 मिनट Integration diagram को critique करें

नीचे दिया गया diagram एक multi-tenant SaaS company पर एक customer-service agent के लिए एक Claude deployment दिखाता है। Labeled components और connections को review करें, और हर एक को select करें जो एक integration problem को represent करता है।

सभी problems को select करें। Sound components को unselected छोड़ें।

Problem: Claude Code को customer-facing chat product के लिए backend के रूप में use किया गया

↓↓↓

Problem: सभी tenants के लिए shared API key use किया गया

Problem: Capability check "मैं एक premium customer हूं" user message में based

Problem: Account number और email user message में pass किए गए और request logs में appear करते हैं

↓↓↓

Sound: API से पहले server-side authentication layer

Sound: Tenant data storage layer पर per tenant isolated

Problem: Claude response को downstream CRM में pass किया गया बिना integration layer पर logging के

उत्तर जांचें अभी के लिए छोड़ें

स्क्रीन 14: A/B testing और observability at scale

शिक्षण A/B & Obs17 मिनट A/B testing और observability at scale

Integration patterns Claude को enterprise stack में get करते हैं। सवाल जो follow करता है वह है क्या यह perform कर रहा है जैसा कि यह should एक बार वहां है। Observability monitoring question को answer करता है। Structured A/B testing improvement question को answer करता है। दोनों के बिना, आप या तो blind fly कर रहे हैं या changes बना रहे हैं जिन्हें आप measure नहीं कर सकते।

Live Claude systems के लिए structured A/B testing एक Claude system के लिए एक A/B test किसी भी experiment के समान structure को follow करता है: एक hypothesis, एक treatment group, एक control group, एक metric, और एक sample size large enough ताकि result statistically meaningful हो। Traditional software A/B testing से difference यह है कि LLM outputs probabilistic हैं, जो results को noisier बनाता है और interaction effects को harder to control बनाता है।

Hypothesis को specific और testable होना चाहिए। "नया prompt बेहतर है" दोनों पर fail करता है: यह कोई treatment, कोई metric, और कोई threshold name नहीं करता है। एक usable hypothesis ऐसा reads: "Summarize करने के लिए instruction को replace करना extract करने के लिए instruction के साथ three most important action items को task success rate को कम से कम 5% से increase करेगा बिना response latency p95 को degrade किए।" यह treatment, metric, success के लिए threshold, और constraint को name करता है।

Component | इसे क्या require करता है | जब यह missing है तो क्या गलत जाता है

Hypothesis | एक specific, falsifiable statement naming treatment, expected direction of primary metric, और secondary metrics पर कोई भी constraints। | एक hypothesis के बिना, कोई भी result को एक win के रूप में interpret किया जा सकता है। आप हमेशा एक metric find कर सकते हैं जो सही direction में move किया है यदि आप fact के बाद काफी देखते हैं। Treatment और control assignment | Random assignment of requests को treatment (new version) या control (current version) में। Assignment को एक given user या session के लिए consistent होना चाहिए contamination से बचने के लिए। | Non-random assignment का मतलब है groups comparable नहीं हैं। यदि treatment group को happen करता है अधिक complex queries को receive करने के लिए, एक apparent win input distribution का एक artifact हो सकता है। Primary metric | एक single metric defined experiment को run करने से पहले। Task success rate, cost per completion, latency p95, या user satisfaction proxy। Experiment को run करने के बाद metric को choose करना outcome-shopping है। | एक unspecified primary metric एक experiment को एक retrospective correlation में turn करता है, जो एक decision के लिए एक much weaker basis है। Sample size | Minimum detectable effect, baseline metric value, और required confidence level से calculated। LLM systems के लिए, outputs में variance deterministic systems से higher है, जिसका मतलब है required sample size larger है। | एक underpowered experiment results produce करता है जो real effect को noise से distinguish नहीं कर सकते। एक team जो तब तक run करता है जब तक वे नहीं देखते कि वे चाहते हैं find करेंगे जो वे looking के लिए थे, चाहे यह real हो या नहीं।

Results को overclaiming के बिना पढ़ना Statistical significance का मतलब है result को chance द्वारा occur होने की unlikely है given sample size। क्या यह matter करने के लिए large enough है एक separate question है। एक change statistically significant हो सकता है लेकिन अभी भी बहुत small ताकि new version को maintain करने की operational cost को justify किया जा सके।

दो questions जो declare करने से पहले पूछने के लिए एक winner हैं: क्या effect large enough है new version को maintain करने की operational overhead को justify करने के लिए? और क्या कोई भी secondary metric degrade किया? एक prompt change जो task success rate को improve करता है जबकि cost को 30% से increase करता है एक net win नहीं हो सकता है, deployment की budget constraints पर depend करता है।

LLM experiments के पास एक additional failure mode है जो classical A/B tests के पास नहीं है: interaction effects treatment और specific input types के बीच। एक prompt change जो typical inputs पर performance को improve करता है edge-case inputs पर performance को degrade कर सकता है जो test period में rarely appear करते हैं लेकिन एक future seasonal spike में frequently। Test period और input distribution alignment LLM experiments में most other software contexts से अधिक matter करता है।

Shadow testing: किसी भी user को देखने से पहले एक change को validate करना एक live A/B test real users को new version में send करता है, जिसका मतलब है एक regression उन्हें experiment close होने से पहले कुछ fraction तक reach करता है। एक तरीका है real traffic के विरुद्ध test करने के लिए बिना उस exposure के। आप new version को current के साथ parallel में run करते हैं, इसे live requests की एक copy send करते हैं, और हर user को current version का response serve करते हैं। New version के outputs logged हैं rather than returned, और आप उन्हें offline के बाद score करते हैं। Deployment decision को एक single user ने new version को देखने से पहले किया जाता है। वह pattern shadow testing कहलाता है।

दोनों patterns के बीच choice down आता है कितना risk deployment को absorb कर सकता है और कितना traffic यह देखता है।

एक live A/B test use करें जब deployment एक worse version को एक small, bounded amount of exposure को absorb कर सकता है और traffic volume high enough है एक statistically meaningful sample को reach करने के लिए एक reasonable window में। Payoff यह है कि आप new version को real user behavior के विरुद्ध measure करते हैं, जिसमें downstream signals शामिल हैं एक live response produce करता है, जैसे कि क्या user ने answer को accept किया या follow up किया।

Shadow testing use करें जब एक single bad output बहुत अधिक risk carry करता है, या जब traffic बहुत low है एक live split को support करने के लिए change को needed होने से पहले। Shadow testing की cost downstream signal की loss है: कोई users shadow output को receive नहीं करते हैं, इसलिए scoring एक offline rubric या golden answers पर rely करता है rather than real user behavior। एक regulated-industry deployment के लिए, जहां unvalidated model change को users को expose करना permissible नहीं हो सकता है, shadow testing अक्सर change को validate करने का एकमात्र acceptable तरीका है।

Scale पर observability: instrumentation design, dashboards, anomaly detection एक Claude system के लिए production observability को चार questions का जवाब देने की आवश्यकता है: system क्या कर रहा है, यह कितनी अच्छी तरह perform कर रहा है, यह कब change किया, और यह क्यों change किया? प्रत्येक question को instrumentation की एक different layer की आवश्यकता है।

Request-level tracing। हर request को एक trace produce करना चाहिए जो model, model version, input token count, output token count, latency, stop reason, और कोई भी tool calls को include करता है। यह raw material है सब कुछ के लिए। Metric aggregation। Request-level data को aggregate करें dashboard display करने वाले metrics में: cost per request, latency p50 और p95, task success rate (यदि downstream system एक acceptance signal provide करता है), और error rate by error type। Per-request decomposition critical है: aggregate metrics healthy दिख सकते हैं जबकि requests का एक small fraction most of the budget को consume करता है। Anomaly detection। Metrics पर threshold alerts set करें जो deployment के लिए matter करते हैं। एक cost spike जो 7-day average के 150% को exceed करता है एक alert deserve करता है। एक latency p95 जो SLA threshold को cross करता है एक alert deserve करता है। Model drift (model outputs के distribution में gradual change over time) threshold alerts के साथ detect करना harder है और periodic distribution comparison से benefit करता है। Change attribution। जब एक metric move करता है, instrumentation को distinguish करने में सक्षम होना चाहिए Model drift (stable inputs पर model का behavior changed), data drift (input distribution changed), और model update effects (model version changed और new version existing inputs पर differently behave करता है)। ये तीन causes different mitigations हैं और उन्हें mix करना गलत fix produce करता है।

एक failure taxonomy: classify करना कि आप क्या देख रहे हैं Instrumentation आपको बताता है एक metric moved; diagnosis आपको बताता है किस kind की failure ने इसे moved किया। Attribution से पहले, failure को classify करें। Common classes distinct हैं और different fixes के लिए call करते हैं:

Prompt failure। Instruction ambiguous या underspecified था और model ने gap को fill किया। Fix prompt में है, model में नहीं। Hallucination। Model ने confident, fluent content produce किया जो input या एक reliable source में grounded नहीं है। Fix retrieval, tool use, या verification के माध्यम से grounding में है। Stronger instruction इसे resolve नहीं करेगा। Model mismatch। Chosen tier task के लिए गलत था या swap किया गया था बिना re-evaluation के। Fix model selection में है, एक eval द्वारा gated। Orchestrator-workers failure। Multi-agent systems में, orchestrator और subagents में trace करें: एक recoverable subagent failure (retry या flag) एक unrecoverable orchestrator failure से अलग दिखता है। Attribution को एक trace की आवश्यकता है जो दोनों को span करता है।

Discernment: output की quality को judge करना Discernment चार AI Fluency competencies में से एक है, defined के रूप में model द्वारा produce किए गए quality को judge करने की discipline rather than इसे face value पर accept करना। एक production system पर applied, Discernment हर output को classify करने की habit है acceptable, needs revision, या needs override के रूप में, और वह judgment को evals और monitoring में feed करना। एक team बिना Discernment के metrics को move करते हुए देखता है और कभी नहीं पूछता कि क्या underlying outputs actually अच्छे थे।

Observability data को business value से connect करना जिन लोगों ने Claude deployment को fund किया वे request-level trace को नहीं read कर रहे हैं। वे एक KPI dashboard को read कर रहे हैं जो deployment को improve करने के लिए design किया गया था outcome को measure करता है। Observability stack को एक translation layer की आवश्यकता है जो technical metrics को business metrics से connect करता है जो वे drive करते हैं।

एक customer service agent के लिए, business metric average handle time, first-contact resolution rate, या customer satisfaction score हो सकता है। Observability stack latency, task success rate, और error rate को measure करता है। Translation layer task success rate को first-contact resolution में map करता है और latency को handle time में, इसलिए business owner देख सकता है कि क्या system numbers को move कर रहा है जो वे care करते हैं।

यह translation layer को build करें जब system को design किया जाता है, deployment के बाद नहीं। यदि technical और business metrics को build time पर map नहीं किया जाता है, first business review एक question raise करेगा कि क्या handle time में change को drive कर रहा है। उस mapping के बिना, इसका जवाब देना एक retrospective reconstruction को require करता है rather than एक live query को run करना।

Cost · Complexity · Risk Cost: एक A/B test को pre-specify किए बिना primary metric चलाना का मतलब है आप हमेशा एक result find कर सकते हैं जो आप चाहते हैं। Underpowered experiments false positives produce करते हैं। एक change जो एक improvement जैसा दिखता है deployed हो जाता है, और team एक version को maintain करने में end up करता है no better than क्या यह replaced जबकि full operational overhead को absorb करता है। Complexity: Observability instrumentation जो first production incident के बाद add किया जाता है root cause question को existing log data से answer नहीं किया जा सकता। Incremental complexity को build करने की instrumentation को correctly first time से lower है than retroactive log reconstruction की complexity। Risk: एक LLM system जिसके पास केवल aggregate-only observability metrics हैं healthy दिख सकता है जबकि requests का एक small fraction most of the budget को consume कर रहा है और wrong outputs produce कर रहा है। Aggregate metrics obvious failures के विरुद्ध protect करते हैं। Per-request decomposition non-obvious ones के विरुद्ध protect करता है।

स्क्रीन 15: 50-session winner जो नहीं था

Watch Out A/B & Obs5 मिनट 50-session winner जो नहीं था

यह गलती क्यों करना आसान है एक proper A/B test को run करना time लेता है, sample size calculation को require करता है, और significance तक reach करने में days या weeks ले सकता है। 50 sessions के new version को 50 sessions के old version के विरुद्ध compare करना एक afternoon लेता है। एक 50-session comparison जो positive दिखता है एक confirmation check है rather than एक meaningful test। Sample बहुत छोटा है signal को noise से distinguish करने के लिए।

Prompt change जो एक win जैसा दिखता था निम्नलिखित एक composite है जो एक pattern को represent करता है जो teams में surface करता है जिनके पास एक working system है और इसे improve करना चाहते हैं लेकिन एक formal experimentation process नहीं है।

एक team एक customer service agent को run कर रहा था system prompt में एक revised instruction को test करना चाहता था। उन्होंने 50 customer sessions के विरुद्ध new version को run किया और 50 sessions के विरुद्ध old version को। Task success rate 68% था new version पर और 62% old version पर। उन्होंने new version को एक winner declare किया और deploy किया।

दो हफ्ते बाद, task success rate new version पर 61% पर settled हो गया था। Apparent 6-point gain disappear हो गया था।

क्या break हुआ Comparison के तीन problems थे, जिनमें से कोई भी result को invalidate करने के लिए sufficient था।

Sample size बहुत छोटा था। एक 6-point difference को एक metric पर high variance के साथ reach करने के लिए statistical significance को hundreds में एक sample size को require करता है। 50 sessions per group पर, observed difference noise floor के अंदर था। Input distribution को control नहीं किया गया था। 50 sessions treatment group में happen करते हैं edge-case inputs को fewer contain करने के लिए than 50 sessions control group में। Apparent improvement input distribution का एक artifact था जो कौन से inputs कौन से group में routed किए गए थे। Primary metric को pre-specify नहीं किया गया था। Team ने task success rate को compare किया क्योंकि यह सही direction में move किया। यदि यह गलत direction में move किया होता, वे दूसरे metric को देखते। Result के बाद metric को select करना एक test को एक search में turn करता है जो metric happen करता है move करने के लिए।

क्या Watch Out करें एक underpowered experiment metric selection के बाद confirmation produce करता है rather than evidence, क्योंकि result hypothesis को reflect करता है जो आप के साथ start किए। Result noise था, और noise एक signal जैसा दिखता था क्योंकि sample signal को noise से tell करने के लिए बहुत छोटा था।

स्क्रीन 16: Experiment-design plane पर place करें

Checkpoint A/B & Obs5 मिनट Experiment-design plane पर place करें

नीचे दिया गया plane दो axes है: expected effect size (कितना बड़ा difference आप expect करते हैं देखने के लिए) और confidence requirement (कितना certain आपको होने की आवश्यकता है result पर act करने से पहले)। प्रत्येक scenario को सही zone में place करें। सही placement experiment design को determine करता है, जो बदले में minimum sample size को determine करता है।

A. एक low-stakes FAQ chatbot में एक clarification message के लिए एक minor wording change। B. एक medical intake summarizer के लिए एक prompt architecture change जहां एक error treatment को delay कर सकता है। C. एक routing model में एक नई classification category को add करना expected को capture करने के लिए 30% of incoming volume। D. Testing करना क्या Sonnet से Haiku में switch करना एक simple formatting task पर cost को save करता है बिना quality को degrade किए। E. एक RAG system में retrieval prompt में एक small change 200 requests per day को process करते हुए।

← Small effect Large effect →

↑ Confidence increases

High confidence

Small effect · High confidence

Large effect · High confidence

Moderate confidence

Small effect · Moderate confidence

Large effect · Moderate confidence

Low confidence

Small effect · Low confidence

Not used in this exercise

उत्तर जांचें अभी के लिए छोड़ें

स्क्रीन 17: Exercise: evaluation framework को define करें

Exercise Evals8 मिनट Exercise: evaluation framework को define करें

Brief एक regional insurer एक Claude system को deploy कर रहा है जो एक submitted claim को read करता है, structured fields को extract करता है (claimant, policy number, loss amount, date of loss), adjuster के लिए narrative को summarize करता है, और claims को flag करता है जो fraud review को warrant कर सकते हैं। System को कुछ seconds में respond करना चाहिए, एक defined per-claim cost के अंदर रहना चाहिए, कभी एक claimant के data को दूसरे के summary में leak नहीं करना चाहिए, और कभी एक claim को auto-deny नहीं करना चाहिए।

इस system के लिए evaluation framework को draft करें। नीचे दिए गए पांच dimensions में से प्रत्येक के लिए, metric को write करें, grading method (code-based eval, LLM judge, या human review), और एक one-sentence reason आपकी choice के लिए। अपना framework write करें, फिर नीचे model answer को reveal करें।

[Click-to-reveal एक answer को compare करने के लिए; feel free को Claude को ask करने के लिए compare करने के लिए जो आपने written है provided answer के साथ।]

  • Accuracy: field extraction
  • Latency: response time
  • Safety: summary faithfulness और no auto-deny
  • Security: no cross-claimant data leakage
  • Cost: per-claim spend

Model answer को reveal करें अभी के लिए छोड़ें

Model answer

  • Accuracy, field extraction: Code-based eval। Expected values (claimant name, policy number, loss amount, date of loss) known हैं और exact या schema match द्वारा verifiable हैं। कोई interpretation required नहीं; एक function output को ground truth के विरुद्ध check करता है।
  • Latency, response time: Code-based eval। Latency p95 एक number है। Check यह है कि क्या यह target के नीचे fall करता है। कोई judgment involved नहीं है।
  • Safety, summary faithfulness और no auto-deny: दो methods required हैं। Code-based eval deny action के लिए (binary, या एक denial issue किया गया था या नहीं)। LLM judge summary faithfulness के लिए (क्या narrative accurately source claim को represent करता है बिना fabrication के एक interpretive task है एक function encode नहीं कर सकता)।
  • Security, no cross-claimant data leakage: Code-based eval। Cross-claimant leakage को check किया जा सकता है प्रत्येक summary को scan करके identifiers के लिए जो किसी भी input claim में appear करते हैं except जो एक को being processed। Deterministic check, कोई interpretation needed नहीं।
  • Cost, per-claim spend: Code-based eval। Cost एक numeric value है derived input tokens, output tokens, model tier, और क्या prompt caching apply करता है से। Check यह है कि क्या यह ceiling को exceed करता है।

Correct: आपने सभी पांच dimensions में grading ladder को correctly apply किया। Distinguishing moves: Dimension 3 को code में action के लिए और judge में faithfulness के लिए split करना, और Security को code में रखना क्योंकि check deterministic है despite high stakes। Partially correct: कम से कम एक dimension mismatched है। Most common errors: field extraction के लिए एक judge का उपयोग करना (यह deterministic है, code इसे handle करता है), और auto-deny check को interpretive के रूप में treat करना (यह binary है, code इसे handle करता है)। Model-based बनाम code-based evals table को re-read करें और revise करें। Incorrect: Grading ladder को re-read करें: code का उपयोग करें जहां behavior unambiguous है (known value, binary check, numeric threshold); एक judge का उपयोग करें जहां correctness को interpretation को require करता है। प्रत्येक dimension पर वह question को apply करें और redo करें। Skip: यदि आपको आवश्यकता है तो move on करें, लेकिन cumulative task आपको scratch से एक eval strategy को design करने के लिए पूछता है। इससे पहले return करें।

Mark this complete जब आप अपने self-assessment से satisfied हों।

Mark complete

स्क्रीन 18: Production readiness builder

Cumulative Module13 मिनट Production readiness builder

Brief एक 600-person management consulting firm एक internal knowledge assistant को deploy करना चाहता है। Assistant को consultants को past engagement reports से relevant excerpts को retrieve करने में help करना चाहिए, firm methodologies के बारे में सवालों का जवाब देना चाहिए, और client RFPs के लिए first-draft responses को generate करना चाहिए past work को एक source के रूप में उपयोग करते हुए। Knowledge corpus 12,000 documents है, 5 से 80 pages तक ranging। Consultants active engagements के दौरान assistant को use करते हैं, peak usage के साथ 800 requests per day। Firm के पास एक $3,000/month cost ceiling है। Response time को 8 seconds पर p95 के अंतर्गत होना चाहिए। Firm के पास एक existing SSO system है (जैसे Okta या others) और एक document management system (जैसे SharePoint या others)। कई documents client-confidential information को contain करते हैं जो NDA के अंतर्गत है।

अपने responses को compare करने के लिए click-to-reveal answers; feel free को Claude को ask करने के लिए compare करने के लिए जो आपने written है provided answer के साथ।

Decision 1, Eval strategy। किसी भी build शुरू होने से पहले, आपको define करने की आवश्यकता है कि इस system के लिए success क्या दिखता है। आपका primary eval task क्या है? आप retrieval relevance और RFP draft quality को कैसे measure करेंगे? आप प्रत्येक के लिए कौन सा eval type use करते हैं? एक behavior को name करें जो एक code-based eval के लिए suited है और एक जो एक model-based eval के लिए suited है। आपकी golden dataset strategy क्या है? आप client-confidentiality constraint को कैसे handle करते हैं?

Decision 2, POC-to-production checklist। आपके पास एक working prototype है। Firm को deployment के लिए commit करने से पहले आपको क्या verify करने की आवश्यकता है? Scenario के लिए एक cost model build करें। क्या projected cost $3,000/month ceiling के अंदर fall करता है? कौन सा reliability pattern इस deployment के लिए सबसे critical है और क्यों? इस architecture type के लिए specific failure mode को name करें जो आप सबसे अधिक concerned हैं।

Decision 3, Use-case sizing और feasibility। Use case को चार AI properties के माध्यम से run करें। क्या यह feasible as scoped है? Working memory axis को apply करें। क्या 80-page documents एक constraint को present करते हैं? Knowledge axis को apply करें। Firm की proprietary methodologies training data में नहीं हैं। Mitigation क्या है? Feasibility verdict को state करें और load-bearing boundary condition को name करें।

Decision 4, Integration pattern selection। Firm के पास Okta SSO और SharePoint है। Okta identity boundary कहां sit करता है Claude call के relative में? SharePoint में documents client-confidential files को include करते हैं, आप data handling constraint को कैसे handle करते हैं? आप observability के लिए क्या instrument करते हैं?

Decision 5, A/B testing posture। Firm एक नई retrieval configuration को test करना चाहता है जो वे believe करते हैं RFP draft quality को improve करेगा। Hypothesis को frame करें। Treatment क्या है, primary metric क्या है, और secondary metrics पर constraint क्या है? Required sample size का अनुमान लगाएं। Baseline task success rate 70% है और आप एक 5-point improvement को detect करना चाहते हैं। Experiment duration के लिए यह क्या मतलब है 800 requests per day पर? Corpus structure को given input distribution control को आपको क्या चाहिए?

Model answer को reveal करें अभी के लिए छोड़ें

Model answer Decision 1, Eval strategy: Code-based eval: extracted citations पर schema compliance (document name, page number, section)। Model-based eval: RFP के लिए draft response की relevance और appropriateness। Golden dataset: past RFPs से built redacted client names के साथ, firm के historical engagements से sourced। Client-confidential documents eval set से excluded हैं जब तक client को explicit approval न दे।

Decision 2, POC-to-production checklist: Cost model: 800 requests/day × 30 days = 24,000 requests/month। System prompt ~2,000 tokens (methodology overview + instructions) + retrieved context ~3,000 tokens + query ~200 tokens = ~5,200 input tokens। Output ~600 tokens। Sonnet tier पर caching के साथ stable system prompt पर, projected cost range के अंदर है। ये flat averages हैं, और answer को यह clear करना चाहिए: projection assume करता है requests mean के पास cluster करते हैं बिना heavy tail के। Corpus 5 से 80 pages तक चलता है, इसलिए retrieved-context size likely bimodal है, और large excerpt requests की एक tail input cost को flat average model पर understate करता है। Within-ceiling verdict sensitivity analysis के साथ check किया जाना चाहिए इससे पहले कि यह relied on हो। Reliability: Sonnet से Haiku में fallback chain latency spikes के लिए। Most critical failure mode: retrieval quality drift जैसा documents को corpus में add किया जाता है। यदि नए documents को inconsistently index किया जाता है, retrieval precision degrade होता है और draft quality degrade होता है इसके साथ।

Decision 3, Use-case sizing और feasibility: Working memory: इस corpus में कोई भी individual document (up to 80 pages, approximately 29,000 tokens) current Claude context window को exceed नहीं करता है 1 million tokens के most current models पर। Binding working memory constraint corpus scale है: 12,000 documents को simultaneously context में load नहीं किया जा सकता। यह RAG architecture को drive करता है जो। एक single document पर maximum length पर, यह context window के अंदर comfortably fit करता है और अपने आप पर एक working memory constraint को present नहीं करता है। पूरे 12,000-document corpus में, 5–80-page range में average document length roughly 14,600 tokens है। Total corpus approximately 175 million tokens को represent करता है। यह किसी भी context window को far exceed करता है। Individual documents के लिए इस size range के अंदर, chunking strategies available हैं लेकिन load-bearing constraint नहीं हैं। Architecture को एक retrieval layer को require करता है relevant excerpts को query time पर surface करने के लिए। Knowledge: Firm की proprietary methodologies Claude के training data में नहीं हैं, जिसका मतलब है model उन्हें memory से supply नहीं कर सकता। Mitigation retrieval है: methodology documents को same RAG layer में indexed हैं engagement reports के रूप में और query time पर context window में surface किए जाते हैं जब एक question उनके लिए relevant है। System को Claude को methodologies को know करने की आवश्यकता नहीं है; इसे Claude को reasoning करने की आवश्यकता है methodology excerpts पर जो retrieval layer इसके सामने रखता है। Constraint यह है कि retrieval quality को high enough होना चाहिए सही methodology content को एक given query के लिए surface करने के लिए। Retrieval precision methodology queries पर को separate metric के रूप में track किया जाना चाहिए eval suite में और production observability में। Feasibility verdict: Feasible with constraints। सभी चार AI property axes scoped architecture के अंदर addressable हैं। Working memory constraints RAG layer द्वारा resolved हैं। Knowledge gaps proprietary methodologies पर retrieval corpus में उन documents को index करके resolved हैं। Steerability requirement system prompt structure और RFP drafts के लिए output schema द्वारा met है। Load-bearing boundary condition retrieval index coverage और freshness है। System काम करता है जब तक retrieval index relevant methodology और engagement documents को contain करता है और corpus changes के रूप में current रखा जाता है। यदि index incomplete है, यदि documents को बिना indexed किए add किया जाता है, या यदि index document management system के साथ sync से drift करता है, knowledge axis mitigation fail करता है और system responses produce करेगा जो firm methodology को omit या misrepresent करते हैं। वह condition को statement of work में एक explicit operational constraint के रूप में document किया जाना चाहिए।

Decision 4, Integration pattern selection: Identity: Okta token को server-side verify किया जाता है। User role और authorized document sets को server द्वारा system prompt में inject किए जाते हैं, user द्वारा supplied नहीं। Client-confidential handling: SharePoint में confidential के रूप में tagged documents केवल consultants द्वारा retrievable हैं जिनके role relevant client engagement को include करता है। Retrieval layer इस access control को enforce करता है content को Claude में pass करने से पहले। PII: documents में client names anonymized identifiers के साथ replace किए जाते हैं context window में document enter करने से पहले। Observability: log model version, input token count (caching hit/miss के साथ), retrieval precision per request (grounding documents के विरुद्ध measured), और output token count। Log user role और session ID context layer पर। Capture करें consultant की acceptance या RFP draft का revision outcome signal के रूप में। Alert करें latency p95 को 8 seconds को cross करने पर और retrieval precision को threshold के नीचे drop करने पर।

Decision 5, A/B testing posture: Hypothesis: नई retrieval configuration RFP draft task success rate को (consultant acceptance के रूप में measured बिना major revision के) 70% से कम से कम 75% तक increase करेगा, latency p95 को 8 seconds के ऊपर increase किए बिना या cost per request को 10% से अधिक increase किए बिना। Sample size: एक 5-point improvement को 70% baseline पर detect करना 80% power और 5% significance के साथ approximately 1,500 sessions per group को require करता है। 800 requests per day पर evenly split, यह approximately 4 days लेता है, जो feasible है। Input distribution control: ensure करें treatment और control groups को RFP complexity के similar distributions हों (proxy: document length और retrieved documents की number)। यदि possible है RFP type द्वारा segment करें।

अपने answers को सभी पांच decisions में compare करें। Strong answers specific numbers को name करते हैं (cost model, ~175M-token corpus, 1,500 sessions/group, ~4 days) और प्रत्येक decision को next से tie करते हैं: sizing model cost ceiling को feed करता है, feasibility boundary condition (retrieval coverage और freshness) observability को feed करता है, और eval A/B primary metric को feed करता है।

Mark complete

स्क्रीन 19: Glossary

Reference Wrap-up Glossary

इस module में use किए गए key terms, alphabetical order में। एक term को expand करने के लिए click करें।

5xx errors HTTP status codes (500–599) का class जो एक server-side failure को indicate करता है एक otherwise valid request को fulfill करने में। Common examples में 500 (Internal Server Error), 503 (Service Unavailable) और 529 (Overloaded Error) शामिल हैं। Usually transient और retry और backoff के साथ resolvable, 4xx codes से distinct जो एक client-side problem को indicate करते हैं।

BAA (Business Associate Agreement) HIPAA के अंतर्गत एक contract required है एक covered entity (या business associate) और एक vendor के बीच जो protected health information को handle करता है इसकी ओर से। BAA specifies करता है safeguards जो vendor PHI को apply करेगा। Claude deployments के लिए, BAA coverage configuration-specific है: same surface एक delivery route पर BAA-covered हो सकता है और दूसरे पर नहीं। Coverage को configuration द्वारा check करें, product द्वारा नहीं।

Caching Caching reusable prompt content को store करता है इसलिए system को हर request पर इसे reprocess करने की आवश्यकता नहीं है। यह सबसे effective है जब system prompt long और stable है, दोनों token cost और response latency को reduce करते हुए; response जो आप receive करते हैं identical है जो आप बिना caching के get करते।

Circuit breaker एक reliability control जो एक downstream dependency पर error rate को monitor करता है और, जब errors एक defined threshold को exceed करते हैं, उस dependency को further requests को block करता है एक cooldown window के लिए ताकि एक degraded component calling system की capacity को consume न करे। Service boundary पर sit करता है, retries से distinct (API call के close) और fallback chains (orchestration layer में)।

Data-residency pinning Integration को configure करना ताकि model execution एक specified geographic region में happen हो, typically sectoral regulations को satisfy करने के लिए, या internal data-residency policy। Pinning को delivery route level पर CSP-mediated integrations पर implement किया जाता है। Pin को integration layer पर set किया जाना चाहिए और हर request पर verified किया जाना चाहिए, entry point choice द्वारा assumed नहीं।

DPA (Data Processing Agreement) एक contract एक data controller और एक data processor के बीच defining कि कैसे personal data को controller की ओर से handle किया जा सकता है, जिसमें processing scope, security obligations, sub-processor terms, और breach notification शामिल है।

Eval एक structured test set used को measure करने के लिए क्या एक model एक defined task पर well perform कर रहा है। एक eval inputs को expected outputs या quality criteria के साथ pair करता है, उन्हें model के विरुद्ध चलाता है, और एक score produce करता है जिसे आप model versions, prompts, या configurations में compare कर सकते हैं। Evals वह हैं कि कैसे teams decide करते हैं क्या एक change एक improvement है या एक regression production तक reach करने से पहले।

Exponential backoff Exponential backoff एक retry strategy है जहां, एक failed request के बाद, system फिर से try करने से पहले wait करता है और प्रत्येक successive wait पहले से longer है, typically हर बार doubling।

GDPR (General Data Protection Regulation) European Union (EU) regulation governing EU और European Economic Area (EEA) में individuals के personal data को processing। Establishes करता है lawful-basis requirements, data subject rights, controller और processor obligations, और fines up to 4% of global annual turnover।

Generator-verifier loop एक two-stage pattern जिसमें एक model-generated output को एक second pass द्वारा check किया जाता है downstream में use किए जाने से पहले। Verifier एक deterministic code-based check (schema validation, comparison against एक authoritative value) या एक second model call scoped को evaluation में हो सकता है। Used को एक compensating control के रूप में जहां underlying task को अधिक precision को require करता है than single-pass generation reliably provide करता है।

Hallucination rate Responses का percentage जिसमें model invented, inferred, या stated information को जो input, source data, या allowed logic द्वारा supported नहीं था।

Live state Data जो एक conversation या process के lifetime के दौरान change करता है: एक order status, एक inventory count, एक price, एक calendar slot, एक user का current session। Live state static reference material से distinct है क्योंकि 10:00 a. m. पर सही answer 10:05 पर गलत हो सकता है। Systems जिन्हें live state को require करता है source of truth के विरुद्ध एक direct lookup को require करते हैं, एक stored snapshot नहीं।

Median latency Median latency वह time है जो distribution में middle request को complete करने के लिए लेता है, जिसका मतलब है 50% of requests faster हैं और 50% slower हैं। इसे p50 latency भी कहा जाता है।

p95 (95th-percentile latency) Latency value जिसके नीचे 95% of requests complete होते हैं, slowest 5% के साथ इसके ऊपर fall करते हुए। Production design target के रूप में use किया जाता है क्योंकि SLA breaches और user abandonment distribution के slow tail द्वारा driven होते हैं rather than median।

PHI (Protected Health Information) Individually identifiable health information held या transmitted एक covered entity या business associate द्वारा US Health Insurance Portability and Accountability Act (HIPAA) के अंतर्गत। PHI को processing करना एक Business Associate Agreement को require करता है किसी भी third party के साथ जो इसे handle करता है।

RAG (retrieval-augmented generation) एक pattern जिसमें एक knowledge corpus को preprocessing step में chunk और index किया जाता है, और query time पर chunks जो user input के लिए सबसे relevant हैं retrieve किए जाते हैं और model के context में pass किए जाते हैं। Static या slow-moving knowledge जैसे manuals, internal documentation, और regulatory text के लिए suited। Live transactional state के लिए suited नहीं, जहां retrieval एक snapshot return करता है जो पहले से ही stale हो सकता है और एक tool call को system of record में सही mechanism है।

Rate limit एक server-enforced cap requests की number पर जो एक client एक defined time window के अंदर बना सकता है। जब cap exceed किया जाता है, server further requests को reject करता है (typically HTTP 429 के साथ) जब तक window reset न हो। एक rate-limit response throttling को indicate करता है rather than failure और एक transient condition है जो backoff के साथ resolve होता है।

Regex Regex regular expression के लिए stand करता है। यह एक rule-based pattern है used को find या validate करने के लिए text जो एक specific format को match करता है, जैसे email addresses, phone numbers, Social Security numbers, या credit card patterns।

Schema Required structure, format, और rules output के लिए। यह define करता है कि कौन से fields appear करने चाहिए, उनके data types, allowed values, और कैसे response को organize किया जाना चाहिए।

SSO (single sign-on) एक authentication arrangement जिसमें एक user एक बार एक central identity provider को sign in करता है और multiple downstream applications को access gain करता है बिना re-authenticating के।

Structured fields Structured fields वे हैं जिन्हें specific outputs को require करता है जिन्हें एक defined format में populate किया जाना चाहिए, जैसे customer name, invoice number, date, amount, या policy ID। ये discrete data elements हैं, free-form narrative text नहीं।

Timeout एक failure mode जिसमें एक request को client के या server के configured wait period के अंदर एक response को receive नहीं करता है और terminated है। Typically caused transient server load, network latency, या एक downstream dependency under stress द्वारा, rather than एक permanent fault।

Tool use Capability जो Claude को external functions, APIs, या services को call करने देता है एक response के दौरान instead of केवल text generate करने के। Model decide करता है कब एक tool को invoke करना है, कौन से arguments को pass करना है, और result को अपने next step में कैसे use करना है। Tool use Claude को एक text generator से एक system में turn करता है जो files को read कर सकता है, databases को query कर सकता है, web को search कर सकता है, या अन्य software में action ले सकता है।

Transient error एक transient error एक temporary failure है जो अपने आप को resolve करने के लिए expected है बिना किसी permanent fix के, जिसका मतलब है यदि आप एक short wait के बाद same request को फिर से try करते हैं, यह likely succeed होगा।

स्क्रीन 20: पांच takeaways

Recap Module3 मिनट पांच takeaways

Key takeaways:

01

Evals को acceptance criteria के रूप में Production code की पहली line लिखने से पहले eval suite को write करें। Golden dataset को हर system change के साथ current रखें। Eval को हर model swap या prompt revision के लिए gate के रूप में use करें।

02

POC से production Production volume पर cost और latency को architecture को commit करने से पहले estimate करें। Retries, fallback chains, और circuit breakers को initial design में build करें। अपने architecture type के लिए specific failure mode को name करें और mitigation को document करें।

03

Use-case sizing और feasibility किसी भी नए use case को feasibility verdict issue करने से पहले चार AI properties के माध्यम से run करें। Verdict को तीन forms में से एक में state करें: feasible as scoped, feasible with constraints, या not feasible। हर constrained verdict के लिए load-bearing boundary condition को document करें।

04

Enterprise integration patterns Identity को server side पर enforce करें। केवल minimum necessary data को context window में pass करें। Observability instrumentation को design time पर build करें, first incident के बाद नहीं।

05

A/B testing और observability Experiment को run करने से पहले primary metric और sample size को identify करें। Scale पर, per-request observability को aggregate dashboards से separate करें। एक translation layer को build करें जो technical metrics को business metrics में map करता है जो stakeholder care करता है।

Module 3 responsible AI, safety, और risk को Architects के लिए cover करता है: guardrail design, regulated-industry considerations, और human-in-the-loop validation strategies।

Sources

Anthropic Skilljar, Building with the Claude API: eval workflow stages, model-based vs. code-based evals, cost और latency modeling, caching, tool use, streaming। Anthropic Skilljar, Claude 101: model family overview, context windows, general Claude capabilities। Anthropic Skilljar, Claude Code in Action: Claude Code को एक integration entry point के रूप में, agentic patterns in practice। Anthropic Skilljar, AI Capabilities और Limitations: four AI properties framework, foundational concepts। platform. claude. com/docs: models overview page (capability और context-window figures), pricing page (per-token price points for cost modeling), और prompt caching page (caching mechanics और consistency risk guidance)। Anthropic, Building Effective Agents: workflow और agent design patterns; कब agents को use करें vs. simpler architectures। Anthropic Cookbook: failure-mode candidates और worked patterns।

स्क्रीन 21: बधाई! आपने successfully इस module को complete किया है।

Module Complete · Architect · 2 मिनट बधाई! आपने successfully इस module को complete किया है।

Module 2 integration patterns, production infrastructure, और operational decisions को cover करता है जो एक Claude deployment को prototype से enterprise scale तक ले जाते हैं।

Production reliability एक architecture decision है, एक operational नहीं; आपके पास अब design time पर इसे बनाने के लिए patterns हैं।

0 of 0 checkpoints passed

M1

Claude Platform & Solution Design Model selection, prompt architecture, tool design, और platform-layer tradeoffs।

M2

Enterprise Integration & Production Deployment patterns, integration architecture, और production reliability।

You Are Here

M3

Responsible AI, Safety & Risk Safety frameworks, risk identification, और governance practices।

Up Next

M4

Stakeholder Engagement, Lifecycle & Go-to-Market Stakeholder communication, lifecycle management, और go-to-market strategy।

M5

Team Enablement और Operational Productivity Team tooling configuration और operational support practices।

Module को review करें Start over

फ्लैशकार्ड 0 कार्ड

No flashcards for this lesson.

ज्ञान जाँच 0 प्रश्न

No quiz for this lesson yet.