Outcome. You can translate one product use case into measurable quality, language, modality, context, tool, privacy, safety, latency, cost, and lifecycle requirements.
Model selection should begin before model names enter the discussion. Describe the job, user, input distribution, required output, source of truth, actions, consequence of failure, and operating environment. Then convert those facts into constraints and measures.
A requirements card includes task quality and severe-failure thresholds; English/Khmer needs; input/output modalities; effective context and retrieval; tool and structured-output behavior; data residency and retention; permission boundaries; p50/p95 latency; throughput; cost per successful task; availability; versioning; and migration needs. Mark each requirement as must-have, target, or preference.
Provider categories create different trade-offs: closed hosted APIs, cloud-platform managed models, open-weight hosted endpoints, and self-hosted models. Use current vendors as candidates only after requirements are set. Otherwise teams unconsciously rewrite the problem to justify the model they already prefer.
Mental model. The use case defines the scorecard; the scorecard filters candidates; the eval decides among survivors.
Evidence trail — reviewed 23 July 2026. NIST’s Generative AI Profile supports lifecycle risk measurement: https://doi.org/10.6028/NIST.AI.600-1. Use live model, pricing, data-control, and deprecation documentation only after the requirements are frozen.
Task: classify and summarize inbound tickets for human agents. Must: no autonomous customer action; EN/KM pass rate ≥92%; zero cross-customer data leaks in the test set; JSON schema ≥99.5%; p95 ≤4s; evidence quote for urgency; approved data retention. Target: ≤$0.03 per successful ticket. Preference: provider-hosted tools. Define 50 representative cases and severe failures.