What is being assessed: a model, or a system that uses a model?
For procurement or research comparison, specify the model, interface, tools, permissions, environment, and recovery loop—not only the model name.
Read comparison →Should a system receive information at inference time, or have its behaviour adapted through training?
Choose based on evidence requirements: retrieval can make the immediate source context inspectable, while fine-tuning targets behaviour or task adaptation. Evaluate both against the same task, corpus, and update assumptions.
Read comparison →Does access to a tool establish autonomous task completion?
Treat tool access as one capability of a system. For higher-stakes workflows, ask for measured success, failure recovery, supervision, permissions, and rollback—not a tool-call demonstration.
Read comparison →How does the delivery mode change what can be evaluated and governed?
For reproducibility, privacy, and auditability, record the deployment mode, model artifact or API version, date, configuration, and surrounding system—not just the prompt and output.
Read comparison →What can a benchmark score establish about dependable work?
Use benchmark results as scoped measurements. Before deployment, require task-relevant evaluation, repeated trials, error analysis, drift monitoring, and a plan for human oversight or recovery.
Read comparison →