Skip to content

Evaluation API ​

Import supported APIs from @umwelten/evaluation or the bundled umwelten package.

Suite execution ​

ExportPurpose
EvalSuiteDeclarative tasks, model execution, caching, scoring, and run output
runFullEvalCompose the standard language, coding, and tool-calling suites
makeLanguageSuiteConstruct the standard language suite
makeCodingSuiteConstruct the standard coding suite
makeToolCallingSuiteConstruct the standard tool-calling suite

Types: EvalSuiteConfig, EvalTask, VerifyTask, JudgeTask, VerifyResult, TaskResultRecord, FullEvalOptions, FullEvalResult, SuiteRunResult, LlmEvalSuiteName, and the three suite option types.

typescript
import { EvalSuite, runFullEval } from '@umwelten/evaluation';

EvalSuite.run() accepts { signal?: AbortSignal }. Its script-level flags are --all, --new, and --run N.

For an ad-hoc comparison, use umwelten eval run --prompt … --models … --id …. This command caches raw responses but does not score them.

Ranking ​

PairwiseRanker ranks existing responses through cached LLM-judge comparisons. Supporting exports include expectedScore, updateElo, buildStandings, allPairs, swissPairs, and evaluationResultsToRankingEntries.

See Pairwise Ranking API.

Aggregation and reports ​

ExportPurpose
loadSuiteLoad and normalize persisted evaluation dimensions
findLatestRunDir / loadDimensionLower-level persisted-run loading
buildSuiteReportBuild a structured multi-dimension report
buildNarrativeReportBuild a Markdown methodology/results narrative
ReporterRender structured reports to console, Markdown, or JSON

Related types include EvalDimension, SuiteResult, ModelScorecard, DimensionScore, report options, and report section types.

Not part of the API ​

There is no runEvaluation, EvaluationRunner, MatrixEvaluation, or BatchEvaluation API. Use an array of tasks for a batch, ordinary data iteration for a matrix, and an executable report script for aggregation.

See Evaluation architecture and Model Evaluation.

Released under the MIT License.