Skip to main content
An evaluator pack is a versioned definition of a workflow and the decision or outcome to assess. It describes what the evaluator may use as evidence, the question it must answer, and what each answer means. Packs are designed to keep business meaning, evidence, and model judgments inspectable. A proposed pack is a draft: connecting a trace source or creating a pack does not activate evaluator-pack scoring. Existing operational views may still summarize captured event signals such as success, latency, or token use.

The pack lifecycle

  1. Connect evidence. Bring in traces from a supported integration or instrument your agent with an SDK. Each source exposes different fields and history.
  2. Review workflow findings. Confirm or correct what the workflow is trying to complete, which states count as success or failure, what actions matter, and which source facts prove an outcome. Unsupported or contradictory claims remain unresolved.
  3. Inspect evaluator proposals. A proposal should identify its unit of evaluation, question, answer criteria, evidence recipe, source support, and unresolved assumptions. Definitions are versioned.
  4. Review eligible examples. Read the human question and answer meanings alongside the case evidence. Source and trace provenance are shown when recorded. If a pack has no eligible examples, there is no example queue to label.
  5. Assess readiness. Before a pack can be trusted for production use, its labels, calibration, holdout isolation, independent review, coverage, and owner approval must meet the applicable gates.

Review an example

When an example is available, the reviewer should be able to:
  • read the specific decision question and the criteria for each answer;
  • inspect the decisive facts and attributed trace evidence;
  • save work in progress and return to it;
  • lock an answer and continue through the eligible queue; and
  • compare with the model’s answer only after locking the human answer.
The reviewer should not have to infer the question from technical field names or from the order of model-generated questions. Missing evidence should remain visible rather than being treated as a successful or failed result.

Draft review is not live scoring

The current workspace supports draft pack definitions and human review for eligible cases. The production evaluator remains separate from that review experience. An unapproved pack’s judgments must not be treated as approved evaluator results or drive customer routing, alerts, or automatic actions. Promotion requires an independently supported validation path. Depending on the pack, this includes disjoint calibration labels, an untouched sealed set, adequate answer-class coverage, separate reviewer checks, and approval from the authorized owner. A model’s own judgments are not independent labels.

Source-specific setup

The source system and trace identity shown in a review must come from recorded provenance. Tensile does not infer a source from an organization or customer name.