Skip to main content
Assessments provide scores and labels for LLM response policies. Formal scores tool calls for alignment with the user’s intent and risk. It assigns labels to LLM responses based on conditions evaluated against the conversation.

Configuration

Enable assessments on the Assessment page in the Control Plane. Assessments are included in LLM logs. Under Risk and misalignment checks, create conditions with a name, type, severity, and enabled status. Risk checks identify harmful consequences. Misalignment checks identify departures from the user’s expressed or reasonably implied intent. You can edit, disable, or delete defaults, and add your own checks. Deleting a default does not recreate it when you enable assessments again. Verify: Save a check and confirm its condition, type, severity, and status in the checks table. On the same page, create labels for topics or behaviors. Each label has a condition that completes “The current conversation …” and a probability threshold. Formal assigns the label when it is active and the assessed probability reaches the threshold. For example, labels can identify a broad topic or a specific behavior:

Policy Evaluation

LLM response policies can use the following fields in input.assessment: Alignment and risk scores are available for LLM responses with tool calls. If assessment data is unavailable, conditions that reference input.assessment do not match. A check contributes nothing when its assessed probability is 0.5 or lower, and its full configured severity when the probability is 0.8 or higher. Between those points, its contribution rises linearly from 0 to the severity. For each tool call, risk is the highest risk contribution. Alignment is 1 minus the highest misalignment contribution. With zero contributions, risk is 0 and alignment is 1. Across a response, Formal uses the highest tool call risk and lowest alignment. Missing or invalid assessment answers make assessment data unavailable. Repeated identical requests can reuse an assessment for up to 30 seconds. Changes to checks may take that long to appear for those requests. With the example engineering label above, this policy blocks an LLM response when that label applies and a tool call risk score is at least 0.7: