Configuration
Enable assessments on the Assessment page in the Control Plane. Assessments are included in LLM logs. Under Risk and misalignment checks, create conditions with a name, type, severity, and enabled status. Risk checks identify harmful consequences. Misalignment checks identify departures from the user’s expressed or reasonably implied intent. You can edit, disable, or delete defaults, and add your own checks. Deleting a default does not recreate it when you enable assessments again. Verify: Save a check and confirm its condition, type, severity, and status in the checks table. On the same page, create labels for topics or behaviors. Each label has a condition that completes “The current conversation …” and a probability threshold. Formal assigns the label when it is active and the assessed probability reaches the threshold. For example, labels can identify a broad topic or a specific behavior:Policy Evaluation
LLM response policies can use the following fields ininput.assessment:
Alignment and risk scores are available for LLM responses with tool calls. If assessment data is unavailable, conditions that reference
input.assessment do not match.
A check contributes nothing when its assessed probability is 0.5 or lower, and its full configured severity when the probability is 0.8 or higher. Between those points, its contribution rises linearly from 0 to the severity.
For each tool call, risk is the highest risk contribution. Alignment is
1 minus the highest misalignment contribution. With zero contributions, risk is 0 and alignment is 1. Across a response, Formal uses the highest tool call risk and lowest alignment. Missing or invalid assessment answers make assessment data unavailable.
Repeated identical requests can reuse an assessment for up to 30 seconds. Changes to checks may take that long to appear for those requests.
With the example engineering label above, this policy blocks an LLM response when that label applies and a tool call risk score is at least 0.7: