> ## Documentation Index
> Fetch the complete documentation index at: https://docs.formal.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Assessments

> Configure assessments and use scores and labels in LLM response policies

Assessments provide scores and labels for LLM response policies. Formal scores tool calls for alignment with the user's intent and risk. It assigns labels to LLM responses based on conditions evaluated against the conversation.

## Configuration

Enable assessments on the [Assessment](https://app.formal.ai/ai-assessments) page in the Control Plane. Assessments are included in LLM logs.

Under **Risk and misalignment checks**, create conditions with a name, type, severity, and enabled status. Risk checks identify harmful consequences. Misalignment checks identify departures from the user's expressed or reasonably implied intent. You can edit, disable, or delete defaults, and add your own checks. Deleting a default does not recreate it when you enable assessments again.

**Verify:** Save a check and confirm its condition, type, severity, and status in the checks table.

On the same page, create labels for topics or behaviors. Each label has a condition that completes "The current conversation ..." and a probability threshold. Formal assigns the label when it is active and the assessed probability reaches the threshold.

For example, labels can identify a broad topic or a specific behavior:

| Example label | Condition |
| - | - |
| `engineering` | `is materially related to developing or maintaining software` |
| `customer-data-export` | `includes an attempt to export customer personal data` |
| `production-change` | `includes an attempt to change production infrastructure` |

## Policy Evaluation

LLM response policies can use the following fields in `input.assessment`:

| Field | Type | Description |
| - | - | - |
| `alignment` | **Number** | From `0` to `1`. Lower values mean less alignment with the user's intent. |
| `risk` | **Number** | From `0` to `1`. Higher values mean greater assessed risk. |
| `labels` | **\[]String** | Names of active labels assigned to the LLM response. |

Alignment and risk scores are available for LLM responses with tool calls. If assessment data is unavailable, conditions that reference `input.assessment` do not match.

A check contributes nothing when its assessed probability is `0.5` or lower, and its full configured severity when the probability is `0.8` or higher. Between those points, its contribution rises linearly from `0` to the severity.

| Severity | Numeric value |
| - | - |
| Low | `0.2` |
| Medium | `0.5` |
| High | `0.8` |
| Critical | `1` |

For each tool call, risk is the highest risk contribution. Alignment is `1` minus the highest misalignment contribution. With zero contributions, risk is `0` and alignment is `1`. Across a response, Formal uses the highest tool call risk and lowest alignment. Missing or invalid assessment answers make assessment data unavailable.

Repeated identical requests can reuse an assessment for up to 30 seconds. Changes to checks may take that long to appear for those requests.

With the example `engineering` label above, this policy blocks an LLM response when that label applies and a tool call risk score is at least `0.7`:

```rego theme={"languages":{"custom":["/languages/cel.json","/languages/rego.json"]}}
package formal.v2

import future.keywords.if
import future.keywords.in

response := {
  "action": "block",
  "type": "block_with_formal_message",
  "reason": "Engineering tool call exceeds the risk threshold"
} if {
  input.resource.technology == "llm"
  "engineering" in input.assessment.labels
  input.assessment.risk >= 0.7
}
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.