Skip to main content

Understand Safety Guardrails

Safety Guardrails (AI Reply Safety Guardrails) check an AI response for risk before it is sent to the customer, then act according to the rules you configure. They control whether the AI’s answer content itself is appropriate — even if the knowledge base has the answer, the model may reply with promises, overreach, or non-compliant content. That is what the guardrails catch. When a rule is hit, you can have the AI do three things: reply with specified wording, stay silent (do not reply), and notify a Leader. These are independent — for example, you can stay silent while notifying the Leader about the risk.

Where to find it

In the left navigation, click Assistant Settings → Safety Guardrails.

When to configure

  • You have seen promise-style replies (e.g., “I already checked”, “I already added you to the group”) that were not actually executed.
  • Replies involving sensitive commitments, competitors, violations, or content customers may not understand need to be intercepted.
  • You want the AI to stay silent in uncertain or high-risk scenarios, or hand off the issue to a Leader for manual handling.

Choose the detection method

Each rule must first set its detection method, which cannot be changed after creation, so consider it before adding. Keyword matching ignores spaces and line breaks so it does not miss replies with stray whitespace (e.g., “已 建群”). There can be only one model semantic judgment rule per communication window — put multiple scenarios in that single rule.

Add a guardrail rule

1

Select a communication window

First choose which AI assistant window this set of guardrails applies to.
2

Add a rule

Click “Add rule” and choose the detection method (keyword hard match / model semantic judgment).
3

Configure the trigger condition

Keyword: enter the words or phrases to intercept one by one, press Enter to add; model: describe the high-risk scenario clearly.
4

Configure the action

Choose “Reply” or “Do not reply”. If replying, enter the reply content to send to the customer; you can insert images or files.
5

Configure Leader notification

Whether to sync the Leader when the rule is hit. Notification is independent of the action.
6

Save the rule

Each rule is saved separately and takes effect immediately.

Rule fields

How rules are applied

  • Each rule is edited and saved separately; matching stops after the first hit (first match wins).
  • With no rules configured, the guardrails intercept nothing; add rules to make them active.

View results in Detail Review

In the data dashboard’s Detail Review, each message’s “message processing chain” shows the Safety Guardrail step result (passed or intercepted by a rule), helping you confirm whether rules are effective and what they hit.
Safety guardrails directly affect real customer conversations. Validate rule hits in a test chat or testing group before using them in production.