Guardrails and redaction
Sensitive Info Detection scans what people send, what tools return and what models answer, and warns, redacts or blocks card numbers, credentials, phone numbers and your own patterns before they leave your workspace.
People paste real things into AI: a customer’s card number in a refund request, an API key in a code snippet, a phone number in a CRM export. Guardrails catch that automatically. Sensitive Info Detection scans for sensitive information and warns, redacts or blocks it before it leaves your workspace, and every time it acts, it leaves a record in the Privacy Log.
The guardrail policies
On the Workspace page, open Guardrails. Two policies protect your workspace’s chats, prompts and agents:
| Policy | Status | What it does |
|---|---|---|
| Sensitive Info Detection | Active | Detects and redacts sensitive information shared in chats, prompts and agents. |
| Prompt Injection Detection | Coming soon | Will detect and block attempts to hijack your agents’ and prompts’ instructions. |
Set up Sensitive Info Detection
Select Configure on Sensitive Info Detection. The settings have four parts: where to scan, what to look for, your own patterns, and a preview.
1. Scan surfaces
Choose which parts of a request or response get scanned. Turn off all surfaces to disable detection.
| Surface | What it covers |
|---|---|
| User input | What people type and send. |
| Tool data | What tools and integrations hand back to the model, such as CRM contacts or database rows. |
| Model output | What the model answers. |
2. Built-in detectors
Each detector can be enabled or disabled on its own, and reacts the way you choose.
| Detector | What it finds |
|---|---|
| Payment data | Credit card numbers and IBANs. |
| Credentials | API keys, tokens, private keys and passwords found in text or code. |
| Email address | Email addresses. |
| Phone number | Phone numbers, in international or NANP formats. |
| IP address | IPv4 addresses. |
| Government ID | Government identification numbers. |
For each enabled detector, pick an action:
| Action | What happens |
|---|---|
| Warn | The text goes through unchanged, and the detection is recorded in the Privacy Log. |
| Redact | The sensitive value is replaced by a placeholder, so the model never sees it. |
| Block | The content is stopped and does not go through. |
3. Custom patterns
Under Custom Patterns, add your own regular expressions for anything specific to your business that the built-in detectors cannot know about - internal account numbers, employee IDs, project code names. For example, a pattern for Northwind customer account numbers:
NWC-\d{6}
4. Preview
Use Preview to try your settings on sample text before you rely on them, so you can see what gets caught and what gets through.
What the model receives
With Redact, the model gets a placeholder where the value was, and works with that. In this chat, the card number and IBAN became [payment_data] and the phone number became [phone]. The model still wrote a useful refund email, but it never saw the real numbers.
The same applies to data from your integrations. With Model output or Tool data scanned, a phone number from your CRM comes back as [phone] in the answer - see the deals table in CRM.
A recommended starting configuration
| Setting | Start with | Why |
|---|---|---|
| User input | On | Catches what people paste. |
| Tool data | On, if you use integrations | Scans CRM and database results before the model sees them. |
| Model output | On | A last check on everything shown and sent. |
| Payment data | Redact | The model rarely needs the real number to do the job. |
| Credentials | Block | A model never needs a real key, and a leaked one is a real risk. |
| Phone number | Redact | Personal data the model can usually work without. |
| Government ID | Redact or Block | High-risk personal data. |
| Email address | Off or Warn | People often need addresses to send email and look up contacts. Redacting them breaks those requests. |
| IP address | Off or Warn | Turn on if your teams handle logs or network data. |
Run with this for a week, then read the Privacy Log. Frequent Warn entries for a detector are a sign to move it to Redact; complaints that answers are missing something needed are a sign to relax it.