> For the complete documentation index, see [llms.txt](https://docs.zeroauthority.xyz/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.zeroauthority.xyz/nova-bot/nova-ai-agent/community/nova-guardian.md).

# Nova Guardian

**Nova Guardian** is a passive community-protection shield. Once an admin turns it on, Nova silently screens **every message** in your Discord server or Telegram group for **scam intent and social engineering** — not just links — and takes graduated, configurable action. Turning it on protects the **whole** server/group.

The hard part of community safety isn't banning links (servers already do that) — it's the **text-based** lures that carry no link at all. Guardian reads **intent** with an LLM, so it catches paraphrases, typos, leetspeak and letter-spacing tricks that keyword blocklists miss. It also watches **behavior over time**, so it shuts down repeat-spam and flood blasts (the same message posted over and over) that a single-message check can't see.

***

## What It Catches

| Pattern                              | Example                                                                                                                                                                      |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **🎣 Fake hiring / recruiting**      | "We're onboarding mods/testers/ambassadors — DM me", fake remote web3 jobs, pay-per-task lures                                                                               |
| **🪂 Fake airdrops / claims**        | "Claim your airdrop", "verify your eligible wallet", "check eligibility", fake free-mints                                                                                    |
| **🔑 Seed / wallet harvesting**      | Asking for or "buying old/inactive wallets", "paying for seed phrases", "import/sync/restore your wallet"                                                                    |
| **🎭 Team / admin impersonation**    | "Official support here, DM me to resolve", or an account whose name clones a real admin's — automatically detected from who actually holds admin/mod permissions             |
| **📨 DM-bait / off-platform lures**  | "Let's continue in DM", "message me on WhatsApp/Telegram/Signal"                                                                                                             |
| **💸 Send-to-receive / advance-fee** | "Send 0.1 get 0.2 back", "double your crypto", "first 50 only", guaranteed-return giveaways                                                                                  |
| **🔗 Malicious / phishing links**    | Lookalike domains, wallet-drainer sites, fake claim/mint pages, hidden URL shorteners                                                                                        |
| **🎰 Gambling / betting-tips spam**  | "Insider picks" / tipster groups, casino recruiting, "earn 3k a day" betting lures that funnel members to a bot or channel handle — in **any language**                      |
| **💴 Illegal-goods spam**            | Counterfeit currency, forged IDs/documents, stolen accounts — spam regardless of what's being sold                                                                           |
| **🚫 Repeat-spam / flood**           | The same message (incl. a **forwarded** one) posted over and over — in a rapid burst **or** drip-fed minutes apart with a couple of emoji swapped between copies — see below |

Guardian judges **intent**, not keywords — and not just in English. Its quick pre-filter knows common non-English lure phrases (e.g. Chinese betting-spam and counterfeit-goods slang), and the AI verdict is made in whatever language the message is written in. Warning others about a scam, asking how airdrops work, or legitimate OTC trading talk are recognized as **clean** — and you can allowlist phrases that are normal for your community.

***

## Repeat-Spam & Flood Protection

The intent radar judges one message at a time, so Guardian adds a second, **behavioral** layer that watches *repetition over time* — the classic promo blast forwarded 9–10 times in a row, often with no link and not in English:

| Trigger              | When it fires                                                                                                                                                                                                        |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Repeated message** | The same (or near-identical) content posted **3×** by one account — in a rapid burst **or** drip-fed up to **45 minutes** apart (spammers space copies 10–15 min apart to dodge burst filters; that no longer works) |
| **Repeated forward** | A **forwarded** message repeated — fires sooner (**2×**), since a re-forwarded blast is a stronger spam tell                                                                                                         |
| **Rapid-fire**       | One account sending **8+** substantive messages within a minute                                                                                                                                                      |

"Near-identical" is deliberate: the content fingerprint ignores emoji, punctuation and spacing, so swapping 🦊🌷 for 🙈👧 between copies still counts as the same message. It keys on **repetition, not on forwarding** — **a single, normal forward is never touched**; only when the *same* account *repeats* itself does the detector react. Moderators and admins are exempt.

When a flood is caught, Guardian raises **one** review card for the whole burst (not ten), tagged **🚫 Repeat-spam / flood** — e.g. *"Repeat-spam flood: author posted the same message 3× (near-identical copies — emoji/punctuation swaps ignored) in under 29 min."* — and, with enforcement on, removes **every copy** in the burst at once. It's content- and language-agnostic, so it works on non-English and link-less spam that the intent classifier would miss.

***

## How It Decides — the action ladder

Every message gets a verdict, and each tier has its own action:

| Verdict                    | What Nova does                                                                                                                                                                                                                                                                                                                                                                                              |
| -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Clean**                  | Nothing.                                                                                                                                                                                                                                                                                                                                                                                                    |
| **Suspicious**             | Silently posts a **mod-review card** to your review channel — author, age, reason, message link. No public action. Never deleted or muted, even with auto-moderation on.                                                                                                                                                                                                                                    |
| **Scam** (high-confidence) | Posts the review card **and**, with **auto-moderation** on: **1)** deletes the message immediately, **2)** mutes the author for **24 hours** (Discord timeout / Telegram restrict — change it with *"mute scammers for 2 hours"*), **3)** privately tells the author which pattern they matched, how long the mute lasts, and how to appeal, **4)** hands your mods a report with **Ban / Revert** buttons. |

### Sensitivity modes

| Mode                     | Behavior                                                                              |
| ------------------------ | ------------------------------------------------------------------------------------- |
| **Lenient**              | Only the most blatant scams are flagged — fewest false positives.                     |
| **Balanced** *(default)* | Sensible middle ground.                                                               |
| **Strict**               | Catches more, flags more for review — a few more false positives for mods to dismiss. |

Trusted, established members are held to a lighter standard than brand-new accounts, so a long-time member and a fresh account posting the same line aren't treated the same.

***

## Auto-Moderation

Flag-only mode still needs a moderator to be awake. **Auto-moderation** closes that gap: when Guardian is certain a message is a scam, it acts immediately (the four actions above) and *then* asks a human — instead of leaving the scam up while it waits for one.

{% hint style="warning" %}
**Auto-moderation is the default when you enable Guardian** — on Discord only after you've nominated a mod target (Nova prompts you for one), so a false positive can never become an unappealable mute. Prefer to watch first? Say ***"activate Guardian without auto-moderation"*** and it starts in **flag-only** mode — it never deletes, mutes or bans anyone until you switch auto-moderation on later with *"@Nova turn on auto-moderation"*.
{% endhint %}

### The mod report

Guardian acted *without* a human, so a human must always get the final say — in one click. The report is a flag card (see the sample below) with the action line *"✅ 🗑️ message deleted · 🔇 author muted 1d"* and two buttons:

| Button        | What it does                                                                              |
| ------------- | ----------------------------------------------------------------------------------------- |
| **🔨 Ban**    | Guardian's call was right and the account is a throwaway — ban it from the community.     |
| **↩️ Revert** | Guardian was wrong — lift the mute immediately. The flag is recorded as a false positive. |

* **Discord:** the report is posted to your **flag-report channel** (chosen when you turn auto-moderation on; defaults to where you enabled Guardian) and **@mentions your mod role**. Nova **refuses** to arm auto-moderation until you've set a mod target — *"set the mod role to @Moderators"* — otherwise a false positive would leave a member muted with nobody able to undo it.
* **Telegram:** the report is **DM'd to each of your moderators** — not posted in the group, which would just show the scam to the very people it targeted. Nova DMs your nominated mods, or falls back to the group's own admins. **Mods must have opened a chat with Nova at least once** (Telegram forbids a bot from messaging first); if no mod is reachable, the report is posted in the group rather than lost.

***

## AI Scan Budget

Most messages are cleared for free by Guardian's instant pre-filters — the **AI intent verdict** only runs on the small minority that actually look risky. Those AI scans draw from a daily budget set by your community's plan:

|                           | **Free** | **Community PRO** |
| ------------------------- | -------- | ----------------- |
| **AI intent scans / day** | 20       | **500**           |

For context: 500 scans/day is far more than even a large, actively-attacked community typically needs — the pre-filters absorb the ordinary chatter, so only genuinely suspicious messages spend a scan.

**When the day's budget is spent, Guardian does not go dark.** Flood/repeat-spam protection, the deterministic spam short-circuits, and all the built-in heuristics keep running at full strength. Only the per-message AI intent analysis pauses until the budget resets at **00:00 UTC**. The first time it happens each day, Guardian posts a single heads-up to your mod-review channel so you're never wondering why the flags went quiet.

Check today's usage anytime — the status includes it:

> *"@Nova is Guardian on?"* → *"…AI scans today: 137/500."*

***

## What a Flag Looks Like

When Guardian flags a message, it posts a clean, scannable card to your mod-review channel — designed so a moderator can judge it at a glance.

{% tabs %}
{% tab title="Discord" %}
A **color-coded embed** — red for a likely scam, amber for "please review":

```
🚨 Nova Guardian
Likely Scam — Seed / wallet harvesting

Newcomer with zero contribution history is offering to buy
existing wallets with transaction history — a classic
seed-harvesting lure.

Confidence     Category                  Channel
95%            🔑 Seed / wallet harvest   #general

Member
Cafu · joined 6d ago
1510524330151510066

❝ Flagged message ❞
> I am looking to purchase a phantom wallet with a history
> of frequent transactions. I am willing to pay 3-4 SOL.

Action
🛡️ Flag only — enforcement off.

🔗 Jump to message
──────────────────────────────────────────────
Enforcement OFF · auto-delete & auto-mute disabled · Today at 8:48 AM
```

The verdict, confidence, category, **origin channel** and member sit in scannable rows; the flagged text is quarantined in a quote block; a **Jump to message** link takes you straight to it; and the enforcement status lives quietly in the footer. The channel mention keeps working even after the flagged message is deleted (when the jump link goes dead), so you always know where the spam landed.
{% endtab %}

{% tab title="Telegram" %}
Clean formatting with bold headers and an **expandable quote** for the flagged text:

```
🚨 Nova Guardian — Likely Scam

🔑 Seed / wallet harvesting · 95%

📍 Zero Authority DAO · topic 42

👤 Cafu · joined 6d ago
1510524330151510066

📝 Newcomer with zero contribution history is offering to buy
existing wallets with transaction history — a classic
seed-harvesting lure.

❝ I am looking to purchase a phantom wallet… ❞  (tap to expand)

🛡️ Flag only — enforcement off.
──────────
🔗 View flagged message
Enforcement OFF · auto-delete & auto-mute disabled
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
Guardian **never pings anyone** in a flag card — the quoted scam text can't tag or notify members — and it **never re-publishes a working lure**: links and bot handles in the quoted text are defanged for display (`https://scam.xyz` → `hxxps://scam[.]xyz`, `@handles` made non-clickable), so flagging a scam doesn't give its link a second audience. Mods still see exactly what the link was; the audit log keeps the original text. Cards are visible to whoever can see your mod-review channel.
{% endhint %}

***

## Set It Up — just talk to Nova

There's no wizard — you configure Guardian conversationally. (Setup is admin-only, in your server/group — not DMs.) @mention Nova or reply to it:

1. **Turn it on:** *"@Nova turn on Guardian / scam protection for this server"* — auto-moderation arms by default (Discord asks for a mod target first); add *"without auto-moderation"* for flag-only.
2. **Point it at a mod channel (optional):** *"send Guardian flags to #mod-review"* — defaults to the enable channel.
3. **Impersonation protection is automatic** — Guardian knows your real admins/mods from their server permissions. Optionally protect extra names or whitelist a permissionless helper: *"the real admins are @alice and @bob"* (also never flags them).
4. **Add allowlist phrases:** *"allowlist 'buying wallets' — this is an OTC channel"*
5. **Nominate your moderators** (required before auto-moderation arms on Discord): *"set the mod role to @Moderators"*
6. **Tune / check anytime:** *"make Guardian stricter"* · *"is Guardian on?"* · *"how many scams did Guardian catch this week?"* · *"turn off auto-moderation"* · *"turn off Guardian"*

**Bot permissions for auto-moderation.** On **Discord** Nova needs *Manage Messages*, *Moderate Members* and *Ban Members* (all in the current [invite link](/nova-bot/getting-started/adding-nova-to-discord.md)), and its role must sit **above** the members it moderates. On **Telegram**, make Nova an **admin** with *delete messages* and *ban/restrict users* rights. If a permission is missing, Guardian falls back to flag-only and tells you exactly what to grant. (Communities set up before auto-moderation existed keep their chosen mode; on Discord, re-invite Nova with the current link if the permissions are missing.)

***

## Configuration Options

| Option                               | What it does                                                                                                                                                                                     |
| ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Mode**                             | `strict` / `balanced` / `lenient` detection sensitivity                                                                                                                                          |
| **Auto-moderation**                  | The full module: delete + 24h mute + notify the author + a **Ban / Revert** report to your mods. **Default ON for new setups** (explicit opt-out: *"activate Guardian without auto-moderation"*) |
| **Mod-review / flag-report channel** | Where flag cards and auto-moderation reports are posted (defaults to where you enabled it)                                                                                                       |
| **Mute duration**                    | How long an auto-mute lasts (default **24 hours** with auto-moderation on)                                                                                                                       |
| **Auto-delete**                      | *(granular)* Remove scam messages but do nothing else — no mute, no author notice, no actionable report. Prefer auto-moderation.                                                                 |
| **Auto-mute**                        | *(granular)* Mute the author but do nothing else. Prefer auto-moderation.                                                                                                                        |
| **Allowlist**                        | Phrases that must never be flagged (normal-for-your-community talk)                                                                                                                              |
| **Known admins**                     | *(optional)* Real admins are detected automatically from permissions; use this to protect extra named identities or whitelist a trusted helper who has no permissions                            |
| **Thresholds**                       | Advanced numeric overrides for the suspicious / scam cut-offs                                                                                                                                    |

***

## Weekly Recap

Every **Monday at 12:00 UTC**, Guardian posts a short recap card to your mod-review channel: how many messages it flagged that week, how many were high-confidence scams, what actions it took, and the top patterns it saw. (It stays quiet on weeks with nothing to report.)

```
🛡️ Nova Guardian — Weekly recap

Guardian reviewed the week and flagged 12 messages — 5 high-confidence scams.

Flagged          Auto-actions
12 (5 scams)     🗑️ 3 · 🔇 2

Top patterns
🔑 Seed / wallet harvesting (4)
🎣 Fake hiring / recruiting (3)
📨 DM-bait / off-platform lure (2)

✅ Removed 3 messages and muted 2 members.
──────────────────────────────────────────────
Last 7 days
```

***

## Privacy Note

Guardian classifies messages **in real time** to reach a verdict. It only writes a record to its audit log when it actually **flags or acts on** a message (suspicious or scam) — that record includes the verdict, category, reason, action taken, and a short snippet of the offending message for the mod-review card. Ordinary clean messages are never stored by Guardian.
