Under the hood

How Danger Stamp is built.

For the technical crowd: the detection engine, the AI model we use, where it looks and how we kept it private.

139rulebooks, one per scam family
8,900+brands and banks it knows, including every US bank and credit union
27,878test cases, scams and the real pages they copy
1AI call at most per check, about 350 ms

01 · The engine

Facts in, a verdict out, in six steps

One small engine does all the deciding. It has no dependencies, and the exact same file runs in the extension, on our server and in our test harness, so the tests measure what users actually get.

  1. 1

    Read the facts

    Web address, title, visible text, form fields, sender, messages. Plus behaviour caught by page hooks: clipboard writes, full screen, wallet requests.

  2. 2

    Work out flags

    Is the domain new, on a free host, a look-alike of a brand? Does the sender use free email? Look-alike letters are folded to plain ones first.

  3. 3

    Pick rulebooks

    Each rulebook has a trigger. Only the ones whose trigger fires are in play. On an ordinary page, none fire and nothing else happens.

  4. 4

    Instant rules

    Some scams give themselves away by behaviour alone. 33 rulebooks can stamp from code, with no AI call at all.

  5. 5

    One AI call

    Every rulebook still in play adds its yes/no questions to a single call. One call, however many rulebooks fired.

  6. 6

    Decide

    Each rulebook turns the answers into stamp, warn, badge or nothing. The loudest action wins, then the highest score.

In your browserOn our server, only when needed

Rulebooks

A rulebook is a small JSON file: a trigger, a few plain yes/no questions, rules that combine the answers, and the stamp and warning to show. There are 139, from fake logins and wallet drainers to fake escrow sites and bot comment threads. Most stamp; a few only warn or add a small badge.

Triggers are written in a tiny condition language over signals, flags, keywords, fields, domain age and surface, with negation and nested any-of groups.

What the extension knows, and what it doesn't

The extension carries a stripped copy of every rulebook: the trigger, the instant rules and the stamp. The questions, thresholds and decision rules stay on the server. The build fails if a question ever ends up in the extension bundle.

That keeps the part scammers would love to read out of their hands, and it means the extension can update its rulebooks without a new release.

Brands and domains

988 hand-picked brands with 1,719 official domains, plus every US bank and credit union from the FDIC and NCUA registers (7,949 more). Bank names are only trusted in a page's title, login box or logo, because names like "First State Bank" are everywhere.

Brands are often shown only as a picture, so on password, card and code pages the extension also reads the logo: the words attached to it, where it loads from, and a fingerprint of its shape and colour. 18,553 real-logo fingerprints cover 7,608 brands. It all happens in the browser, and in a test on 1,000 ordinary websites only one was taken for another brand.

Domain age comes from RDAP, and it is only looked up when a rulebook in play actually needs it.

02 · The AI

A model that answers yes or no, fast

We use Jev, the "System One" model from TypeSafe. You don't chat with it. You send it a state and typed questions, and it sends back numbers: the probability that each statement is true.

That suits scam detection. Our questions are literal and narrow, like "the web address is not owned by the brand this page shows". Code handles what the model is bad at, and the model handles what code is bad at: reading a page the way a person would.

It is quick and cheap enough to run on every risky page: about 350 ms and about a hundredth of a cent per call at list price. Our test cache holds 0+ answers.

What we learned about asking it

  • One condition per question. Two ideas in one question blur the answer.
  • It reads words literally. "Only" and "the company that charged you" change the answer, so wording matters.
  • Don't ask it to count, date or spell-check. Look-alike domains, domain age and counts are done in code and passed in as flags.
  • Less is more. Padding the page text lowers accuracy, so what we send is trimmed hard.
  • Batches blur. Single comments go five to a call, each question pinned to its own item. Threads and ads go one at a time.
  • Scores drift a little between runs, so no test case is allowed to sit close to a threshold.

03 · Where it looks

Pages, inboxes, chats, comments, posts and ads

Websites

Every page, with 25+ detectors for form fields and page shape: password, card, CVV, one-time code, recovery-phrase grids, ID uploads, countdown timers, fake browser windows, tiny recurring-charge text.

Webmail

Gmail, Outlook, Yahoo and AOL, plus Roundcube, SnappyMail and other webmail. Sender, reply-to, look-alike sender domains and free-email "companies". Events opened in Google Calendar are read like emails.

Chats

WhatsApp Web, Messenger, Facebook Marketplace, Instagram DMs, Discord and Telegram Web.

Comments

YouTube comments and reply threads, plus page-wide bot-network signals: the same "mentor" or coin pushed by several accounts.

Feed ads

Sponsored posts on Facebook and Instagram, found by their "Sponsored" label in five languages.

Social posts

Posts on X, LinkedIn and Reddit get a soft AI SLOP tag when they read like AI-written filler. Code counts the tell-tale signs, then one Jev question with examples decides.

Behaviour

Hooks that run before the page does: clipboard writes that look like commands, full screen without a click, keyboard lock, exit traps, notification prompts and wallet approval requests. Risky downloads are paused while the page is checked.

04 · Privacy

Built so there's nothing to leak

The extension reads everything you read. That's the only way it can work. So the design goal was simple: send as little as possible, and save none of it.

Gated in the browser

A page is only sent when a sensitive signal fires or a rulebook's trigger matches locally. Emails and chats go through a stricter gate. A message is only sent when the other person's words match money, payment, code, urgency or link patterns, or when it's the first few messages of a chat from someone you haven't talked to (an unsaved contact you haven't replied to). Emails are also sent when Gmail itself flags them, when they use your own name as the sender (only a yes or no leaves the page, never your name or address), when they're one big picture, when they contain a QR code, or when they contain explicit adult-lure words. An event you open in Google Calendar is treated like an email, with the same gate. Chats with people you know never leave your computer unless they show those warning signs. You can also select or paste a message and choose "Check this": then only that text is sent, redacted the same way.

Redacted before sending

Card numbers, SSN and SIN-shaped numbers, account and routing numbers, page titles and other people's email addresses are blanked out in the browser. For another person's address on a free webmail service, only the service is kept ("[email]@gmail.com"), because "write to me at a Gmail address instead" is a scam sign; company addresses are blanked completely. Web addresses are sent without their query strings. For an email's links, only the site name each link goes to is sent (paypal.com, never the page or anything after it), because a link that goes somewhere the sender doesn't own is the strongest sign of phishing. QR codes in an email are read on your computer, and again only the site name the code leads to is sent. On document sites like SharePoint, Notion or Canva, only the site names a page's links go to are sent, and for a sign-in window drawn inside a page, only the address its fake address bar shows. Phone numbers are kept on purpose, because the number is the scam. In a chat, the name or number shown for the other person is sent too, because "a stranger wrote first" is part of the check.

A stateless server

The server keeps no database of checks and logs nothing about them. Its only log line for a check is a credit counter. Email and chat checks can be switched off entirely, and browser tests prove that normal emails and switched-off inboxes are never sent. "Ask someone you trust" is built on your computer and Danger Stamp never sends it. Reporting a false alarm is opt-in and sends only the web address without its query.

Who sees anything, and what they get: your browser sees everything you read; the Danger Stamp server sees a redacted excerpt, judges it and forgets it; Jev at TypeSafe sees that same redacted excerpt; and rdap.org sees only the domain name, to check how old it is.

05 · Research and testing

Where the scams come from, and how we prove it works

Research

Scam videos, scam-baiter channels, social comments, the Facebook Ad Library, Reddit, live phishing feeds, the top 50,000 real sites, every US bank and credit union, and fraud agency lists.

How we research scams →

Testing

27,878 test cases, holdouts we never tune on, and 17 blind rounds of fresh real scams, each item labelled by two independent AI labellers who never see the rulebooks. Every round is published, including the bad ones.

Every test result →

Seen a scam we missed?

Reports become test cases, and missed scam families become new rulebooks.

Report a scam