Phishing-URL research, done in the open

Read a link the way your browser does.

Phishing links hide their real destination in plain sight. SecureMind Labs builds and openly evaluates URL-based phishing detection—starting with a dissector that shows what a link actually points to, right in your browser.

  • Runs entirely in your browser
  • The link is never opened or sent
  • Facts about the URL, not verdicts
Example Fictional link, dissected
Subdomain
Any text its owner chooses—including a bank’s name.
Registrable domain
account-review.co is who you would really be visiting.
“northwind-bank.com” appears in the link, but only as a subdomain of account-review.co. The dissector finds this using the Public Suffix List—the same list browsers use.

URL dissector

Paste a link. See where it really goes.

The dissector parses the address with your browser’s own URL rules and lists what is factually true about it. It does not visit the site, look it up, or decide whether it is safe.

Up to 2,048 characters. Processed only on this page; nothing is stored or sent.

Try an example:

Nothing dissected yet.

Results show the link split into its parts, the domain you would actually be visiting, and any structural tricks such as hidden login text, disguised IP addresses or look-alike characters.

How it works

Three layers, kept deliberately separate.

A single “safe/unsafe” badge hides what each step actually knows. We keep parsing, prediction and judgement apart, and label where each one runs and how mature it is.

  1. 1

    Structure

    What the URL literally says: the scheme, the real registrable domain, ports, embedded credentials, encodings.

    Where
    This page, in your browser
    Status
    Live
    Cannot tell you
    Who runs the site or what it does
  2. 2

    Model

    A machine-learning estimate of whether the URL text resembles known phishing links.

    Where
    A separate research demo (Streamlit)
    Status
    Research
    Cannot tell you
    Anything reliable about brand-new campaigns

    The demo runs our original prototype, trained on synthetic URLs. The 98.6% accuracy it displays comes from an offline benchmark on a different dataset, not from the model that analyses your link, and a “safe” result from it is not reliable. Treat it as a demonstration only.

  3. 3

    Judgement

    Context no URL contains: who sent it, whether you expected it, what the page asks you to do.

    Where
    You
    Status
    Always needed
    Cannot be replaced by
    Layers 1 or 2

Research

What we learned trying to break our own evaluation.

We trained URL-only models on two public, CC BY 4.0 datasets, split by registrable domain so no site appears in both training and testing.

Finding 01

One dataset gave the answer away.

In PhiUSIIL, every legitimate URL is a bare https://www. homepage. A one-line rule with no machine learning scores F1 0.996 there—and 0.754 on a different dataset. So our main model reads only the hostname.

Finding 02

Strong on familiar data, much weaker on new data.

Our hostname model ranks phishing above legitimate URLs clearly better than the current demo model on held-out domains (ROC-AUC 0.919 vs 0.848). On a dataset it never saw, the gap nearly closes (0.741 vs 0.717), and its false-positive rate rose from 0.9% to 6.2%.

ROC-AUC on two test sets, hostname candidate versus current demo model Dots show ROC-AUC; horizontal bars show 95% intervals from a domain-level bootstrap. The same values are listed in the table below the chart. 0.5 0.6 0.7 0.8 0.9 1.0 PhiUSIIL · held-out domains 0.919 0.848 Hannousse · different dataset 0.741 0.717 ROC-AUC (0.5 = chance)
Hostname candidate (experimental) Current demo model Bars: 95% intervals, resampling whole domains.
ROC-AUC with 95% intervals
Test setHostname candidateCurrent demo model
PhiUSIIL, held-out domains0.919 (0.902–0.935)0.848 (0.809–0.885)
Hannousse, different dataset0.741 (0.711–0.772)0.717 (0.682–0.751)

Read the full evaluation, limitations and method

About

An early-stage, independent research project.

SecureMind Labs grew out of a machine-learning course project (ECE 569A) at the University of Victoria. It is now an independent effort to build phishing-URL tools whose limits are as visible as their results: every number on this site comes from a reproducible evaluation, and every capability is labelled with where it runs and how mature it is.

We are not a commercial security service, and nothing here should be the only thing standing between you and a suspicious link.

Contact
[email protected]
Source code
github.com/Chandhu03/AI-Based-Phishing-URL-Detection (opens in a new tab)