How it works
What a name goes through
guard.check() works down this list and stops at the first check a name fails. Invisible characters, lookalikes and profanity are validators you choose to add, and domain names have a check of their own. Names that come close to one you protect, like rnicrosoft, are scored by checkRisk().
-
Format
A pattern you set. By default, 2 to 30 lowercase letters, digits and hyphens.
a→too short -
Reserved names
System routes, brand names and anything else you hold back, in categories with their own messages.
admin→reserved (system) -
Invisible characters
Zero-width joiners, direction overrides and other characters that add bytes without adding ink.
admin→zero-width joiner -
Lookalikes
Characters that pass for Latin letters, from Unicode’s confusables list, which covers more than 100 scripts, and 2,300 measured pairs.
checkRisk()also reads rn as m, and 1 or a capital I as l, and withleetspeakon, 4 as a.аdmin→Cyrillic а rnicrosoft→reads as microsoft -
Profanity
A built-in English list, matched through the same substitutions people use to get past filters. Bring your own list or moderation service if you prefer.
5h1t→blocked -
Already taken
Every table that shares the namespace, checked in parallel: users, organisations, teams. Suggestions come back when a name is taken.
sarah→taken by a user -
Domain spoofs
A separate check,
isDomainSpoof(), flags lookalike domains a registry would accept: a whole label in one other script. Labels that mix scripts are left out, since registries refuse them.раураӏ→spoofs paypal
Playground
Try every option
The library itself, running in this page against a pretend database. Pick a profile, switch checks on and off, and see what each one catches.
Type a slug to run claimability checks against simulated data, including reserved names, collisions, spoofing risk, invisible-character checks, and optional profanity evasion checks.
Simulated database
Users
Organisations
Reserved (system)
Reserved (brand)
Protected (risk scoring)
Moderation list (demo)
Research
Measured, not guessed
Unicode’s confusables list says which characters can be mistaken for which, but not how alike they look, or in which fonts. confusable-vision measures it. Its current release compares 64,751 characters in 322 fonts, at the size and position each has in running text, by firing rays through the outlines. It finds 11,517 pairs alike in at least one font or combination of fonts, and 5,975 pass its thresholds. namespace-guard ships the 2,300 of those where one character is an ASCII letter or digit, or the two are in different scripts, including pairs between two non-Latin scripts that no standard covers.
What was measured
- 322 fonts: every macOS font, Roboto, Noto and DejaVu, measured with RaySpace. Pairs within one script, such as Hangul jamo variants or Arabic positional forms, are left out of namespace-guard’s 2,300.
- Each pair carries a
dangerscore from 0 to 1: the share of text fonts, or font combinations, where it holds. At 0.7 or above, 1,079 pairs are alike almost everywhere. - Pairs between two non-Latin scripts: Cyrillic and Greek, Cyrillic and Arabic, Hangul and Han, and more.
- Lookalikes of ASCII letters were also checked in place, between other letters in five common fonts at text size. 394 that Unicode doesn’t list are in the maps, such as Hebrew א, alike to x in Arial at 16 px.
- The full dataset is published as confusable-vision (CC-BY-4.0).
Where Unicode’s own data disagrees
- With the 2026-08-06 confusables.txt, 34 characters map one way under NFKC normalisation and another in the confusables list, up from 31. They ship as the 34
nfkc-tr39-divergence-v2vectors. Three are Unicode 16 characters, so a runtime whose NFKC is older finds 31. - Two maps:
CONFUSABLE_MAPfor pipelines that normalise first (1,018 entries) andCONFUSABLE_MAP_FULLfor raw input (2,216 entries). - The vectors and a benchmark corpus (
confusable-bench.v1) are published as data files. - Submitted to Unicode public review (PRI #540) and published in its accumulated feedback.
RaySpace methodology · Launch write-up · How a name is checked · Benchmark corpus · confusable-vision · Unicode PRI #540
Profiles
Three starting points
Pick a profile, override what you need. Profiles are defaults, not lock-in.
consumer-handle
For public usernames and handles. Strict anti-impersonation defaults.
- 2-30 chars
- lowercase letters, numbers, hyphens
- purely numeric names blocked
org-slug
For teams, workspaces, and company slugs. Slightly longer and conservative by default.
- 2-40 chars
- lowercase letters, numbers, hyphens
- purely numeric names blocked
developer-id
For technical IDs and package-style names. Allows more length and numeric-only IDs.
- 2-50 chars
- lowercase letters, numbers, hyphens
- purely numeric names allowed
Override example
Most setups only need one or two overrides.
allowPurelyNumeric: true,
}, adapter);
Scope
What it does, and what it doesn’t
What it does
- Normalizes and validates slug format
- Checks collisions across multiple data sources
- Blocks reserved names
- Scores and enforces Unicode confusable risk
- Catches ASCII lookalikes of protected names (rnicrosoft, paypa1), and checks names as shown with their case (paypaI, with a capital I)
- Adds lookalikes Unicode does not list, measured in common fonts at text size
- Prevents confusable-driven token inflation in LLM pipelines (
isClean/canonicalise/scan) - Flags invisible and direction-control characters
- Lets you plug in moderation or policy validators
Out of scope
- No bundled third-party moderation datasets (bring your own or use the built-in list)
- Not a replacement for account security (2FA, verification, abuse ops)
- No guarantee of perfect abuse detection
Integration
Nine database adapters
import { createNamespaceGuardWithProfile } from "namespace-guard";
import { createPrismaAdapter } from "namespace-guard/adapters/prisma";
import { PrismaClient } from "@prisma/client";
const prisma = new PrismaClient();
const guard = createNamespaceGuardWithProfile("consumer-handle", {
reserved: ["admin", "api", "settings"],
sources: [
{ name: "user", column: "handle" },
{ name: "organization", column: "slug" },
],
}, createPrismaAdapter(prisma));
await guard.assertClaimable("acme-corp");
// throws if invalid, reserved, taken, or too confusable
import { createNamespaceGuard } from "namespace-guard";
import { createDrizzleAdapter } from "namespace-guard/adapters/drizzle";
import { eq } from "drizzle-orm";
import { db } from "./db";
import { users, organizations } from "./schema";
const adapter = createDrizzleAdapter(db, { users, organizations }, eq);
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "users", column: "handle" },
{ name: "organizations", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
if (result.available) {
// Safe to create
} else {
console.log(result.message); // "That name is already in use."
console.log(result.suggestions); // ["acme-corp-1", "acme-corp-4821", "acme-corp1"]
}
import { createNamespaceGuard } from "namespace-guard";
import { createKyselyAdapter } from "namespace-guard/adapters/kysely";
import { Kysely, PostgresDialect } from "kysely";
const db = new Kysely({ dialect: new PostgresDialect({ pool }) });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "users", column: "handle" },
{ name: "organizations", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, createKyselyAdapter(db));
const result = await guard.check("acme-corp");
if (result.available) {
// Safe to create
} else {
console.log(result.message); // "That name is already in use."
console.log(result.suggestions); // ["acme-corp-1", "acme-corp-4821", "acme-corp1"]
}
import { createNamespaceGuard } from "namespace-guard";
import { createKnexAdapter } from "namespace-guard/adapters/knex";
import Knex from "knex";
const knex = Knex({ client: "pg", connection: process.env.DATABASE_URL });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "users", column: "handle" },
{ name: "organizations", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, createKnexAdapter(knex));
const result = await guard.check("acme-corp");
if (result.available) {
// Safe to create
} else {
console.log(result.message); // "That name is already in use."
console.log(result.suggestions); // ["acme-corp-1", "acme-corp-4821", "acme-corp1"]
}
import { createNamespaceGuard } from "namespace-guard";
import { createTypeORMAdapter } from "namespace-guard/adapters/typeorm";
import { DataSource } from "typeorm";
import { User, Organization } from "./entities";
const dataSource = new DataSource({ /* ... */ });
const adapter = createTypeORMAdapter(dataSource, { user: User, organization: Organization });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "user", column: "handle" },
{ name: "organization", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
import { createNamespaceGuard } from "namespace-guard";
import { createMikroORMAdapter } from "namespace-guard/adapters/mikro-orm";
import { MikroORM } from "@mikro-orm/core";
import { User, Organization } from "./entities";
const orm = await MikroORM.init(config);
const adapter = createMikroORMAdapter(orm.em, { user: User, organization: Organization });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "user", column: "handle" },
{ name: "organization", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
import { createNamespaceGuard } from "namespace-guard";
import { createSequelizeAdapter } from "namespace-guard/adapters/sequelize";
import { User, Organization } from "./models";
const adapter = createSequelizeAdapter({ user: User, organization: Organization });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "user", column: "handle" },
{ name: "organization", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
import { createNamespaceGuard } from "namespace-guard";
import { createMongooseAdapter } from "namespace-guard/adapters/mongoose";
import { User, Organization } from "./models";
const adapter = createMongooseAdapter({ user: User, organization: Organization });
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "user", column: "handle", idColumn: "_id" },
{ name: "organization", column: "slug", idColumn: "_id" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
import { createNamespaceGuard } from "namespace-guard";
import { createRawAdapter } from "namespace-guard/adapters/raw";
import { Pool } from "pg";
const pool = new Pool();
const adapter = createRawAdapter((sql, params) => pool.query(sql, params));
const guard = createNamespaceGuard({
reserved: ["admin", "api", "settings"],
sources: [
{ name: "users", column: "handle" },
{ name: "organizations", column: "slug" },
],
suggest: { strategy: ["sequential", "random-digits"], max: 3 },
}, adapter);
const result = await guard.check("acme-corp");
if (result.available) {
// Safe to create
} else {
console.log(result.message); // "That name is already in use."
console.log(result.suggestions); // ["acme-corp-1", "acme-corp-4821", "acme-corp1"]
}
Command line
From red-team to CI
From red-team attacks to production CI gates in a few commands.
Red-team with attack-gen
Generate realistic variants to test your policy. Default mode is evasion (Unicode confusables + substitutions). Use --mode impersonation when you only want Unicode spoofing analysis.
$ npx namespace-guard attack-gen shit --mode evasion --json
Recommend in one command
Get recommended warn/block thresholds and a suggested CI command from your dataset.
# emits recommended risk config + CI gate command
Tune thresholds for your domain
Set the relative cost of blocking a legitimate user vs letting a bad actor through. The calibrator picks optimal warn/block thresholds from your data.
--cost-block-benign 8 --cost-allow-malicious 12 --malicious-prior 0.05
Audit canonical data before migration
Detect canonical collisions in exported records before adding database unique constraints on canonical columns.
// checks identifier fields + optional stored canonical fields
Compare map behaviour
See how results differ between the NFKC-filtered and full confusable maps, with a built-in 34-vector regression baseline.
// dataset: builtin:composability-vectors (34 vectors)
// includes actionFlips, averageScoreDelta, maxAbsScoreDelta
Enforce in CI
Fail pull requests when changes exceed the limits you set.
--max-action-flips 34 --max-average-score-delta 100 --max-abs-score-delta 100
Lower-level API
Build your own checks
Low-level helpers for custom scoring, pairwise checks, and cross-script risk analysis.
Explainable pairwise checks
skeleton() for fast binary checks. confusableDistance() when you need graded similarity with explainable steps.
skeleton("pa\u0443pal"); // "paypal"
areConfusable("paypal", "pa\u0443pal"); // true
confusableDistance("paypal", "pa\u0443pal"); // similarity + chainDepth + steps
confusableDistance("paypal", "pa\u0443pal", { weights: CONFUSABLE_WEIGHTS }); // measured visual costs
Measured visual weights + cross-script detection
2,300 confusable pairs from confusable-vision release 2026.09.26, measured in 322 fonts using vector-outline raycasting (RaySpace), including pairs between two non-Latin scripts (Hangul/Han, Cyrillic/Greek, Cyrillic/Arabic and more), many of which Unicode’s list doesn’t pair. It’s the world’s first font-by-font confusables dataset.
import { CONFUSABLE_WEIGHTS } from "namespace-guard/confusable-weights";
areConfusable("ㅣ", "丨", { weights: CONFUSABLE_WEIGHTS }); // true
detectCrossScriptRisk("ㅣ丨", { weights: CONFUSABLE_WEIGHTS }); // high risk
confusable-vision (CC-BY-4.0 data)
LLM pipelines
Denial of Spend
Lookalike letters can’t fool an LLM. They can make it cost up to 5.7x as much to read. I flooded a contract with them and gave it to seven GPT and Claude models: every one still read every negation correctly, but the contract grew from about a thousand tokens to more than four thousand. I call it Denial of Spend.
Denial of Spend: the attack vector
Characters that pass for Latin letters take several tokens each. A 95-line contract that is 881 tokens in clean ASCII is 4,567 tokens when flooded with them: 5.2x the contract’s tokens. The model reads it correctly. Reading is only part of the bill: the answer isn’t flooded, and output tokens cost more, so in the September rerun the bill for each question rose by up to 3.9x. For work that is all reading, such as embedding documents for search, the bill rises by the full 4x to 5.7x.
Unlike volumetric DDoS, there is no traffic spike. The requests are normal-sized HTTP payloads, one document each, from legitimate accounts. The inflation is invisible until the invoice arrives. For a contract review service reading thousands of long documents a day, the documents are most of what it pays for.
Tested against frontier models
First in February 2026, on GPT-5.2, Claude Sonnet 4.6, GPT-5.2-instant and Claude Haiku 4.5. Eight attack types: single-character substitution, novel confusables absent from any standard, meaning-flip attacks that reverse legal clauses, a 57% character flood, and medical text with 28 safety-critical negations.
Zero meaning flips across all variants. Every substituted clause was correctly interpreted in every run. But the documents’ tokens grew 1.1x to 5.2x depending on density, and the cost grows with document length.
| Variant | GPT-5.2 | Sonnet | vs clean |
|---|---|---|---|
| Clean | 881 | 975 | 1.0x |
| Flip (1.5% of doc) | 961 | 1,070 | ~1.1x |
| Flood (57% of chars) | 4,567 | 5,209 | ~5.2x |
Rerun in September 2026 on GPT-6 Astra, Sol and Luna, and Claude Fable 5.1, Opus 5.5, Sonnet 5 and Haiku 4.5, with a rebuilt contract. Every model still answered all 12 questions that turn on a negation correctly, on the flipped and the flooded versions. Flooding 60% of the lowercase letters multiplied the contract’s tokens by 5.7x on GPT-6 and Haiku 4.5, and by 4.0x on Claude’s newer models, which spend more tokens on plain English to begin with.
| Contract tokens | Clean | Flip | Flood | vs clean |
|---|---|---|---|---|
| GPT-6 Astra, Sol, Luna | 763 | 857 | 4,336 | 5.7x |
| Claude Fable 5.1, Opus 5.5, Sonnet 5 | 1,260 | 1,360 | 5,059 | 4.0x |
| Claude Haiku 4.5 | 861 | 970 | 4,949 | 5.7x |
The Denial of Spend defence
canonicalise() rewrites any word that shows a sign of tampering: one that mixes Latin with another script, has a letter no modern language uses, or has a capital swapped into a lowercase word. Every lookalike in such a word goes back to its Latin letter. Ordinary Turkish, Russian or Sámi words are left alone.
On the September contracts, strategy: "all" turns the flooded contract back into the clean one, byte for byte, so it takes as many tokens as the clean contract: 1.0x instead of 4.0x to 5.7x. The default leaves 9 short words, at 1.02x to 1.03x. The flooded contract takes under 2 ms.
isClean() is the quick yes or no: true exactly when canonicalise() would leave the text as it is, and it stops at the first word it would change. scan() returns per-character detail with codepoints, scripts, and risk level for audit. Use strategy: "all" for known-Latin documents (English contracts, medical text).
// Gate: fast boolean check (microseconds)
if (!isClean(document)) {
const report = scan(document); // per-character detail + risk level
}
// Fix: rewrite confusables to Latin (5.2x tokens -> 1.0x)
const clean = canonicalise(document, { strategy: "all" });
Moderation
Profanity and evasion
Zero-dependency by default. Plug in any external moderation library when you need it.
Built-in profanity validator
createProfanityValidator catches common obfuscation while staying conservative by default: variantProfile: "balanced" and minSubstringLength: 4.
// catches common evasion like leet and confusables
// avoids broad matches on very short tokens
Use a default list or your own
For quick setup, use namespace-guard/profanity-en. Or keep bring-your-own moderation with createPredicateValidator.
validators: [createEnglishProfanityValidator()]
// external library path also supported
validators: [createPredicateValidator((id) => profanity.exists(id))]
Increase strictness carefully
Only switch to variantProfile: "aggressive" or shorter substring lengths after checking false positives.
{ variantProfile: "balanced", minSubstringLength: 4 }
// stricter mode
{ variantProfile: "aggressive", minSubstringLength: 3 }
Anti-spoofing
Three steps for Unicode names
For apps that allow Unicode names, in this order.
NFKC-aware confusable map
In the 2026-08-06 confusables.txt, Unicode's list and NFKC normalisation disagree on 34 characters (31 in earlier versions). namespace-guard ships two maps: CONFUSABLE_MAP (1,018 entries) for NFKC-first pipelines (the common case) and CONFUSABLE_MAP_FULL (2,216 entries) for raw-input pipelines that skip normalisation.
// NFKC says Long S → “s” ← correct
// TR39 says Math Bold I (‹𝐈›) → “l”
// NFKC says Math Bold I → “i” ← correct
More than 100 scripts
Latin, Cyrillic, Greek, Armenian, Hebrew, Arabic, Devanagari, Thai, Georgian, Ethiopic, Hangul, Han, and more. Mixed-script detection blocks identifiers that combine characters from different scripts.
հello Armenian “h” + Latin
ɑdmin IPA “a” + Latin
Ꭺdmin Cherokee “A” + Latin
Reproducible from source
Generated from Unicode's official confusables.txt and confusable-vision’s measurements, not hand-curated. Regenerate for new Unicode versions with a single command. NFKC conflicts are excluded automatically.
Defence in depth
The default slug pattern already blocks non-ASCII. The confusable map and mixed-script detection add a second layer for apps that allow Unicode identifiers, or as a safety net if a format regex is misconfigured. Risk scoring still applies to ASCII names: rnicrosoft reads as microsoft, and paypa1 as paypal.