WUBBERY SUBSTRATE ENGINE & COMPANY / SUB-2.35μs ENTERPRISE ROUTER

WUBBERY on the rubber Earth

Intelligence in motion.

WHAT IS WUBBERY? — SUB-2.35μs AI SUBSTRATE ENGINE

WUBBERY is the world's first sub-2.35μs AI Substrate Engine. Powered by the Harkin Theorem (8 patent families), WUBBERY sits between applications and AI models to route 425,000 decisions/sec, slash data centre megawatt power draw by 88%, and eliminate token cost waste on CPUs and GPUs.

Choose the right intelligenceKeep the useful contextPrepare the next move
00

Let's talk about our future

Any data centre not routing like this — do they care about the Earth?

Harsh question. Fair question. AI's power draw is the defining infrastructure problem of this decade, and most of it is pure waste: premium models answering requests a small model handles perfectly. The waste isn't theoretical — it's metered, billable, and entirely optional.
CHAPTER 00 · THE ENVIRONMENTAL STORY

The Data Centre Grid Crisis vs WUBBERY Substrate Salvation

A two-act comparison of how unrouted cloud AI scorches electrical grids vs how WUBBERY on-device WASM preserves Earth.

−88% MEGAWATT POWER REDUCTION
ACT 1: THE CRISIS (Cloud SaaS Overload)PARCHED PLANET
Act 1: Dead Earth scorched by unrouted cloud data centers

Data Centre Grid Scarcity & Latency Lag

Unrouted AI models send simple queries to distant GPU server farms. Every request consumes megawatts of cooling electricity, adds 2,400ms of lag, and inflates cloud bills.

Cloud Latency Penalty:2,400ms Lag
Facility Energy Waste:High Grid Power
Cloud SaaS Token Fee:$0.06 - $0.30 / turn
ACT 2: THE SALVATION (WUBBERY Substrate)SUSTAINABLE FUTURE
Act 2: Vibrant Happy Earth holding WUBBERY banner on clean white background

On-Device Zero-Marginal-Cost Execution

WUBBERY runs WebAssembly bytecode natively on the user's browser or edge CPU. Simple intent queries resolve in 2.35μs with 0% cloud server overhead, keeping Earth green.

Routing Latency:2.35μs Instant
Cloud Server Load:0% Cloud Waste
On-Device Execution Cost:$0.00 Token Fee
Monthly Environmental Calculator: 50,000 Monthly Turns
Grid Power Saved:100.0 kWh / mo
CO2 Emissions Avoided:40.0 kg CO2
Trees Saved:~2 Trees / yr
01

Initial module set

Eight modules first.
Fifteen layers all up.

0.00 µs

per routing decision

425,000 decisions/sec on one CPU thread. Other routers spend an AI call to route AI calls. This is infrastructure.

0.0 %

on 34,734 real requests

9.68B tokens of the founder’s own usage, shadow-replayed and priced three ways. Nothing simulated.

+0.0 %

memory lift · ambiguous regime

BEIR fiqa, the field’s own metric — while scanning 19.7× fewer candidates.

0.0 %

within-one-tier accuracy

Held-out routing suite. Misses printed, not hidden. Re-runnable in front of you.

0 %

predictive prefetch hit-rate

On structured usage — and ~3% on random. We publish where it does NOT help.

0 %

inference energy

Versus all-premium baselines. The routing dividend, measured in watts as well as dollars.

0

modular layers · released in stages

WUBBERY is a 15-layer modular platform. The initial release brings the six modules shown here, plus Governance and Anomaly Detection. The remaining seven layers will be introduced in stages.

02

Against the field

The industry's best, at their best. Then ours.

Every router below decides which model should answer each request. The published numbers are theirs; the framing is charitable — these are genuinely the fastest approaches the field has.

Test V/01 · published vs live-measured

Decision latency per request · lower is better

EV · Embedding semantic router

5ms

the field’s "fast" tierimpressive — five thousandths of a second

RL · RouteLLM classifier

84ms

the published state of the arta model routing models — research-grade

ND · Not Diamond

100ms

the venture-backed leader"worthwhile tradeoff for the savings"

W · ROUTING MODULE — measured live, one CPU thread

0.00235ms

2.35 microseconds. The bar is three pixels wide because it has to be visible at all.

Bars drawn on a log scale, or every competitor would be off the page.

Now ours

0×

faster than the published state of the art

42,553× faster than the leader's best case · 2,127× faster than the field's "fast" tier.

Test V/02 · same numbers, true scale

NOT DIAMOND
100ms best case
ROUTELLM
84ms published
EMBEDDING
~5ms
WUBBERY
0.00235ms — 1/42,000th of the leader's

At true scale our bar is thinner than anything this screen can draw. That's what infrastructure looks like next to an AI call. Theirs need a GPU-backed service; ours runs air-gapped on anything.

Fair fight: their figures are their own published/reported decision overheads; ours is measured live on a laptop CPU — run GET /v1/bench/overhead and watch your own number.

03

Memory · public data

Three benchmarks. Three domains. The lift scales with ambiguity.

BEIR datasets, dense-retrieval baseline, nDCG@10 — their data, their metric. The ambiguity prediction was written down before fiqa was run. It held.
+12.5%

SCIFACT

science · unambiguous

13.6× less work
+9.9%

NFCORPUS

medicine · hard

12.4× less work
+61.2%

FIQA

finance · ambiguous

19.7× less work

Falsify this

The boundary, published on purpose: with no state signal the routing module gates to zero influence — identical to baseline to four decimal places, on all three datasets. Set β=0 and the lift must vanish. Set noise=1.0 and the routing module must equal baseline. It does.

Test B/01–03 · third-party harness · run 2026-07-10 · seed 7 · deterministic

04

Cost · real traffic

One machine. Every request. Three worlds, priced.

The founder's complete Claude Code history replayed through the router. Cache pricing identical in all worlds. This card claims cost only — quality lives in 01 and 03.

Test S/01 · live-measured logs

34,734 requests · 9.68B tokens · 3,079 sessions

All-premium · naive ceiling

$14,792

Actual spend · as logged

$13,017

Module-routed · shadow

$8,734
−32.9%

module-routed vs actual spend, on the same 34,734 requests

Test S/02 · pick what you run today

Premium

Fable-5 / GPT-5-class

Balanced

Sonnet-5 / GPT-4o-class

Budget

Haiku / Llama-class

−42.4%

projected cost

89.8%

of requests

94.4%

tier-hold accuracy

Hard requests keep the best model — the saving never comes out of quality.

05

Coding agents · Fable 5 baseline, priced both ways

Claude Code on Fable 5. Same session, 80.7% cheaper.

Without routing · every turn on Claude Fable 5

$6.41

With routing · right-sized model per turn

$1.24

The top-tier baseline, not a strawman — hard turns stay capable, no quality cliff.

One session · every turn, routed

read a file · balancedrename a symbol · fastfix a typo · balancednull check · balanceddiagnose flaky test · balancedformat file · fastcommit message · balanceddocstring · balancedbig refactor (reasoned) · balancedexplain a regex · balancedadd a unit test · balancedsecurity review · balancedclassify a ticket · bulkimplement rate limiter · balancedlist files · fastsummarise a diff · balanced

An Anthropic request never routes off-family. Per-model prices are representative published tier rates — the dollars are illustrative; the tier mix and the rate are measured. Prove it on your own prompts at zero spend with X-LEWM-Dry-Run: 1.

06

What this means for you

Six dials, translated into your world.

2.35µs

Your users never feel the router.

Every other approach adds 5–200ms to every single request. Ours adds nothing a human — or an SLA — can detect.

−32.9%

A third of your AI bill, back.

Measured on real mixed traffic, not a marketing mix. Hard requests still get the best model.

+61.2%

Better answers where it’s hardest.

Ambiguous questions — the ones your support and search actually struggle with — is exactly where the memory lift is largest.

94.4%

Savings without the quality tax.

Routing accuracy means the cheap model only gets what the cheap model can do. That is the whole trick, held to a published number.

94%

Feels faster than it has any right to.

Prefetch means the answer is often warming up before the request lands — 2.7× faster with use, on structured workloads.

−71%

Your GPUs go 3.4× further.

If you run in-house models, energy IS capacity: the same rack serves multiples of the traffic — or the same traffic on a third of the power.

+ more

and that's only what we show. Eight patent-pending families, each holding more than one claim. The rest is ledgered too — it just needs an NDA to see.

07

Modular by design

Choose the layers your business needs.

WUBBERY has 15 layers in total, designed to work independently or together. The initial release combines the six modules shown here with Governance and Anomaly Detection. The other seven will follow in stages.

ROUTE

2.35µs

RECALL

+61.2%

PREFETCH

94% warm

TIER HELD

94.4%

LEDGER

−32.9%

ENERGY

−71%

GOVERNANCE

INITIAL

ANOMALY

INITIAL

Because the platform is modular, a business can start with the layers it needs and add more over time. When selected modules run together, the right answer can be ready before the request arrives — 100% of the time in our ledgered test, versus 0% when the same capabilities run as disconnected point-tools. Customers get a staged path into the platform without giving up the compounding benefit.

08

How it actually plugs in

No rewrite. No SDK. One address change.

Each selected module can run entirely on your side — laptop, server, or air-gapped rack. Your apps don't integrate with it; they simply talk through it. Three ways in, in order of commitment:
A

Watch first — shadow mode.

Zero risk, zero code. We replay your existing logs and show the ledger of what you would have saved. This is the free tier, and it’s how every engagement starts.

$ replay --dir your-logs/

your number, zero spend

B

Route live — change one address.

Every app that calls AI has a setting for where to send its requests. Today it points at your provider. Change that one setting to the routing module’s local address — your keys, your providers, your data path. One settings line, no code.

$ OPENAI_BASE_URL=http://localhost:8787/v1

that’s the whole integration

C

Ask before you spend — the decision API.

Keep full control: your code asks "which model should this request use?" and gets an answer in microseconds, then makes its own call. For in-house fleets, this turns −71% energy into 3.4× capacity.

$ POST /v1/route {"text": "..."}

sealed decision, nothing leaves your machine

What a company rollout actually looks like · one engineer · hours, not months

DAY 1

Shadow your history.

Your platform team exports existing AI logs. We replay and hand back a ledger: what you spent vs what you would have. Nothing installed, nothing leaves your side.

WEEK 1

One workload goes live.

Pick something low-stakes — an internal tool, a summarisation job. One engineer changes its provider address. The ledger starts measuring for real. Rollback is changing the address back.

WEEK 2–4

Expand team by team.

Each workload is its own one-line switch, each shows its own ledger. Security reads the audit log; finance reads the savings. Stop any workload any time.

ENTERPRISE

Inside your walls.

For regulated environments selected WUBBERY modules install on your own servers — your IT, your network rules, air-gapped if needed. Per-workload calibration and sealed capabilities come with the NDA engagement.

GET YOUR NUMBER

Shadow access is free and private — tell us roughly what your traffic looks like and we'll set you up. No card, no commitment, no data leaves your side.

Request module access →
09

Security & privacy

Built like it's guarding something. Because it is.

The platform's protected mathematics are patent-pending trade secrets — so every module is engineered never to leak anything, including your data. The same seal that protects us protects you.

Your prompts stay yours

Decisions happen on your machine in microseconds. In decision mode nothing leaves your network. In gateway mode requests go to your providers with your keys — the platform stores none of it.

No phone home

The service binds to localhost and refuses the network by default. Exposing it is an explicit operator act that hard-requires its own token — and TLS in front.

Keys pass through, never stored

Provider credentials ride the request and are forwarded only to the fixed upstream you configured. No key ever touches disk.

Every decision auditable

An append-only audit log records every route, gateway call, and failed auth — outcomes only, timestamped, safe to ship to your SIEM.

Hardened by default

Timing-safe token checks, brute-force lockout (10 failures → 5-minute ban), per-tier rate limits, request-size caps. On by default, not opt-in.

Nothing to supply-chain

The sealed service has zero third-party dependencies — no packages to poison, no CVE surface to babysit. And it runs fully air-gapped: banks, health, and government keep everything inside the perimeter.

The bank case — your machines, improved

Regulated industries can't ship prompts to a routing API — so no other router can even bid. WUBBERY installs on your own hardware, inside your own perimeter, and makes the machines you already own better: 3.4× effective capacity from the energy dial, answers 2.7× faster with use, a third off any external spend — with the audit log your compliance team already wants. Nothing new to trust. Your metal, upgraded.

AIR-GAP READYZERO-DEPENDENCY BUILDAUDIT-LOG NATIVEGDPR-ALIGNED BY DESIGNSOC 2 · PATHWAYISO 27001 · PATHWAY

Badges marked pathway are in progress, not held — certification lands with the first enterprise pilots. The architecture was built for their audits from day one.

The seal

The seal works both ways: responses carry the model, tier, and outcome scores — never the internal state that produced them. Reverse-engineering the service tells you what it decided, not how. Your data can't leak what we never transmit.

10

At scale

For data centres and big business, this is an energy story.

Routing isn't a niche optimisation — at fleet scale it changes the physics of the business. When every request runs on the cheapest model that can genuinely handle it, the waste comes straight off the power meter.

For data centres

Your bottleneck is power, not silicon. We hand you power back.

  • 3.4× the inference from the same grid connection. −71% inference energy isn't a saving line — it's headroom. The megawatts you already have suddenly serve more than three times the billable load.
  • Cooling follows compute. Less wasted compute means less heat: PUE improves without touching the plant room.
  • Deferred capex. The next hall, the next substation, the next grid negotiation — pushed out years, because the building you have does more.
  • Grid-constrained sites win. Where you can't get more power, efficiency is the only growth path. This is that path.

For big business

A third of the AI bill, back on the P&L — not on a roadmap.

  • Measured, not promised. The same ledger you watched in shadow becomes the invoice basis — savings are counted from your own traffic before anyone is paid.
  • Procurement leverage. When any workload can take any model, no single provider owns you. Negotiate from the ledger.
  • Sovereignty kept. Runs inside your perimeter, air-gapped if required — regulators, auditors, and security teams all read the same audit log.
  • Capacity planning flips. Buy hardware for today's real load, not 3.4× of it. Growth rides the efficiency curve instead of the purchase order.

−71%

inference energy — off the meter, at fleet scale

×3.4

effective capacity from the hardware you already own

0

new racks, new grid connections, or new code required

100%

auditable — every routed decision, timestamped

11

Your number

What would it mean for you? Move the dials.

You don't need to know your token counts. Pick what you look like, and we project from our measured rates. Then run shadow on your real logs for your true number — free.

GPUs serving inference

1 GPU → serves like 3

What that means for a single GPU

One card suddenly serves the traffic of three and a half — the energy dial turns −71% inference energy straight into headroom. No new hardware, no bigger power bill: the same GPU you already own takes 3.4× the users, or the same users with the fans barely spinning. And routing costs it nothing — decisions run on the CPU in microseconds, so the GPU only ever does model work.

×3.4

effective capacity — −71% inference energy: same racks, same power budget, or the same load on a third of the power

2.7×

faster with use — measurably self-improving on structured workloads: the longer it runs, the quicker it gets

0 µs

GPU spent on routing — decisions run on CPU in microseconds: the router never queues behind your models

Projected from measured rates on our ledgered runs. Your mix will differ — shadow mode replays your own logs and shows your true number before you spend a cent.

12

Module Access

Request the layers your business needs first.

WUBBERY is not taking payment for a finished OS. The first eight modules are being released in stages, public pricing is not final, and access is discussed module by module.
Efficiency ModulesStaged Access

Request

pricing confirmed with scope

Modules and release group

Routing · Memory · Prefetch · Tiering · Ledger · Energy

  • Choose individual modules
  • Discuss your workload and environment
  • Confirm availability before onboarding
  • No combined-OS requirement
Request module access →
Trust ModulesStaged Access

Request

pricing confirmed with scope

Modules and release group

Governance · Anomaly Detection

  • Scoped permissions and policy controls
  • Approval and audit requirements
  • Operational anomaly visibility
  • Security requirements assessed first
Request module access →
Future ModulesRoadmap

Register

interest in the next releases

Modules and release group

Seven further layers · Fifteen total

  • Tell us what capability is missing
  • Join the relevant release list
  • Request more than one module
  • No purchase or account commitment
Register module interest →
13

Accessible proof

Don't take our word for it. Open the outputs.

The modules are not publicly downloadable yet, so the site does not offer commands that cannot run. The published aggregate results and methodology are available now.
OPEN JSON

structured aggregate outputs

DOWNLOAD CSV

every published benchmark row

READ METHOD

scope, method, privacy limits and caveats

Ready when you are

Start with shadow — free, private, your own logs, your true number. Or come straight to an NDA briefing for the sealed dials.

14

The rest of WUBBERY

One modular platform. Nine ways in.

This page shows the initial module set and its receipts. The rest of the site shows where those layers can be applied — every lane below is its own front door: